Benchmark tests whether financial chart models' buy and sell calls track their evidence

single source· 1 articles · confidence: medium · first seen 2026-09-12 20:00 UTC

What this means for you

If you evaluate finance models, a single hallucination score will not surface this: the paper's lowest-ranked model on its coverage-aware metric issued a direction on only 6.4% of questions. Code and data are public, so the harness can be rerun; there is no leaderboard or hosted API.

Researchers have published E2A-Bench, a 969-query test set for vision-language models that read financial charts (models taking an image as well as text). It scores 20 of them on whether a stated conclusion traces back to the chart, across three input formats built from 323 HS300 constituents. One model placed near the bottom on the coverage-aware metric because it committed to a direction on only 6.4% of questions — a failure its unsupported-claim score did not show. Financial fine-tuning raised buy-to-sell ratios by 4.21 to 4.68 times. Code and data are public; no evaluation date is given.

Key facts

  • ·E2A-Bench contains 969 queries constructed from 323 HS300 constituents across three input modalities. source
  • ·The benchmark evaluates 20 vision-language models on four metrics: UCR, RCI, ECI and NDR. source
  • ·The lowest-UCR model ranks near the bottom by NDR because of 6.4% directional coverage. source
  • ·Financial fine-tuning amplifies the BUY:SELL ratio by factors of 4.21 to 4.68 across strict base-fine-tuned pairs. source
  • ·NDR measures coverage-aware evidence-to-action reliability, not realized trading performance. source
  • ·Code and data are released at github.com/wanng-ide/E2A-Bench. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire