Frontier models answer fewer than half of real business-intelligence questions

single source· 1 articles · confidence: high · first seen 2026-09-15 20:00 UTC

What this means for you

Nothing to act on. No API, no weights, no pricing, and no model names to compare against — the frontier scores are reported without saying which systems produced them. For anyone building a BI assistant on top of Power BI or Tableau, the finding to carry is that unaided language models get fewer than half of these dashboard questions right.

Researchers posted BI-Bench, described as the first benchmark for end-to-end business intelligence — finding the right tables, transforming them and joining them before a business question can be answered. Its questions and answers were extracted by hand from real user dashboards. The authors report frontier language models score below 50% accuracy; they name no models and give no evaluation date. Their BI-Agent splits the job into search, join and transform steps and calls specialised data tools, gaining up to 40 percentage points over plain models and up to 30 more after post-training on synthesised examples of real BI work.

Key facts

  • ·BI-Bench is described by the paper's authors as the first benchmark for end-to-end business intelligence, covering table search, transformation and joins as well as the final answer. source
  • ·Frontier large language models score below 50% accuracy on BI-Bench, with no model names and no evaluation date given. source
  • ·The tool-augmented BI-Agent improves accuracy by up to 40 percentage points over the vanilla LLMs it was compared with. source
  • ·Post-training BI-Agent with supervised fine-tuning and reinforcement learning on synthesised trajectories from real BI projects yields further gains of up to 30 points. source
  • ·Benchmark questions and ground-truth answers were harvested from public real-world BI projects and extracted by hand from user dashboards. source
  • ·The paper is arXiv 2609.20886, posted 15 September 2026. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire