Researchers release a 2,001-question spoken benchmark for Telugu

single source· 1 articles · confidence: medium · first seen 2026-09-16 20:00 UTC

What this means for you

If you build Telugu voice interfaces, there is now a public test set to check against. Read the scoring caveat first: open-weight judges marked correct answers wrong when the wording differed from the reference, so score by hand, or with a stricter judge, before trusting a number. For everyone else, nothing to act on.

A Telugu spoken question-answering benchmark was posted to arXiv on 16 September: 2,001 factoid question–answer pairs across six domains, 2.53 hours of audio, bilingual transcriptions and human-verified reference answers. The paper's second contribution is an audit of automatic scoring. Gemini-as-a-judge — a model grading answers in place of a person — tracked human ratings most closely but was unevenly strict; open-weight judges penalised correct Telugu answers that differed in surface form from the reference. The authors report that speech input introduces phonetic confusions that change a question's meaning, and that errors compound through speech recognition and translation. The benchmark is public; the paper is a preprint.

Key facts

  • ·VākQA contains 2,001 Telugu factoid question–answer pairs across six domains. source
  • ·The benchmark includes 2.53 hours of speech audio, bilingual transcriptions and human-verified reference answers. source
  • ·Gemini-as-a-judge best approximated human ratings but was non-uniformly strict, per the paper. source
  • ·Open-weight judges systematically penalised correct Telugu answers whose surface form differed from the reference. source
  • ·The benchmark is publicly released; the paper was posted to arXiv on 16 September 2026. source

What the sources say

  • Hugging Face Daily Papers — Single preprint introducing a Telugu spoken QA set and auditing whether automatic judges agree with people.

Sources

The original reporting. Follow these — they did the work.

← the wire