Audio reasoning benchmark reports 92%
An Artificial Analysis report records a 92% score on Big Bench Audio for an audio model and compares time to first token with and without reasoning.
Source: Gemini 2.5 Native Audio Thinking no Big Bench Audio (x.com). Text prepared with AI from this source.
What happened and what to do
In a report published on October 13, 2025, Artificial Analysis says Gemini 2.5 Native Audio Thinking scored 92% on Big Bench Audio, a benchmark of 1,000 audio questions adapted from Big Bench Hard. The report gives average time to first token as 3.87 seconds for the reasoning version and 0.63 seconds for the version without reasoning.
For a company assessing voice interfaces, this comparison points to running its own tests on representative business tasks, measuring both quality and latency. A controlled pipeline for collecting interactions, response metrics, and dashboards can track accuracy, wait time, and failures before the company decides how to integrate audio into its processes.
How the consultancy can help
Wendelmaques can diagnose voice use cases and their quality and latency requirements, scope an evaluation, and implement the collection, testing, and dashboards needed to monitor the system in operation.
Next step
Send a short description of your audio use case to receive a scoped proposal for diagnosis and implementation.
Consulting for your project
Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.
Quoted per project
Request a proposal