Monday, 16 June 2025
29.3 C
Singapore
28.2 C
Thailand
20.1 C
Indonesia
28.7 C
Philippines

Did xAI mislead the public about Grok 3’s benchmarks?

xAI is under scrutiny for allegedly misleading AI benchmark results, with OpenAI employees questioning its claims about Grok 3’s performance.

Debates over AI performance benchmarks—and how they are presented—have sparked controversy in the tech world.

xAI accused of misrepresenting Grok 3’s performance

An OpenAI employee has accused Elon Musk’s AI company, xAI, of publishing misleading benchmark results for its latest AI model, Grok 3. Igor Babushkin, a co-founder of xAI, has strongly defended the company’s claims, insisting that the data was reported accurately. However, as is often the case with AI benchmarks, the truth appears more complex.

The controversy centres around a post on xAI’s blog, where the company shared a graph displaying Grok 3’s performance on AIME 2025. This benchmark consists of difficult maths questions from a recent invitational mathematics exam. While some experts question AIME’s suitability as a true measure of AI ability, it is still widely used to assess AI models’ maths skills.

In the graph, xAI claimed that two versions of Grok 3—Grok 3 Reasoning Beta and Grok 3 mini Reasoning—outperformed OpenAI’s best available model, o3-mini-high, on AIME 2025. However, OpenAI employees quickly pointed out that xAI failed to include an important detail: o3-mini-high’s score at “cons@64.”

The missing metric: What is cons@64?

The term “cons@64” stands for “consensus@64.” This means an AI model is given 64 chances to answer each problem, and its most frequently chosen answers are considered final. This method can significantly improve a model’s benchmark scores, making it a crucial factor in evaluating AI performance. Omitting this detail in a comparison graph can create a misleading impression that one model is superior when that may not be the case.

When looking at AIME 2025 results at “@1” (which only considers the model’s first attempt at answering a question), both Grok 3 Reasoning Beta and Grok 3 mini Reasoning scored lower than OpenAI’s o3-mini-high. Grok 3 Reasoning Beta also slightly underperformed compared to OpenAI’s o1 model running at a “medium” computing setting. However, xAI has marketed Grok 3 as the “world’s smartest AI.”

The broader issue of AI benchmarks

In response to the criticism, Babushkin argued on X (formerly Twitter) that OpenAI has also presented benchmark results in ways that could be considered misleading. However, OpenAI’s past comparisons mainly involved its models rather than direct competition with other companies.

An independent AI researcher compiled a more comprehensive chart showing nearly every model’s performance at cons@64 to provide a clearer picture. This chart provided a more accurate comparison but highlighted another issue: the computational cost behind these results.

As AI researcher Nathan Lambert pointed out, one of the most critical factors remains unknown—the amount of computing power and money required for each model to achieve its highest score. This raises a larger concern about AI benchmarks in general. While they can provide useful insights, they often fail to fully capture a model’s true strengths and weaknesses, leaving room for misinterpretation.

Hot this week

Amazon taps nuclear power to boost AWS cloud energy supply

Amazon signs a 1.92 GW nuclear energy deal with Talen to power AWS cloud and explore new small modular reactors in Pennsylvania.

Apple to end macOS updates for Intel Macs after 2025

Apple says that MacOS 26 will be the final update for Intel Macs, ending new feature support and keeping security updates until around 2028.

Resident Evil Requiem returns to Raccoon City with new story and hero, coming February 2026

Resident Evil Requiem, which launches on February 27, 2026, takes you back to Raccoon City with a new lead and chilling story.

Meta in talks to invest over US$10 billion in Scale AI

Meta may invest over US$10B in Scale AI, marking one of the biggest private AI funding deals and Meta’s largest external AI investment ever.

Redmagic 10S Pro launches in Singapore with faster gaming performance and exclusive offers

Redmagic 10S Pro lands in Singapore with overclocked performance, S$270 early bird deals, and a free cooling fan for a limited time.

Informatica deepens partnership with Databricks to support new Iceberg and OLTP services

Informatica joins Databricks as launch partner for new Iceberg and OLTP solutions, introducing AI tools to speed up GenAI development.

Hong Kong opens skies to larger drones in bid to grow low-altitude economy

Hong Kong will allow the testing of larger drones to boost its low-altitude economy and improve logistics, following mainland China's lead.

Hong Kong to build new AI supercomputing centre in bid to lead global tech race

Hong Kong plans a new AI supercomputing centre to boost its tech hub status and support growing start-ups across the Greater Bay Area.

Steam adds full native support for Apple Silicon Macs

Steam runs natively on Apple Silicon Macs, ditching Rosetta 2 for smoother performance and better gaming on M1 and M2 devices.

Related Articles

Popular Categories