Council Post: ​Are You Being Benchmaxxed?

Katy Wigdahl, CEO of Speechmatics, drives voice AI innovation, implementation and ethical deployment with finance and scaling expertise.

getty

Choosing an AI platform has never been harder. The difficulty of that decision comes not only from the number of new models entering the market. The benchmarks buyers rely on to compare them are also becoming less trustworthy.

Earlier this year, we ran an internal hackathon. The brief was deliberately provocative: Game an AI benchmark the same way vendors do, and see how far it gets you. We climbed 10 places with only a few hours of work and one retrain, placing alongside models from other participant companies.

What I took from the exercise had nothing to do with rankings. It was about what the process revealed: that optimizing for a benchmark and actually improving a product have become two very different activities. In AI, these two activities diverged some time ago.

There's even a term emerging for it: benchmaxxing, optimizing for the test rather than meaningfully improving the product.

The Leaderboard Is A Well-Known Game

The problem is structural; the moment a test set becomes the standard measure of quality, it also becomes the optimization target. Teams fine-tune on data that resembles the test conditions, select inference settings tuned for evaluation and submit. Often, they will even train on the test data itself, a practice that purists frown upon for good reason.

The headline number improves, but whether the model has gotten better at the underlying task is a different question.

A 2025 study found that giving developers even limited additional access to test data could boost leaderboard scores by up to 112%. A February 2026 paper found that nearly half of widely used benchmarks have hit saturation, meaning top models score so similarly the tests can no longer meaningfully separate them.

Vendors are not being dishonest. They are optimizing rationally for the metric that drives procurement decisions. That is the problem.

What The Hackathon Revealed

The ranking was not the interesting part of our experiment. What surfaced were the structural limits of any evaluation run outside your own conditions.

The score improved, but when we fine-tuned even further, the model broke in ways the leaderboard would never have shown. Error rates more than doubled at a later checkpoint, meaning the submission would have been fine, but a live deployment would not have been.

We also discovered a feature that existed in the product but wasn't actually working. It was configured but simply not connected, and the hackathon only surfaced it because we were stress-testing the pipeline end to end rather than trusting the spec sheet. Before someone dug in and fixed it, there was no effect whatsoever; after, it produced a meaningful improvement.

That kind of issue never shows up in a benchmark. They measure what a model does under test conditions, not whether the product actually does what it claims to do in practice.

The Market Is Specializing Faster Than Buyers Are Adapting

The AI model market has changed. A year ago, the dominant question was which foundation model was the most capable overall. That is no longer the right question because the answer changes too quickly and the performance gap between top models is narrowing.

Models are now competing to become the best model for a specific use case rather than the best model in the abstract. Speed, accuracy, cost, reasoning ability, multilingual performance and vertical domain expertise are all being optimized differently by different vendors. The headline benchmark is often a score the vendor is most confident winning on. But this often leaves out which user populations were excluded, which test conditions were not disclosed and which use cases never made it into the evaluation set.

A global enterprise workforce is not a controlled evaluation environment, and a contact center with regional accents, background noise and overlapping speakers looks nothing like a vendor demo. The gap is wide enough that an entire category of companies now exists purely to help buyers evaluate AI systems independently, a market that barely existed three years ago.

That gap between what benchmarks measure and what your business actually requires is where most procurement decisions go wrong.

Your Data Is The Only Benchmark That Counts

The next competitive advantage will not come from tracking leaderboards. It will come from building the internal discipline to test AI against your own data, your own users and your own definition of success before someone else's benchmark makes the decision for you.

That means running products and models against your actual conditions rather than the vendor's curated demonstration. It means investing the time up front to build your own test sets from real data, real users and real failure cases—work that will return more than it costs, often significantly more.

The hackathon reinforced something I already suspected: There is no universal winner in AI; there is only the model that performs best for your specific use case. The only way to identify it is to test it under your own conditions, against your own data and definition of success.​​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?