Every week, a new AI model ships with a scorecard.
The headline reads something like "3% better than the previous model on the standard tests." My LinkedIn feed fills with reaction posts, "this is the new state of the art", and capability comparisons that rearrange the same four company names in a different order.
I have stopped reading them.
Not because the scorecards are wrong. They measure real things. But what they measure is not what determines whether your AI will actually work in your business.
After 18 months of running production AI deployments in healthcare, legal, and insurance, those scorecards have never once predicted which model my customers preferred.
That's because the scorecard is answering a different question than the customer is actually asking.
The public scorecards measure how smart a model is on standard tests. The customer is asking whether the model will work on her data, in her workflow, the day a regulator walks in. Almost no team has stopped chasing the first to start measuring the second.
The four things I actually look at instead
1. Does the AI know when it doesn't know? This is the most underrated quality in any AI you put into a business that matters. An AI that's right 92% of the time and confident every single time is more dangerous than an AI that's right 85% of the time but says "I'm not sure, please double-check" on the borderline ones. In a hospital, that "I'm not sure" is the entire reason the deployment is safe. The public scorecards don't measure this at all. We do, every time.
2. Does the AI behave the same way after the next model update? Every few weeks, the company that makes the AI ships an update. A "better" model. The trouble is, the updated model often answers slightly differently than the one it replaced. Sometimes it's more cautious. Sometimes it's less. Sometimes its tone changes. Sometimes it picks a different option on a borderline case. To a customer who signed a contract based on how the AI behaved last month, "slightly different" is not an improvement. It's a problem. We measure how much the AI's behavior changes between updates. If it changes too much, we don't ship the update.
3. Does the AI get things right on your specific data? The public scorecards test AI on widely-known information, encyclopedia knowledge, common code patterns, school math. Your business doesn't run on that. Your business runs on your specific drug list, your specific case law, your specific underwriting rules. The model that wins on the public test may be confidently wrong on the things that actually matter to you. We test on the customer's own data, in the customer's own files, before we ever ship.
4. What does it cost when the AI is wrong? The scorecards treat every right answer and every wrong answer as the same weight. In a real business, especially a regulated one, one wrong answer can cost 100 times what a right answer earns. A model that's right 94% of the time but catastrophically wrong 6% of the time is worse than a model that's right 89% of the time and appropriately says "I don't know" the other 11%. We measure the cost of being wrong, not just the rate of being right. That changes the answer of which model is "best."
Why this matters
If a vendor is leading their pitch with leaderboard rankings, they are telling you they don't know what your real buyer actually needs.
If your engineering team is celebrating a 3% improvement on a public scorecard, they are optimizing for a number that does not correlate with whether your customer renews the contract at the 12-month mark.
If your CIO is comparing vendors on those public rankings, she is comparing them on the wrong axis. She doesn't know it yet, but she will at her first audit, at the first behavior change after a model update, or the first time a vendor can't explain in plain English why the AI gave a confidently wrong answer on a specific case.
The shift from "how capable is the AI" to "how safe is the AI in my business" has already happened in the buyer's mind. The vendors who haven't caught up are still pitching against last year's evaluation criteria, and losing deals they don't realize they're in.
The vendors leading with leaderboards are talking past the only buyer with budget. The deal is being scored on a different set of numbers, and the teams that build for them are about to compound for a decade.
Three things to do instead
1. Build your own test list that looks nothing like the public scorecards. Your business doesn't care about general capability. It cares about your workflows, your data, your decisions. Build a test list with your operations team, your compliance team, and your most experienced staff. That list, not the public ones, is the bar an AI vendor has to clear.
2. Measure how the AI behaves across model updates, not just during the demo. The vendor who passes your test once and then quietly drifts every time their underlying model upgrades is a vendor you will have to renegotiate with every release. The vendor who guarantees the AI will behave the same way after every update is a vendor you can sign a multi-year contract with.
3. Make the pitch about the cost of being wrong, not the rate of being right. If the vendor cannot tell you, in plain English, what happens when their AI is wrong, who catches it, how fast, with what consequences, they have not designed a product. They have designed a demo.
Final statement
The last time a customer asked me about a public benchmark score was eight months ago.
I get questions about our test list every week.
I get questions about behavior consistency every other week.
I get questions about explaining decisions to a regulator on every single deployment.
The vendors who hear those questions and build for them, win.
The vendors still leading with leaderboards are talking past their buyer.
Stop reading the leaderboards. Start measuring what your business actually pays for.
