Why we publish our misses
A results table with misses beats a page with no table. Disclose the flaws, including the ones in your own methodology.
Most AI product pages show one of two things: no numbers at all, or numbers so good they'd make a statistician nervous. We publish our misses. When an evaluation round comes back below target, that goes in the table — with the target next to it, marked plainly as a miss. No hiding it, no re-running until it looks better.
This isn't humility as marketing. It's the only evaluation posture that survives contact with reality.
Why it matters
A page with no table asks for trust and offers no way to verify it. A page with a perfect table asks for trust and insults your intelligence — everyone in the field knows what real eval numbers look like, and they're never perfect. A table with misses is the only one a serious engineer believes, because it's the only one that looks like actual measurement.
There's a second reason, and it's the more important one: publishing misses is what makes them get fixed. A miss in a public table is a commitment. It creates the pressure that turns 'we should improve data quality' from a backlog item into a roadmap item. Private misses get deprioritized. Public misses get worked on.
And then there's methodology honesty — the part almost nobody does. When we found flaws in our own benchmark scoring, we disclosed them alongside the numbers. A benchmark with a known flaw, disclosed, is useful: the reader can discount appropriately. A benchmark with a hidden flaw is misinformation with error bars.
What the fix looks like
For us, it's three habits. First, every number ships with its target and its measurement conditions — what was measured, on what data, with what method. A number without conditions is a rumor. Second, misses are labeled as misses, in the same table, in the same type size. No footnotes, no 'challenges we're excited about.' Third, when the methodology has a known flaw, say so in the same breath as the number. The reader deserves to know what they're discounting.
The objection is always 'but competitors will use it against us.' They might. But the customers worth having — the ones who run their own evals, who know what a miss means and what it takes to fix it — read an honest table as competence, not weakness. Anyone can claim perfection. Publishing the miss with a plan to fix it is the stronger signal.