Publish the number that holds you to account
Drafted with Akora, reviewed and edited by Lee Norvall. We disclose this on every post where Akora helped — see our authoring approach.
Most infrastructure vendors publish the benchmarks that flatter them. The more useful discipline is shipping the measurement a customer would use to prove you wrong.
Every infrastructure vendor publishes performance numbers. Almost none publish the instrument you would need to check them.
That asymmetry is the whole game. A benchmark in a blog post is a claim about a machine you cannot see, running a workload you did not choose, at a moment you cannot reproduce. It is marketing wearing a lab coat. The reader's only options are belief or indifference, and most sensibly choose indifference.
We have tried to build the opposite habit, and it is worth writing down because it has cost us something — which is how you know it is real.
Ship the instrument, not just the reading
Every response from our gateway carries a header — x-air-overhead-ms — reporting the
gateway's own processing time for that specific request, excluding the model
provider.
It is a small thing that changes the shape of the relationship. You do not have to trust our latency claims, or run a synthetic benchmark, or take our word about a percentile. You read the number off your own traffic, in production, on the requests you actually care about. If we get slower, you find out before we tell you.
We are not aware of another gateway that hands you the number you would use to hold it to account. We would like to be wrong about that, because the practice should be normal.
It cost us a claim, which is the point
The discipline earned its keep in the least comfortable way available.
An early internal measurement suggested our gateway was roughly 2× slower on tool-bearing turns — the kind of finding that, in most organisations, gets quietly re-run until it goes away. Because the per-request overhead was published in the response, it could be decomposed rather than argued about. The gateway's own cost was 2–9 milliseconds. The tail was the model provider being slower when handed 59 tools — on both arms of the test, including the one that never touched our infrastructure.
The claim was wrong, and it was withdrawn. Not because someone was persuasive in a meeting, but because the instrument made the question decidable.
That is the actual argument for measuring this way. It is not that it makes you look good. It is that it makes you correctable — and an organisation that can correct itself internally is the only kind that can be trusted externally.
Publish the caveats at the same size as the number
Our current figures put gateway overhead in the single-digit milliseconds on a 59-tool, 63 KB request, and time-to-first-token through the gateway measured marginally faster than going direct — because our cloud's peering to the provider more than pays for the extra hop.
Now the caveats, which belong in the same paragraph rather than a footnote: that is one machine, one domestic connection, a small sample, and our own benchmark document marks the figures provisional. The defensible claim is the shape — parity with going direct, single-digit-millisecond overhead — not the exact milliseconds. Anyone quoting those specific numbers back at us as a guarantee is quoting them wrong, and we would rather say so than enjoy the ambiguity.
We also decline to publish a cost-saving percentage, despite having cache figures that would make a good headline. Our production usage ledger is currently around two-thirds test traffic, and any saving computed from it unfiltered would be fiction. When we have a clean number we will publish it unprompted.
Independence, where you can get it
One more habit worth naming. The parity measurements above were run by a different team inside the business — the client team, as an internal consumer of the gateway — not by the team that built it. Against a real upstream provider. No mocked responses.
That is not third-party validation and we will not call it that. It is one step better than marking your own homework, and the protocol is published so an actual third party can reproduce it, or falsify it, against their own workload.
Why this is a security argument, not a performance one
It would be easy to read this as a post about latency. It is not.
The organisations we build for are being asked to put AI infrastructure between their staff and their data, and to accept a set of claims about cost, privacy and reliability that they largely cannot verify. The industry's answer to that has been trust badges and benchmark charts.
We think the answer is instruments. A claim a customer can falsify is worth more than a claim they have to believe — and if that principle is sound for latency, it is sound for everything else we say. It is the same reason our telemetry is built so that it cannot contain prompt content, rather than promising that it does not.
Ask your vendors for the instrument, not the number. The ones who have thought about this will know exactly what you mean.
- benchmarks
- measurement
- transparency
- engineering-culture