The catalog

AYA Bench Attestation

A public benchmark badge that cannot lie.

A public benchmark badge that cannot lie.

AYA Bench Attestation puts an "AYA-Verified" number on your site, and here is the entire trick: the badge is issued by a fail-closed gate, not by a marketing department. If the run does not survive the gate, there is no badge. Not a smaller badge, not an asterisked badge. No badge. That refusal is what makes the badge worth displaying.

Get your number attested Read how the gate works

THE PROBLEM

Nobody believes your benchmark number, and honestly, nobody should. The typical published score was produced by a harness that fabricates success the moment it hits a tool it does not recognize, against a baseline quietly running a weaker model, graded by the same model being tested. The industry knows: OpenAI walked away from SWE-bench when the score stopped meaning anything, and hundred-point cherry-picking is documented sport. So every vendor's number gets discounted on sight, which means the vendor with a TRUE number has no way to make it count. And from August 2, 2026, EU AI Act obligations give capability claims regulatory weight, so an unverifiable number stops being just embarrassing and starts being a liability.

HOW IT WORKS

The badge is the last step of a pipeline that refuses at every earlier step.

- One sanctioned runner. Your number is produced only by our honest-benchmark runner. No side-channel scores, no numbers you hand us, no exceptions. - The harness fails closed. An unknown or unavailable tool returns failure, never fabricated success. Any stub-mode run contaminates the score and permanently disqualifies it. - One model, both arms. Baseline and treatment share the same pinned model, so a lift is a real lift, not a model swap dressed as progress. - Qualifying benchmarks only. Execution-grounded (the answer runs against a real database) or externally scored. Self-graded percentages are ineligible, full stop. - The gate decides. No number publishes unless the claim gate is GREEN: strict mode on, fair baseline confirmed, zero fabrication detected. A missing field blocks. A RED gate issues nothing, and your fee does not buy a different outcome. - The badge expires. Attestations are quarterly audits. A failed or stale re-audit pulls the badge automatically. A number from last year does not get to pretend it is current.

WHAT YOU GET

An AYA-Verified number with the verdict attached: the score, the benchmark, the model, the run date, the re-audit date, and the gate's GREEN, publicly checkable. Usually it is a smaller number than the one you were about to publish. It is also the only one in your category that survives a skeptic, which is the entire point.

WHO IT'S FOR

AI companies whose benchmark claims face someone who will check: an investor's technical diligence, an enterprise buyer's eval team, a regulator. And on the demand side, platforms and buyers who want vendors pre-attested before the procurement decision. If you are shopping for the biggest possible number, we are genuinely not for you; the gate will refuse it, and we will let it.

PRICING (the ladder)

Every number below is a hypothesis we validate with you, not a commitment. - Attestation run, fixed price: we run your system through the sanctioned pipeline, and if the gate is GREEN, issue the badge. A RED run is reported to you privately and issues nothing. - Quarterly attestation, a standing subscription: re-audit runs keep the badge current. - Attestation program, for platforms and marketplaces that want their listed vendors badged, negotiated.

IP is licensed, never assigned; the gate and the mark stay ours. Payment is gated, human-confirmed, and never influences the verdict. The fee buys the run, not what it finds.

THE PROOF (the origin story is a confession)

We built this gate because our own harness was lying to us. We caught it fabricating a jump from zero to ninety-five percent on tools it never ran, and a twenty-point lift that was really a weaker baseline. We shipped the gate instead of the number, and every AYA eval now runs through it. We steer on determinism and cost-per-solved-task, not leaderboards. So we are not selling cleverness, we are selling the fix to a mistake we made and refuse to repeat, and the first badge this program issues will be on our own system, gate-GREEN, before we ever issue yours.

HONEST NOTE

Straight status: the runner and the fail-closed gate are real, code-resident, and exercised on AYA's own evals today. What does NOT exist yet: the badge itself. No badge has ever been issued, to us or to anyone; the public verification page, the signed badge asset, and the revocation service are declared, unbuilt work; and live runs wait on funded eval keys and a booted mesh. We will not issue customer badge number one before we issue our own, and ours has to survive the same gate. If that sequencing costs us your urgency, we accept the trade. A badge issued faster than that would be exactly the kind of number this product exists to kill.

Get your number attested

*This page is a specification. The capability it describes is not built yet, and nothing here is a claim that it runs today.*