The catalog

Verify-My-Research-AI

Grade your research AI, and see exactly where it misread the paper.

Grade your research AI, and see exactly where it misread the paper.

Verify-My-Research-AI sits behind your existing research assistant, independently reads the actual source paper, and tells you whether each claim is really in there, where it overstated or fabricated, and what the paper actually says. Not a second opinion about the first opinion, a read of the source itself.

Book a verification pilot See a sample disagreement report

THE PROBLEM

Your research AI summarizes a paper, pulls out a finding, and hands you a claim that reads clean. Most of the time it holds up. But sometimes it says the paper found something the paper never claimed, it quotes a number that is not in the text, it drops the one limitation that changes the whole meaning, or it confidently cites a result from a figure that does not exist. You do not know which claims are the good ones and which are the confident-but-wrong ones, because checking would mean re-reading every paper yourself, which is the exact work you bought the AI to avoid.

A confidence score does not help here, because a hallucination arrives with high confidence. The only thing that settles it is reading the source.

HOW IT WORKS

Verify-My-Research-AI does the reading, independently, for every claim your AI makes.

1. It fetches the actual source paper and extracts the full text (deterministic, no model in this step). 2. It reads that paper on its own, with a strong model, producing its own summary, findings, and limitations, uncontaminated by your AI's claim. 3. It then judges your AI's claim against that independent read, and returns either agreement or a specific failure kind (unsupported claim, misquoted source, overstated finding, fabricated result, missed limitation, wrong paper or scope) with the correct answer attached as the receipt.

The verifier sits behind your existing workflow. You change nothing about how your research AI runs. We just check its homework against the source, one claim at a time.

WHAT YOU GET

- A disagreement report, per paper, that says agree or names the exact way the claim missed. - The correct answer as a receipt, grounded in the source paper, so a disagreement is actionable and not just a flag. - An independent read of each paper (summary, key insights, findings, limitations) you can keep. - A running record you can re-run, so you can watch your AI's accuracy over time and per source.

WHO IT IS FOR

- Teams shipping research-assistant AI over the literature (lit-review tools, R&D copilots, science search) who need to prove their summaries hold up. - Analysts, diligence teams, and labs relying on AI-generated paper summaries, where a fabricated finding is expensive to discover late. - Research leads who own the correctness of what the AI ships, not just its fluency.

It is not for casual reading or low-stakes summarization, where a wrong summary costs nothing.

PRICING

We price the ladder, not a number, and every number below is a hypothesis we want to validate with you, not a commitment.

- Pilot (paid proof, fixed price): we run the verifier on a batch of your own paper claims, and you see the disagreement report before you commit to anything further. - Value-metered: once the pilot proves the delta, per verified claim, while your research AI keeps running exactly as it does today. - Platform: license the verifier into your stack as an always-on check behind every AI summary.

Any step that charges a card is gated, and it fails closed, so nothing bills without an explicit human confirmation.

PROOF (the dogfood)

We do not sell a reader we do not use. AYA already runs the deep-read half of this product on itself every single day. Our own research radar (the willow radar) uses the exact same fetch-the-PDF, read-it-with-a-strong-model, parse-it-deterministically atoms this product composes, to read and grade papers for AYA's own R&D. Verify-My-Research-AI is that internal muscle turned outward. The verifier is producer-agnostic, so it grades AYA's own radar summaries against their source papers the same way it will grade yours.

HONEST NOTE

We will not overclaim this, because the whole product is about not overclaiming.

Research has no deterministic ground truth. Unlike a regulation, where a decision tree gives a provably-right answer, a paper's meaning is read, not computed. So the source paper, read independently, is our ground truth, and the judge is a strong model reading that source, not a code oracle that can guarantee correctness. That is honest, and it is still a large improvement over trusting an unread claim, because we ground every verdict in the actual paper rather than in a second guess about the first.

There is a deterministic path, but only for structured decision-tree claims (the regulatory verifier), and it fires only when you supply the tree. For free-text research claims, the check is model-graded by design.

And to be straight about where this build is today: the pattern is composed and structurally verified, but it has not yet been run end to end against a real customer's research-AI export, no live model call or paper fetch was made to produce this page, and the pricing numbers above are hypotheses. The first live batch on your own claims is exactly the pilot.

*This page is a specification. The capability it describes is not built yet, and nothing here is a claim that it runs today.*