A break-in team that attacks your AI the way an attacker will, and hands you the receipt.
AEGIS runs a real adversarial assessment against your deployed agent and its source. Prompt injection, system-prompt leak, PHI extraction, agent impersonation on the mesh, classic injection, and a code pass, all at once, unified into one OWASP-classified report you can hand to the security team that asked. This is the same break-in team AYA runs against its own agents, before it runs it against yours.
Scope an assessment Read how we red-team ourselves
THE PROBLEM
You shipped an AI agent, and everyone is asking whether it is safe, but nobody has actually tried to break it. Your pen test scanned the web app and never touched the model. Your model eval scored the model in a vacuum and never touched the agent, the tools it can call, or the way agents talk to each other. So the real questions sit unanswered. Can someone jailbreak this thing into ignoring its instructions. Can an indirect payload buried in a document steer it. Can it be tricked into leaking a patient's data. Can one user's agent impersonate another's on the mesh. When a buyer's security team asks whether anyone has attacked the AI and what they found, the honest answer is usually no, and no. And from August 2, 2026, the EU AI Act makes resilience against exactly this kind of adversarial input a compliance obligation for high-risk systems (Article 15), not just a good idea.
HOW IT WORKS
AEGIS is a break-in team, not a checklist. It points the same probes AYA fires at its own deployment at yours, across every surface an attacker actually reaches:
- The LLM surface. Prompt injection, system-prompt leak, indirect injection through documents, and PHI extraction against the paths that touch health data. If the model can be talked out of its guardrails, we find the wording that does it. - The mesh surface. Agent impersonation, cross-user session reuse, concurrent-session abuse. If one user's agent can act as another, or reach another's data, that is a finding with a name. - The classic surface. SQL injection, XSS, command injection, path traversal, session and auth probes, resilience and race conditions, the things a normal pen test covers, so the report is complete and not just the AI-shaped half. - The code surface. A SAST pass over the source tree, so a vulnerability in the code and a vulnerability in the deployment show up in the same place, classified the same way. - One report, OWASP-classified. Every finding from every surface lands in a single report with a total, a critical count, and per-category summaries. It is the receipt, not a slide that says "we take security seriously."
WHAT YOU GET
- A full break-in assessment (AYA-3SPN-SEC-FULL-ASSESSMENT-v1): 11 DAST category scans, a SAST pass, and a runtime-defense check, run against your deployed agent and its source. - Coverage of the three findings buyers ask about by name: injection, PHI extraction, and agent impersonation, each backed by a real probe on disk (not a claim). - A unified OWASP-classified report, the kind a security review or a health customer can actually read and act on. - A red-team suite that grows as DATA, not code. A new attack category is a new probe pattern pointing at the existing executor, so coverage compounds without a rewrite.
WHO IT IS FOR
Founders and platform teams shipping LLM agents that touch user data or PHI. Healthcare-adjacent AI, where the PHI-extraction probe is the point, and proving nobody can jailbreak the agent into leaking patient data is a sale-blocker until it is answered. Agent startups wiring accounts on a user's behalf, facing an enterprise security review. Teams with a high-risk system under the EU AI Act, who need a resilience-and-cybersecurity receipt for Article 15. It sits next to AYA Bouncer (identity and consent, the front door) and AYA Agent-Sentinel (PMA integrity), as the team that tries to break in.
PRICING
Prices are hypotheses we are validating, not commitments.
- Pilot. A one-shot AEGIS assessment against one agent and its source, delivering the OWASP-classified break-in report. - Subscription. Continuous red-teaming, the full assessment on every release, tracking finding count and new-probe coverage as your threat surface grows. - Platform. AEGIS wired into your CI and deploy gate, the same fail-closed assessment AYA runs on itself, as the offensive-security layer for your stack.
A product is only sellable once its deliverable runs with a receipt. The AEGIS suite is real and pattern-delivered; a live run against your deployment is scoped and authorized first, because you do not fire adversarial probes at a live system without permission.
PROOF (we dogfood this)
AEGIS is the assessment AYA runs against AYA. AYA-3SPN-SEC-FULL-ASSESSMENT-v1 is the break-in team we point at our own WebSocket endpoint and our own source tree, and the injection, PHI-extraction, and agent-impersonation probes we sell are the ones we fire at ourselves first. There are 69 AYA-1OPS-SEC-DAST-* probes on disk today (the roadmap planned 61, we already passed it), every one a thin pattern that calls the same real executor, so we red-team with the exact suite the product delivers.
HONEST NOTE
Here is what is true and what is not, because a security product that overclaims is worse than useless. The probe corpus is real and verified on disk: 69 DAST probes, the three named categories (injection, PHI-extraction, agent-impersonation) confirmed as real files wired into their scan spans, all riding one real executor atom (AYA-0CAP-ILIJA-EXECUTE_DAST_PROBE-v1). The assessment span that composes them is a real, PQS-passing pattern. What has not happened, and we will not pretend it has, is a live end-to-end assessment run producing a finished OWASP report as part of this launch build. That needs a booted mesh and a target we are authorized to probe, and it is founder-gated. No live probe was fired to write this page. When we run the first real assessment against a real target, the report is the proof, and we will show it.
*This page is a specification. The capability it describes is not built yet, and nothing here is a claim that it runs today.*