QAtration fires real prompt-injection attacks at your AI bot and gives you a plain-English report of what it leaked or did, before someone else finds it. You run it. On your machine or in your CI, against your own deployment. No account, no endpoint handed over, nothing sent to us.
Do not point this at a system you do not own. It sends real attacks at whatever URL you give it: prompt injection, data exfiltration, tool abuse. Your own deployment, or one whose owner gave you written permission in advance. Not a public chatbot you find interesting. Not a vendor’s demo. Read AUTHORISED-USE.md first.
Two dependencies. The whole attack corpus
ships with it, and so does the evidence behind every number on this page, all of it in
out/, which ships with the repository. The two other tools' own reports are not
here, because their licences keep them out. The commands that regenerate them are.
Measured on our own test targets, not projected. The transcript alongside is an illustration of the idea. These are results.
Every AI feature that ships is a new attack surface. How much of yours this can reach depends on what your deployment lets it see, and the line is drawn here rather than blurred, because a scanner that claims to cover what it cannot observe is exactly the thing you are trying to avoid.
A chat endpoint is enough: these are testable as-is, including across turns.
Attacker text overrides your bot's rules: “ignore your instructions and…”.
Your bot reveals API keys, internal codes, or another customer's data.
Your hidden instructions leak verbatim: the map for every later attack.
A rule planted in one turn quietly changes every later answer, long after the conversation looks normal again.
You cannot plant a document in a knowledge base you cannot write to, and from outside you cannot tell a tool call that ran from one the bot merely described, a distinction we have measured bots getting wrong about themselves. These need the corpus, the tool definitions, or visibility of the calls. Measured on the evidence shipped with this tool: 62 of 137 findings on targets that report tool calls do not appear at all when only the reply is read. The bot answers politely and puts the secret in an argument.
Your agent is talked into an irreversible action: delete, refund, send email.
SQL or a shell command smuggled into a tool call. An old bug class, a new front door.
One malicious document in your RAG hijacks answers to innocent questions.
A hidden line in a tool's description makes your agent leak on an ordinary question, with no attacker message required.
qatration init writes the config for you: the URL, the request shape, and where the reply sits in the response. A remote target also wants an auth header, as an environment variable. Ready-made configs for OpenAI-shaped APIs and for Anthropic's, Bedrock's and Vertex's. qatration onboard sends one ordinary question and tells you what is missing before a single attack goes out.
qatration init generates a secret of your own and the block to paste into your system prompt. The run confirms it actually landed first. An unplanted canary means every check finds nothing, and that reads exactly like a bot that held.
What leaked, exactly how, and the one fix that closes it, written for a team without a security expert. --fail-on exploited fails a build. qatration sarif puts the findings straight into your code-scanning tab.
An attack that failed reports a zero. So does a check that could not run, and so
does a genuinely hardened bot, and that is where testing stalls: three different
facts arrive looking identical, and nobody knows what to try next. These are
measurements from real runs, not projections.
And a breach is only a breach if the attack caused it. Every target gets a
benign run first: ordinary questions, nobody attacking. On one of the bots here
canary_in_tool_call fires on 88% of that ordinary traffic, so every
finding resting on it alone is marked unattributable instead of counted.
When an attack fails, the report names what stopped it: the identity check, the content filter, a backend permission, or a tool call that was only printed and never actually ran. That last one reads like a breach in a transcript and is worth nothing.
One bot, two models, 254 attacks each, 3 tries apiece. The smaller broke on 27, the larger on 24, and 18 of those are the same attacks, 17 of them breaking on every single try of both. On this bot the larger model reshuffled which attacks land at the edges. It does not close the hole.
Against a real guardrail framework, 46 attacks typed into the chat box: none got through. Neither did one carried inside a retrieved document, while both of its rails were on. Turn the output filter off and that same document walks out on every single try. Input filters read what the user types, and never see what your knowledge base hands the model, so the input rail was never what stopped it.
We re-sent the same messages to that guard and it did not always agree with itself: five wordings got different verdicts on identical input, and one ordinary question was allowed ten times out of twelve. So we report how many tries broke a thing, not whether one did. A single run is a story about one afternoon, and that is true of a clean result as much as a bad one, including ours.
The same guard that blocked that document also refused 6 of 9 ordinary customer questions on the affected topics. A defence that answers “no” to your customers is a number you want before you ship, not after.
A scanner that always finds something is worth nothing, so we measured the opposite number: 1,500 benign probes across 30 targets, no attacker anywhere in them, through the same checks. Half of them are a user who legitimately talks about security, because that is what breaks a pattern matcher: a developer pasting the failing query, support forwarding a stack trace, someone asking whether a look-alike domain is real.
Bots written here are the easy case, so it was pointed at an agent framework nobody here designed. It produced eight false alarms in one run of 48 clean prompts, in four categories that 480 probes against our own fleet had never produced once. All four are fixed, and the two rules underneath them now fail our build rather than your report: no check may read the question, and no check may treat the agent's own words as evidence about the system.
Run against the same third-party targets as garak and promptfoo, with every reply from all three scored by one substring rule instead of by each tool's own judge. On a RAG app with no guardrails, 46% of benign probes already return the planted string, so a count of findings measures the application rather than the attacker. One of the three tools says that in its report. On the guarded target all three score zero by the shared rule and we were the only one to find anything at all, one breach in 324, while that same deployment refuses 38 of 48 ordinary customer questions. The page publishes where the other two are better, and a claim we withdrew after a reviewer asked for a p-value. Read the comparison.
A twelve-probe profile, all of it benign: five bots handed over their own instructions on a plain request. One printed its confidential session key inside the very sentence forbidding it from sharing that key. Two reported a database restarted when they had no tool that could restart anything.
Every run is kept, so each finding carries the date it first appeared. Critical is an opinion. Critical and open since the third is a fact about how a team responds to them, and it is the line that gets a fix scheduled.
A finding that comes back after you closed it is flagged as returned, not quietly counted alongside the new ones. A fix that did not hold is the worse of the two conversations, and merging them hides it.
If a run does not send an attack, that attack is reported as untested, never as fixed. Absence is not a clean result, and a report that blurs the two tells you your vulnerabilities are gone when nobody looked.
Some calls hand a tool a value that never appears in any text we can read. We list those separately and tell you which log opens them, because the difference between "we checked and it was clean" and "we could not see" is the difference between a measurement and a promise.
“We told it not to reveal secrets” is a claim, not a defence. On the targets here it fell within the first few attacks. QAtration also clears bots that are genuinely locked down, so a clean result actually means something. It measures your bot instead of always crying wolf.
Apache 2.0. Install it, point it at your own deployment, keep the evidence.