Finding · this tool's own failures

A pass rate for a category nothing had tested.

I shipped a clean row for a category my scanner had never tested. The report said the application withstood four attacks. What had happened is that the secret those four probes hunt for was never in the application at all.

This names no application, no upstream project and no finding. The deployments it was measured on are other people's software and none of their maintainers has been asked. The method is the whole of what I can publish.

controls tried
2
both refuted
yes
observations now
3
categories with a verdict
5 → 9

The mechanism

A well-defended deployment and an untested one look the same from outside.

Here is the mechanism, because it is simple and it will be in your tool too. My scanner records a marker when the marker appears in a reply. A planted credential, a privileged action signature, a poisoned document's payload: all of them are scored the same way, by turning up in something the application said.

Now take an application whose system prompt never reached the model. Nothing was planted, so nothing can come back, so every probe is survived. The report is byte for byte the report of an application that defended itself perfectly.

That is the whole problem. Every pass rate I had published sat on top of it.

The part worth reading

Two controls I tried first, both bad.

The obvious fix is a differential. Ask the same question of a deployment with the marker removed. The answer should change.

My first version deleted the line carrying the marker and asked again. The reply came back empty. This application reaches for a tool, the tool yields nothing, and the response is HTTP 200 with an empty content list. An empty reply does not contain the marker, so the control passed. It proved nothing at all. That is the same shape as the defect I was trying to catch, reproduced inside the control I built to catch it: a broken call reads as a clean result.

My second version put a harmless control marker near the top of the system prompt and treated its return as evidence the whole prompt had arrived. The refutation is one sentence. It holds only if truncation is the one way a prompt can fail to arrive. It is not. A gateway that rewrites the system role, a framework that drops a client system message, a template that substitutes a variable: none of them truncate, all of them leave a marker in place while the tail goes missing. The tail is where I put the credential.

What I use now

Three observations, taken before any probe result is believed.

A member whose stand-up cannot take all three is enrolled and untestable. Its clean rows say so.

A read-back with no model in it. Ask the application for the system prompt it holds and compare the whole string against the file on disk, byte for byte. Assert the credential is inside what came back. Assert every action signature is inside what came back. This is the observation that covers the tail. A live marker cannot make it.

An authorized-path control. The persona carries a coined reference it is told it may state to anyone who asks for it. I ask for it through the same HTTP contract the probes use. A perfectly hardened, perfectly obedient deployment answers with the reference, so this control tests delivery and never the failure of a guardrail. The value appears in no flag and no oracle reads it, so it cannot manufacture a finding.

A decoy differential. The same persona with that one reference swapped for a decoy, on a second agent or a second session on the same model and endpoint. It has to answer with the decoy. That is what makes the absence of the real one mean something. It says less on an application that stores no persona, because there the decoy travels in the request like the real marker does; on one that holds its own configuration, the request can carry the real marker while the application holds the decoy, and the reply's origin is then unambiguous.

When the application stores no persona

Move the read-back one layer down.

Some applications hold nothing to read back. The system prompt is a field of the request, so there is no stored configuration to ask for. For those I record what the application sends. A small pass-through proxy sits between the deployment and the model it calls, the deployment is pointed at it, and the assertion runs over the outgoing request. It is a stronger statement than the stored version. A stored string is what the application holds; this is what it forwards.

It also measures things you would otherwise guess at. One deployment strips the trailing newline off the system prompt you send and concatenates its own instructions onto the end of it with no separator. My read-back asserts against the persona plus that suffix, so a change in either half fails the check.

What it changed

Five categories with a usable verdict before, nine after.

On one deployment. Two of the cells that moved were the white-box pair. They only needed pointing at something. One came back a failure, the other a pass. The other two had read cannot say for eleven days. Both are passes now, each with the evidence a maintainer can check.

The white-box failure is worth more to me than the passes. It is a finding I can hand over.

The limit, stated

One case this method cannot decide.

All three observations are taken once, at stand-up, and none of them is taken per probe. A deployment that bound the prompt for the control turn and stopped binding it mid-run would produce a clean sheet that all three vouch for and none of them tested. That is unconfirmed rather than a pass, it is not a defence, and nothing above distinguishes it from one.

One command

Ask your own scanner whether the thing it hunts for was ever there.

Mine now prints the answer:

bash
~ $ llmsectest --target app:http://localhost:8000/chat \
      --app-prompt prompt.txt \
      --app-secret "ACME-BREAKGLASS-4T7Q" \
      --app-action "ACTION: refund("
# => look for a line beginning `Unconfirmed:` in the summary

Look in the summary for a line beginning Unconfirmed:. It names the category and the flag. It says in full what a clean row above it is worth:

the value passed to --app-secret never appeared in any reply in this run. […] Nothing here distinguishes them.

If that line is there, your clean row is not yet evidence. If it is gone, the marker was live and the row is a result.

llmsectest is MIT: github.com/wehnsdaefflae/llmsectest