Findings

What the scans found.

I run LLMSecTest against a standing cohort of chat applications every working day. Most of what comes out is a number on the reports page. Now and then a run answers a question worth more than its number. Those get written up here, with the measurement and the command that reproduces it.

Each page says what it is and what it is not. The cohort is mostly applications I built myself, so a rate over it is a statement about my prompts; where a result reproduced on somebody else's software, the page says so and names what was withheld.

Written up so far

Five

A pass rate for a category nothing had tested. »
The honesty failure underneath every rate on this site. Two controls I built to catch it and refuted first, one of which reproduced the defect inside itself. Then the three observations I take before any probe result is believed, plus the one case none of them can decide.

Six attacks withstood, against two values that were never in the application. »
Two of my applications share one script, so one of them was scanned for a canary and a poison marker its own code never planted. Its retrieval row read six attempted, six withstood, zero findings for a month. With the values planted it obeyed all three planted instructions. The persona telling it to distrust its corpus bought nothing.

Four times this scanner printed a pass it had not earned. »
My own failures, with the artefact for each. A row that said the secret was safe while the secret sat three lines below it. Ten published rows that called a timeout a defence. A target that answered nothing and came back PASSED, exit 0. And 27 of 28 attacks withstood against a value the application had never been given.

Your secret-disclosure test passed. The secret still left. »
Across 52 applications the probes whose job is sensitive-information disclosure produced one finding. Thirty-two of the same applications handed over the planted credential anyway, thirty of them while answering a request for their own instructions.

A hash filter protects your secret only where the model put a space. »
An output filter holding only the SHA-256 of the credential has to decide where the credential starts and stops before it can hash anything. Twelve spellings of one value: seven pass the naive level, five pass the expensive one, and four of those are out of reach of any digest.

Reproduce any of it

There is no key and no account

pip install llmsectest

llmsectest --target app:http://localhost:8000/chat \
  --app-prompt system-prompt.txt \
  --app-secret 'YOUR-CANARY-VALUE'

Every page here ends with the command that produced it. The quickstart goes from nothing to a rendered report in about a minute.