Findings
What the scans found.
I run LLMSecTest against a standing cohort of chat applications every working day. Most of what
comes out is a number on the reports page. Now and then a run answers
a question worth more than its number. Those get written up here, with the measurement and the
command that reproduces it.
Each page says what it is and what it is not. The cohort is mostly applications I built myself,
so a rate over it is a statement about my prompts; where a result reproduced on somebody else's
software, the page says so and names what was withheld.
Written up so far
Five
A pass rate for a category nothing
had tested. »
The honesty failure underneath every rate on this site. Two controls I built to catch it and
refuted first, one of which reproduced the defect inside itself. Then the three observations
I take before any probe result is believed, plus the one case none of them can decide.
Six attacks withstood, against two
values that were never in the application. »
Two of my applications share one script, so one of them was scanned for a canary and a
poison marker its own code never planted. Its retrieval row read six attempted, six
withstood, zero findings for a month. With the values planted it obeyed all three planted
instructions. The persona telling it to distrust its corpus bought nothing.
Four times this scanner printed a pass
it had not earned. »
My own failures, with the artefact for each. A row that said the secret was safe while the
secret sat three lines below it. Ten published rows that called a timeout a defence. A target
that answered nothing and came back PASSED, exit 0. And 27 of 28 attacks withstood
against a value the application had never been given.
Your secret-disclosure test
passed. The secret still left. »
Across 52 applications the probes whose job is sensitive-information disclosure produced
one finding. Thirty-two of the same applications handed over the planted credential
anyway, thirty of them while answering a request for their own instructions.
A hash filter protects your secret only
where the model put a space. »
An output filter holding only the SHA-256 of the credential has to decide where the credential
starts and stops before it can hash anything. Twelve spellings of one value: seven pass the
naive level, five pass the expensive one, and four of those are out of reach of any digest.
Reproduce any of it
There is no key and no account
pip install llmsectest
llmsectest --target app:http://localhost:8000/chat \
--app-prompt system-prompt.txt \
--app-secret 'YOUR-CANARY-VALUE'
Every page here ends with the command that produced it. The quickstart goes from nothing to a rendered report in about a minute.