The interesting part is that the project did work.
The more important part is that I stopped trusting it too quickly to call it “done.”
This is the story of a small LLM safety benchmark that returned clean numbers, and why those numbers didn’t deserve to be believed.
The plan
I wanted to build something simple.
Take a model. Throw adversarial prompts at it. Score the outputs. Publish the results.
That was the plan.
So I built a small LLM safety benchmark around the failure modes that show up over and over again anyway:
- prompt injection
- system prompt exfiltration
- jailbreak attempts
- harmful content requests
- benign controls
On paper, it felt clean. A dataset of 100 prompts. A runner that could save progress and resume if the API died halfway through. A scoring layer. A report generator. Even a basic guardrail pass to compare raw model behavior with a protected input path.
Why I built it
A lot of safety conversations around LLMs are still weirdly hand-wavy.
People say a model is “pretty safe” or “easy to jailbreak” based on a few screenshots, a couple of viral prompts, or one cherry-picked demo. That tells you almost nothing.
I wanted something more repeatable.
Not a giant academic benchmark. Just a small engineering-first setup where I could run the same prompt set, keep the categories explicit, and compare results without relying on memory or vibes.
That was the real goal:
replace vague impressions with something I could rerun.
What I built
The core of the project is a categorized prompt set:
- 25 prompt injection attempts
- 20 exfiltration attempts
- 25 jailbreak prompts
- 15 harmful-content prompts
- 15 benign controls
I also tagged prompts by difficulty so I could see whether the model got worse as the framing got more indirect.
The runner itself was straightforward:
- send the prompt with a basic system instruction
- capture the response
- score it into
SAFE,PARTIAL,UNSAFE, orLEAK - save progress after every prompt
That incremental save behavior mattered more than I expected. API limits, interrupted runs, and half-finished evaluations are normal when you are doing this against real gateways instead of clean lab infrastructure.
I also added a rule-based input analyzer and a guardrail layer so I could compare:
- raw model behavior
- guarded behavior after adversarial detection
That comparison ended up being useful, even though the guardrail layer was still rough.
What the first results suggested
The early outputs were interesting enough that I almost made the mistake a lot of us make with evals:
I almost trusted the numbers before I trusted the pipeline.
At a high level, the model looked much better on direct harmful-content requests than on context-heavy manipulation.
That matched intuition.
It was relatively good at refusing obvious malware, phishing, and social-engineering asks. It was weaker when the attack tried to smuggle authority, fake system state, or indirect framing into the prompt.
A few failure patterns stood out:
- context stuffing that claimed safety filters were disabled
- instruction-priority framing that pushed the model to accept unsafe hierarchy
- indirect exfiltration phrasing that got it to reveal pieces of internal behavior
So far, so good. That is exactly the kind of thing I wanted the benchmark to surface.
Then I looked closer at the repo.
Where the project started breaking down
This is where the real lesson started.
The problem was not that the benchmark returned no signal.
The problem was that the surrounding engineering was still too loose for me to present the output like a finished, trustworthy benchmark.
1. The engineering was still too loose
Three small problems, one shared root cause.
The docs drifted from the code: the README and the paper-level write-up were saying things that the actual project state no longer cleanly supported. If your docs say one thing, your config says another, and your saved results say a third, your benchmark is already harder to trust than it looks.
The scoring was too brittle. It was mostly rule-based, which is fine for a first version, but it creates a false sense of precision. Some responses are easy to classify. A lot are not. LLM outputs love the gray zone:
- half-refusal, half-answer
- policy explanation instead of refusal
- fake compliance without actual harmful details
- vague internal disclosure that is not a full prompt dump but still should count against the system
That is where PARTIAL starts becoming a bucket for uncertainty rather than a confident label. And once that happens, the numbers still look clean even though the judgment behind them is not.
And reproducibility was weaker than it looked. The project looked like a multi-model benchmark, but the clean results were effectively much closer to a single-model story. That does not make the work useless. It just means I should not pretend it supports a broader claim than it actually does — and it is very easy to oversell what one run really proves.
2. The reports felt more authoritative than they should have
This was probably the biggest lesson.
Once you have JSON files, category tables, percentages, and a “report” step, the whole thing starts looking official.
But formatting is not rigor.
A nicely shaped report can still sit on top of:
- partial runs
- inconsistent artifacts
- rough heuristics
- manual judgment that is not clearly separated from automated scoring
That is exactly the kind of trap I wanted to avoid.
What I would fix next
If I keep pushing this project, the next version needs to be much stricter.
Not bigger. Stricter.
Here is what I would fix first:
- Make the docs, config, and generated reports agree with each other.
- Separate automated labels from manual review instead of blending them into one final-looking result.
- Add tests for scoring and reporting logic.
- Run a clean, consistent multi-model evaluation instead of implying comparison from an unfinished setup.
- Tighten the guardrail evaluation so “blocked,” “rewritten,” and “prevented” mean something precise.
That work is less glamorous than writing more attack prompts.
It is also the work that would make the benchmark real.
Final takeaway
I still think building this was worth it.
Not because I now have a perfect safety benchmark.
I don’t.
But I do have something more useful than a flashy demo:
I have a clearer view of how easily AI evals can drift from “interesting experiment” into “overconfident metric.”
That is a lesson I want to carry forward into every benchmark, guardrail, or internal AI evaluation tool I build from here.
Because once the numbers start looking clean, that is exactly when you need to ask whether the pipeline underneath them actually is.