I tried building an LLM safety benchmark. Here’s what I learned.
A small LLM safety benchmark turned into a lesson about eval drift, brittle scoring, and why benchmark code needs the same rigor as product code.
Software developer in Bengaluru. I build web applications, work with AI tools, and write down what I learn from real projects.
Right now I'm at Doozie Software Solutions, working on LLM and full-stack problems. Outside of work I run a small personal lab where I build real-time speech systems and evaluation tooling — currently MendSpeech and AI Risk Evaluation Workbench.
A small LLM safety benchmark turned into a lesson about eval drift, brittle scoring, and why benchmark code needs the same rigor as product code.
4 hours, no plan, and an autonomous GitHub issue resolver that almost didn't work. Notes from the HSRFC x OpenClaw Builders event in Bengaluru.