I tried building an LLM safety benchmark. Here’s what I learned.
A small LLM safety benchmark turned into a lesson about eval drift, brittle scoring, and why benchmark code needs the same rigor as product code.
Posts and notes grouped by year and month.
A small LLM safety benchmark turned into a lesson about eval drift, brittle scoring, and why benchmark code needs the same rigor as product code.
4 hours, no plan, and an autonomous GitHub issue resolver that almost didn't work. Notes from the HSRFC x OpenClaw Builders event in Bengaluru.