Welcome to Safety in the Layers
What this site is, the 180-day plan behind it, and the four kinds of posts you'll find here.
Illustration: a language model completing “safety in the …”. The next-token dropdown chooses “layers” (0.91) and strikes out “jailbreak” (0.06) and “breach” (0.03).
I'm Prathik, a security engineer spending the next six months locked in on alignment, interpretability and adversarial robustness. Everything goes here: the daily log, the papers, the builds, and the exact roadmap. Follow along, or take the same path.
Day 7 of 180 · Week 2: Projections, eigen-structure, SVD, calculus, Karpathy
The lock-in
Each week of the plan is one token in the context window. The glowing one is where I am right now.
What I post
The same work, cut four ways: the day-by-day log, the papers behind it, what I build, and what I can teach back.
Daily and weekly notes from the lock-in: what I studied, built, broke and learned.
1 postResearch papers broken down: the claim, the method, the numbers, and what I think.
First post soonThings I'm building and measuring: harnesses, attacks, evals, and the results.
First post soonConcepts explained the way I wish they had been explained to me.
First post soonLatest
Logs, paper notes, builds and explainers, newest first.
All postsWhat this site is, the 180-day plan behind it, and the four kinds of posts you'll find here.
Projects
When gradient-based adversarial attacks fail against safety-tuned language models, how much of that failure is real robustness, and how much is the search, the harness, or the grader?
How does the framing of a legitimate security request move a model's refusal boundary?
Which open-weight models are actually useful for security work on a single 16 GB GPU, and how do you measure that honestly?
New here?
Pick a lane, then build foundations, hands-on skills, papers, projects, programs and community, with the free resources that actually matter at each step.