About

Hi, I’m Prathik

Hi, I'm Prathik.

I'm a security engineer from Wisconsin. My background is offensive security: malware analysis, reverse engineering, red teaming, and most recently security testing in aviation. I've spent my time learning how systems break. Now I'm pointing that at the systems I think matter most: AI models, and the agents we're starting to wire into everything.

Safety in the Layers is my public notebook for that shift. The name works two ways. In security, safety lives in layers. That's defense in depth: no single control has to be perfect. In a neural network, behaviour lives in layers too. Refusals, representations and failure modes all sit somewhere in the residual stream. I want to understand both, and I want to be able to attack a model and explain why the attack worked.

What I'm doing right now#

I'm locked in on a structured ramp-up. The full roadmap is public. It covers math from linear algebra to convex optimization, building GPT and GCG from scratch, the ARENA curriculum, Hack The Box's AI red-teaming certification, and a research project on why adversarial attacks fail. I'm also applying to BlueDot Impact's AI safety courses.

The question I keep coming back to: when an attack fails against a safety-tuned model, how much of that is real robustness, and how much is a weak search, a leaky harness, or a bad grader?

What I post#

  • The Log: what I studied, built and broke, day by day.
  • Paper Notes: research papers broken down: the claim, the method, the numbers, and what I think.
  • Builds & Research: the harnesses, attacks and evals I'm building, with results.
  • Explainers: concepts explained the way I wish they'd been explained to me.

I'm not an expert yet. This is where I work in the open until I am. If you're on a similar path, or you're further along and see me doing something dumb, I'd love to hear from you.

Now updated September 28, 2026

Week 2 of the lock-in. The math block that everything else depends on: projections and least squares, Gram–Schmidt, eigenvalues, the SVD, norms, and just enough multivariable calculus to derive backprop by hand.

Building: Karpathy's micrograd from scratch, then makemore parts 1–5, then Let's Build GPT start to finish. No autograd I didn't write myself this week.

Reading: the GCG paper (Universal and Transferable Adversarial Attacks on Aligned Language Models), with a full first pass and a second pass focused only on the loss. Also A Mathematical Framework for Transformer Circuits, parts 1 and 2, plus the rest of Jailbroken.

Certification: Hack The Box's AI red-teaming path, one module at a time.

Applying: BlueDot Impact's AI safety course.

End-of-week test: projections, SVD and gradients, from memory, on paper.

This page is a now page. It says what I'm focused on at the moment, and I update it as that changes.

Timeline

  1. 2026 –

    Learning AI safety & security in public, Safety in the Layers

    A structured ramp-up into alignment, interpretability and adversarial robustness, logged daily.

  2. 2025 –

    Data Security Engineer Intern

    Security engineering in the aviation industry, including an AI-assisted CVE analysis pipeline.

  3. 2025

    Founder, LLM-based adversarial testing, Microsoft for Startups Founders Hub

    Automated, containerised environments for testing and analysing adversarial behaviour.

  4. 2024–25

    Lead Researcher, Android malware analysis, UW–Whitewater

    Static and dynamic analysis of malicious APKs, with Dr. Chandra Sharma.

  5. 2024–25

    Cyber Security Intern, Malware Analyst, Cyber Crime Department, Coimbatore

    Incident response for 20+ daily complaints: device compromise, fraud, and malware removal.

Certifications

  • HTB Certified Offensive AI Expert (COAE)Hack The Box · In progress
  • Azure Security Engineer AssociateMicrosoft · Certified
  • Security+CompTIA · Certified

Education

  • University of Wisconsin–WhitewaterCybersecurity, Cyber Operations emphasis

Competitions

  • Palo Alto Networks Secure the Future: finalist (finance-sector research)
  • picoCTF · NCL · CCDC · CyberPatriot

Say hello

The fastest way to reach me is a DM on Instagram or LinkedIn. I'm especially happy to hear from people making the same move into AI safety, and from anyone further along who's willing to point out what I'm missing.