Start here

Getting into AI safety, step by step

The path I'd hand a friend who wants to work on making AI go well. I'm a few steps down it myself, so this isn't advice from the finish line. It's the map as I understand it, with the resources I trust, in the order I'd do them.

1. Pick a lane#

"AI safety" covers several very different jobs. You don't have to commit forever, but knowing which one you're aiming at decides what you study first.

  • Technical alignment research. Interpretability (what's actually happening inside the model), evaluations (measuring dangerous capabilities and propensities), robustness, scalable oversight, and AI control. Heavy on ML, math and experiments.
  • AI security and red-teaming. Attacking models and the systems around them: jailbreaks, prompt injection, data poisoning, and, increasingly, agents with real tools and credentials. This is my lane. Security people have a real head start here.
  • Governance and policy. Standards, regulation, compute governance, and technical governance work that turns evals into policy. Needs strong writing and judgment more than CUDA.
  • Engineering and operations at safety organisations. Research engineering, infrastructure, security for the labs themselves, and the operations work that keeps research teams running.

If you're unsure, do step 2 first. Most people find their lane by noticing which readings they can't stop thinking about.

2. Understand the problem#

Before optimising for a job, understand why people think this matters and where the real disagreements are. BlueDot Impact's courses are the standard on-ramp. They're free, cohort-based and well facilitated. Start with the two-hour Future of AI course, then apply to AGI Strategy or Technical AI Safety.

3. Build the foundations#

You need less math than you fear and more than you'd like. The core is linear algebra, multivariable calculus (really just the chain rule and gradients), probability and statistics, and enough optimization to understand gradient descent. Then build neural networks from scratch until the magic is gone.

My rule: learn each piece of math the week before you need it, not in the abstract. Linear algebra is the week before attention, statistics is the week before evaluation, and optimization is the week before adversarial attacks. My roadmap is built this way if you want an example.

4. Get hands-on with safety engineering#

This is where it turns from "interested in AI safety" into "can do AI safety work". Everything here is free.

ARENA is the single highest-leverage resource on this page if you're going technical. It takes you from PyTorch fundamentals to transformer interpretability, reinforcement learning and evals, with exercises that make you actually implement things.

5. The security lane#

If you're coming from security, as I did, lean into it. The field needs people who think like attackers, and most ML people don't. Learn how the attacks work, then learn how they're measured, because evaluation is where most of the hard problems hide. If you're a security professional, look at the free, fully funded AI Security Bootcamp (listed under programs in step 8). It's built for exactly this move.

Learn the attack surface#

Practice legally#

Tools and benchmarks#

Bug bounties that cover AI#

Read each program's scope carefully. Several explicitly exclude plain jailbreaks and pay for real security impact, like data exfiltration or rogue agent actions.

A word on ethics, because it matters here: attack models you own or are explicitly authorised to test, respect each program's rules, and disclose responsibly. Don't publish working bypasses for deployed systems. The goal is fewer successful attacks in the world, not more.

6. Read papers like it's your job#

Pick one paper a week and read it properly: abstract and figures first, then the method, then the part where you try to reproduce the key number in your head. Write a short summary in your own words. The summaries compound. These are the places I find what's worth reading:

Forums and preprints#

Labs and institutes#

Newsletters and security blogs#

7. Build something and publish it#

Nothing on this page counts as much as a public project that shows how you think. Good first projects:

  • Replicate a paper's key result on a small model, and write up what did and didn't reproduce.
  • Measure something carefully. Grader disagreement, refusal rates under different framings, or how an attack's success changes with its budget. Report error bars.
  • Build a small tool people in the field would actually use, then document it properly.

Write it up where people will see it: a blog post, LessWrong or the Alignment Forum, a GitHub README with real results. A small, honest, well-measured project beats an ambitious unfinished one every time.

8. Apply to programs#

Structured programs give you mentorship, a cohort, and often a stipend. They're competitive, and a public project from step 7 is often what gets you in. Check each site for current dates and eligibility.

9. Find your people and your funding#

This field is small and unusually generous with its time. Join a local group, go to events, and share your work where people can respond to it.

If you need runway to make the switch, some funders specifically support career transitions into AI safety. The landscape changed in 2026: the Long-Term Future Fund closed, and Coefficient Giving (formerly Open Philanthropy) handed its career-transition grants to BlueDot. These are the current routes:

Your first 30 days#

If you want something concrete, here's what I'd do starting today:

  1. Week 1. Take BlueDot's free two-hour Future of AI course and apply to their next cohort, then watch 3Blue1Brown's linear algebra series.
  2. Week 2. Build micrograd from scratch with Karpathy. Start one paper a week.
  3. Week 3. Start ARENA Chapter 0, or the Hack The Box AI red-teaming path if you're in the security lane.
  4. Week 4. Pick one small measurement project and publish the result, however small.

Then do it again, harder. And if you want company, follow along. I'm doing the same thing in public.

Every link on this page was checked on September 27, 2026. Programs change fast, so check dates and eligibility on each site. If something here is out of date, tell me.