AI safety courses and concepts

AI safety is the work of making sure AI systems do what they are meant to do and not something harmful — covering how models fail, how people misuse them, and the rules and controls placed around both. It is unusual among technical subjects in that its hardest problems are not technical: a system can work exactly as designed and still produce an outcome that is unfair, unexplainable, or illegal in the country it is running in. Learning it means holding two things at once — the specific ways models go wrong, and the governance built to catch them before anyone else does.

The vocabulary of AI safety

Safety vocabulary is deceptive because it is made of ordinary English. Alignment, transparency, fairness and interpretability each name something narrow and specific, and each is used loosely everywhere else — so a policy document or a course outline can deploy one as though it were self-explanatory when it is not. Alongside them sit the names for particular failures: what it is called when a model invents a source, when an instruction is smuggled in through the material it reads, when a safeguard is talked around.

Courses on AI safety

The dividing question is who you are answerable for: your own use of these tools, or a system that other people rely on. One kind of course is about personal judgment — deciding what to hand over, checking what comes back, being straight about where it came from; the other is about the apparatus you put around a system you own, from evaluation and access control through to the paperwork a regulator will ask for.

Safety is the one part of AI that stays relevant whichever side of it you end up on. Learn AI sets out how the systems work, which is what most safety questions eventually rest on; Use AI covers the everyday practice of working with these tools well; Develop AI is where the controls get built rather than described.

← Browse all courses