Category

Keeping AI systems safe and accountable

Every layer of the AI stack can fail, be attacked, or be misused - and every serious team plans for all three. These entries cover the safety and governance layer: what AI safety and alignment actually mean, how attacks like prompt injection and jailbreaking work and how teams defend against them, and the practices and rules - responsible AI programs, model cards, and regulation - that keep AI systems accountable to the people who use them.

Entries in this category

Where this layer fits

Safety work cuts across the whole stack rather than sitting at one level. It starts inside the model with the training techniques covered in core concepts, gets enforced at the application layer where agents add guardrails and human oversight, and gets measured by evals like red-teaming that probe whether the defenses actually hold.

Not sure where to start? AI safety is the entry that frames everything else in this category. Or browse everything at once in the A-Z index.