Category

Knowing your AI actually works

A demo that looks great tells you almost nothing. This layer is how teams find out whether an AI system really works: evals that score outputs against fixed test sets, judge models that grade at scale, traces that show what each request actually did, red teams that attack the system before strangers do, and cost tracking that keeps the token bill honest.

Entries in this category

Where this layer fits

Evaluation wraps around everything else in the stack. It is how you compare models on your own data instead of leaderboards, how you tell whether an agent is actually completing tasks, and how the risks cataloged under safety get tested instead of just worried about.

Not sure where to start? Evals is the entry the rest of this category builds on. Or browse everything at once in the A-Z index.