Overview
An evaluation is a set of configured tasks that you run on a model or a dataset to produce technical evidence with full provenance.
What you can do
Evaluations
Run evaluations, control caching and repeatability, and download evidence.
System Tasks
Use ready-made tasks to evaluate common model and dataset properties.
Benchmark Tasks
Define custom evaluation tasks declaratively in YAML.
Solver
Control how a task interacts with your model, from single- to multi-turn.
Scoring
Score model outputs with built-in scorers or model-as-a-judge.