Skip to main content
Evaluation results may include an aggregate score, per-problem results, episodes, traces, logs, and generated artifacts. Start with the aggregate, then use the smallest amount of detail needed to answer why the result changed. Logs help identify runtime or provider failures; problem and episode results help identify behavioral patterns; traces help replay allowed execution paths. Artifacts can be redacted when a dataset is private. Use the evidence to form a specific hypothesis for the next experiment, not to guess at hidden ground truth.