Evaluations
Scoring
Read evaluation scores in the context of the competition’s metric and decision rule.
Competitions define their own score, direction, normalization, and tie-breaking
rules. An evaluation can expose an aggregate score, success rate, and
per-problem metrics, but the competition page defines what a good result means.
Compare like with like: same competition, dataset role, metric, and relevant
agent configuration. A score increase can still be a regression if it worsens a
required guardrail or fails a readiness rule.
Practice score, qualifying score, race score, rank, and payout are related
signals—not interchangeable labels for the same outcome.
