Leaderboard

Loading…

Register a team and submit →

How the score works

Each task is scored on a handful of weighted biological questions. Every metric is rescaled against two published anchors — the floor (copy_last, or wt_identity for Task 3) and the attainable ceiling, which is half the held-out target scored against its other half. The mapping is hyperbolic: 50 is the floor and 100 the ceiling. Downward it is bounded by construction rather than by clipping, so a model worse than doing nothing lands below 50 and keeps going without the scale running out of room. Upward it is clipped — the ceiling is a half-against-half estimate, not a true optimum, so a submission can beat it, and 100 is where it stops.

A metric a submission cannot produce, or that returns NaN, counts as zero skill inside its question — never dropped from the average. Full derivation →

Submission rules

  • One team per person, one track per team. Ranking uses each team's best score on a task.
  • Submissions are rate-limited per team per task per day, and the limit tightens in the final phase to reduce leaderboard probing.
  • Ground truth is withheld while it can still affect the board. At the start of the final phase the validation answers are released for every task and the test leaderboard opens, since ranking has moved to the hidden test split by then. Test ground truth is never distributed.
  • Duplicate registrations, unauthorised data use, code sharing across teams, or falsified results are grounds for disqualification.

Reference baseline scores are regenerated on every scoring run, so they move whenever the panel does.

staging — copy of production data, not the live sitego to the real site ↗