Waiting for results
How well can an agent discover a training strategy?
Compare models within the same class, task, frozen setup and protocol. Five-minute adaptive pilots include inference and gameplay in their wall-clock budget.
Legacy clips report saved net XP from short, single-response programs. The intended 30-minute benchmark measures the best normalized XP/min in complete 15-second authoritative windows. Initial/final XP cannot establish that peak score.
Model × declared class / task
A protocol-declared cohort is required for the research matrix. Existing run evidence remains below.
Each cell retains planned and attempted counts, valid scores, failures, unknown evidence and unfinished runs. One sample is a pilot observation; no uncertainty estimate or overall model ranking is implied. Candidate ranged and magic fixtures have no results until their native skills and evidence are accepted.
The next run starts here.
This page shows actual full-client attempts as their evidence arrives.
- API model returned
- —
- Input actions
- —
- Live XP change · diagnostic
- —
- Last observed state
- —
Results by frozen inputs
Every group is shown. Matching inputs do not prove identical live monster scenes.
Verified persisted results appear here after normal logout.
Persisted XP includes penalties and is measured after normal logout. Publication evidence is checked separately; a blocked check does not erase the verified score. All displayed runs remain unranked.
Every visible attempt
Failures, recoveries, and unscored integration runs stay in the history.
| Attempt / exact model | Status | Persisted XP | Diagnostic XP | Input actions | Publication evidence | Recording |
|---|
No attempts have been exported yet.