Formats
The supervised, preference, environment, and world-model formats we derive from research trajectories.
Supervised
Preference
RL environments
- Long-horizon research envThe full research problem, over many steps.specRhiaan
- Sparse-reward variantSame problem, binary 0/1 success, no intermediate steps.specRhiaan
- Short-horizon research envA bounded slice of the problem.specRhiaan
- One-step envA single steer, no time to recover. Tests pure instinct.specShlok
- Verifiable envHit the metric. Any solution counts.specShlok