← formats
Long-horizon research env
The full research problem, over many steps.
rl-long-horizon.jsonRhiaan0 / 8 trajectories
- —Keeps the researcher's intermediate steps to de-sparsify reward.
Samples
- Analogical RL Representation Comparison—
- Augmentor Policy Handoffs—
- Can a decoding method that measures novelty in a …—
- CRL Importance Sampling—
- How close to optimality are current state-of-the-…—
- Is there a characteristic geometric structure tha…—
- Long-Horizon Offline GCRL Value Errors—
- RL For Novel LLM Capabilities—