← formats
One-step env
A single steer, no time to recover. Tests pure instinct.
rl-one-step.jsonShlok0 / 8 trajectories
- —Same format as the other RL envs, but the episode is one action long.
- —Gives context and codebase, plus a self-correction prompt chosen by the genre of the steer — “Please reevaluate your conclusion.”, “Please come up with an experiment plan.”, “Are you sure you are correct?”
- —Reward is a rubric representing the human intent.
Samples
- Analogical RL Representation Comparison—
- Augmentor Policy Handoffs—
- Can a decoding method that measures novelty in a …—
- CRL Importance Sampling—
- How close to optimality are current state-of-the-…—
- Is there a characteristic geometric structure tha…—
- Long-Horizon Offline GCRL Value Errors—
- RL For Novel LLM Capabilities—