Pairwise preference
rlhf-pairwise.jsonTwo candidate continuations from one state, with an ordering.
File metadata
- pair_id
- p01-rl-native-substrate
- synthetic_rejected
- false
- prompt
- Formalize the exploration-lever direction into the first GPU experiment.
- critique
- hmm is there something more RL-centric that can solve something more concrete there instead of training an LLM?
- margin
- decisive
context
Opening of the project. The researcher: "I'm curious about how RL can help LLMs get better at doing fully novel stuff. Let's scope out a small experiments/couple of things so we can figure out some promising directions". The knowledge base is empty. The agent framed the RLVR sharpen-vs-expand tension and offered three directions; the researcher picked "B: exploration lever" and "Puzzle/game (novel)" from the offered options. No experiment has run.
rejected
chosen
principle
Run a mechanism question in the smallest substrate that still exhibits the phenomenon; if the substrate supplies its own explanations for the effect you are measuring, no result from it is attributable to the mechanism.
downstream_evidence
Every remaining experiment in the session ran in Crafter/Craftax. The LLM proposal was never revisited and appears in the agent's own end-of-session accounting as dead: 'the LLM Countdown "sharpen vs expand" study (you pivoted away from LLMs)'.