← trajectories

RL For Novel LLM Capabilities

Pairwise preference

rlhf-pairwise.json

Two candidate continuations from one state, with an ordering.

1/10

File metadata

_format:option-b-pairwise-preference
_note:DPO / reward-model shaped. Each pair is a REAL preference event where possible: `rejected` is the proposal the agent actually emitted, `chosen` is the proposal it actually emitted after the researcher's critique, and `critique` is the researcher's verbatim wor…
pair_id
p01-rl-native-substrate
synthetic_rejected
false
prompt
Formalize the exploration-lever direction into the first GPU experiment.
critique
hmm is there something more RL-centric that can solve something more concrete there instead of training an LLM?
margin
decisive

context

Opening of the project. The researcher: "I'm curious about how RL can help LLMs get better at doing fully novel stuff. Let's scope out a small experiments/couple of things so we can figure out some promising directions". The knowledge base is empty. The agent framed the RLVR sharpen-vs-expand tension and offered three directions; the researcher picked "B: exploration lever" and "Puzzle/game (novel)" from the offered options. No experiment has run.

rejected

action:propose_tier1
title:Exploration pressure: does RL expand or just sharpen? (Countdown)
gpu_type:A100
summary:Qwen2.5-1.5B base + GRPO on Countdown with a verifiable exact-match reward. Four arms with the exploration mechanism as the only independent variable: vanilla GRPO (the sharpen control), +entropy bonus, +rollout temperature 1.2, +diversity reward for correct s…

chosen

action:reframe + ask_researcher
summary:Drop the LLM substrate. 'The LLM version tangles the novelty question up with pretraining priors, tokenizer quirks, and reward-shaping hacks, so it's hard to cleanly attribute any "expansion" to RL itself.' Keep the design principle — hold the RL algorithm fix…

principle

Run a mechanism question in the smallest substrate that still exhibits the phenomenon; if the substrate supplies its own explanations for the effect you are measuring, no result from it is attributable to the mechanism.

downstream_evidence

Every remaining experiment in the session ran in Crafter/Craftax. The LLM proposal was never revisited and appears in the agent's own end-of-session accounting as dead: 'the LLM Countdown "sharpen vs expand" study (you pivoted away from LLMs)'.