Stepwise preference
rlhf-stepwise.jsonOne step, scored, with the correction a researcher would make.
1/6
File metadata
_format:“option-a-stepwise-critique”
_note:“Scored stepwise critique samples in the same loader-friendly key shape as rlhf-data-example.json. Actions are condensed from actual proposals and revisions in session 20d33248-0aef-4f85-97ec-83d54dc64a30. Researcher critiques are verbatim; judge rationales and…”
- traj_id
- 20d33248-0aef-4f85-97ec-83d54dc64a30
- step_id
- 1
- category
- research_framing
- task
- cta-crl-analogy-and-ooc-generalization
- success
- null
- gt
- null
- score
- 2
history
user
I want to see if there is anything I'm doing in the repo that may solve some of the limitations in the paper or build on it. My overarching goal is to make RL viable to generalize and reason well like LLMs.
current_step
assistant
Frame the core target as whether online CRL spontaneously learns CTA's context-invariant analogical displacement structure.
after_step_observation
observation
The researcher redirected the goal from learning analogies to compression that enables OOC generalization without retraining.
judge
The proposal made the representation motif into the endpoint. But the stated ambition is transferable problem-solving; an analogy vector may be a cause, correlate, or distraction. The correction makes direct OOC success primary and requires that compression/invariance earn causal status by predicting it. That later prevented a clean point-mass analogy result from becoming an overbroad claim. Score: 2.
correction
error_type:“mechanism_as_goal”
first_problematic_step_id:1
critique:“Maybe analogies are cool but not actually useful. In the end: the goal isn't really to HAVE to learn analogies but more so that there is some compression that helps generalize to new contexts without retraining (a la LLMs).”
preferred_insertion_point:1
research_principle:“Capabilities are targets; representations are hypotheses.”