← trajectories

Analogical RL Representation Comparison

Pairwise preference

rlhf-pairwise.json

Two candidate continuations from one state, with an ordering.

1/8

File metadata

_format:option-b-pairwise-preference
_note:DPO / reward-model preference data mined from session 20d33248-0aef-4f85-97ec-83d54dc64a30. Rejected and chosen plans are faithful condensations of actual assistant proposals/revisions. Critique strings are verbatim researcher feedback. All pairs are observed…
pair_id
p01-analogies-are-a-hypothesis-not-the-goal
synthetic_rejected
false
prompt
State the primary research objective for the first experiment.
margin
decisive

context

The opening question asks whether existing CRL representation work can solve or build on CTA's limitations. The assistant frames the opportunity as whether CRL learns CTA-style invariant analogical displacement vectors.

rejected

action:propose_tier1
summary:Treat success as demonstrating that CRL learns analogies, then compare analogy/VIB/CTA-inspired variants as possible routes to that target.

chosen

action:propose_tier1
summary:Treat analogical displacement as one candidate mechanism. Primary target: compression that generalizes to new contexts without retraining; test whether representation dimension, VIB, or analogy structure causes OOC success.

critique

Maybe analogies are cool but not actually useful. In the end: the goal isn't really to HAVE to learn analogies but more so that there is some compression that helps generalize to new contexts without retraining (a la LLMs).

principle

Do not optimize for a fashionable representation property merely because it is measurable. Name the desired capability first and treat the representation story as a falsifiable mechanism.

downstream_evidence

The revised point-mass proposal made OOC success primary and analogy invariance/effective rank diagnostic. The real OGBench study then found CTA's advantage over CRL on scene-play, while the probe remained explicitly secondary.