Pairwise preference
rlhf-pairwise.jsonTwo candidate continuations from one state, with an ordering.
File metadata
- pair_id
- p01-analogies-are-a-hypothesis-not-the-goal
- synthetic_rejected
- false
- prompt
- State the primary research objective for the first experiment.
- margin
- decisive
context
The opening question asks whether existing CRL representation work can solve or build on CTA's limitations. The assistant frames the opportunity as whether CRL learns CTA-style invariant analogical displacement vectors.
rejected
chosen
critique
Maybe analogies are cool but not actually useful. In the end: the goal isn't really to HAVE to learn analogies but more so that there is some compression that helps generalize to new contexts without retraining (a la LLMs).
principle
Do not optimize for a fashionable representation property merely because it is measurable. Name the desired capability first and treat the representation story as a falsifiable mechanism.
downstream_evidence
The revised point-mass proposal made OOC success primary and analogy invariance/effective rank diagnostic. The real OGBench study then found CTA's advantage over CRL on scene-play, while the probe remained explicitly secondary.