Pairwise preference
rlhf-pairwise.jsonTwo candidate continuations from one state, with an ordering.
File metadata
- pair_id
- p01-improve-the-data-not-the-controller
- synthetic_rejected
- false
- prompt
- Give the first concrete formulation of the augmentor.
- critique
- Yeah I definitely think that better data can beat better algorithms: we should improve the logging policy. I think we should explore trajectory-level data augmentation
- margin
- strong
context
Opening turn. The researcher's question is whether an augmentor trained on limited logged data can enrich a hand-scripted, deterministic logging policy during collection while handing control back safely. Literature check surfaced Trajectory Stitching (Hepburn & Montana 2022) and ORIL. Nothing has been designed yet.
rejected
chosen
principle
When the stated bottleneck is data quality, intervene on the data-generating process at the granularity of a trajectory — do not answer it by bolting a learned controller with its own uncertainty estimator, gate, and auxiliary losses onto the logger.
downstream_evidence
The entire remainder of the session is trajectory-level: detour replacement, the summed-action bridge, the shortcut inequality. No learned augmentor network was ever trained, and the final method has no learned component at all — the researcher's later steer 'no network involved' is the same preference again.