Pairwise preference
rlhf-pairwise.jsonTwo candidate continuations from one state, with an ordering.
File metadata
- pair_id
- p01-derisk-before-build
- synthetic_rejected
- false
- prompt
- Propose the first GPU experiment.
- principle
- De-risk the single load-bearing assumption before building the method that depends on it — and make the de-risking artifact human-inspectable, not a summary statistic.
- margin
- strong
context
Opening turn of the project. The researcher wants to know whether hidden-representation novelty can improve Pass@k. Literature check found no prior work using residual-stream novelty as the diversity signal. The researcher has specified: fork when the model reaches a new thinking point; R1-Distill-Qwen-1.5B on math; scope is just 'compare normal pass@k with this new method'. No experiments have run yet.
rejected
chosen
critique
can we actually first test if this is a reasonable way to detect idea boundaries? maybe if we can see a side by side of the tokens outputted and whether there is an idea happening right now. lets do the forking after we derisk the first idea boundaries
downstream_evidence
The de-risking run returned F1=0.474 (P=0.37) — a real but noisy signal. Had the full method run first, a null Pass@k result would have been uninterpretable: broken detector, or good detector and bad forking?