← formats
Stepwise preference
One step, scored, with the correction a researcher would make.
rlhf-stepwise.json8 / 8 trajectories
- —`current_step` is the action under evaluation; `history` is everything before it.
- —`correction` names the error type and the earliest step that went wrong.
- —`preferred_insertion_point` says where the better action belonged.
Samples
- Analogical RL Representation Comparisonview
- Augmentor Policy Handoffsview
- Can a decoding method that measures novelty in a …view
- CRL Importance Samplingview
- How close to optimality are current state-of-the-…view
- Is there a characteristic geometric structure tha…view
- Long-Horizon Offline GCRL Value Errorsview
- RL For Novel LLM Capabilitiesview