← trajectories

How close to optimality are current state-of-the-…

Title truncated in the source doc — replace with the full research question.

Sparse-reward variant

rl-long-horizon-sparse.json

Same problem, binary 0/1 success, no intermediate steps.

No sample yet. Add data/trajectories/optimality-gap/rl-long-horizon-sparse.json.