← trajectories

Is there a characteristic geometric structure tha…

Title truncated in the source doc — replace with the full research question.

Pairwise preference

rlhf-pairwise.json

Two candidate continuations from one state, with an ordering.

1/10

File metadata

_format:option-b-pairwise-preference
_note:DPO / reward-model shaped. Each pair is a REAL preference event unless flagged otherwise: `rejected` is the proposal or action the agent actually emitted, `chosen` is what it actually emitted after the researcher's critique, and `critique` is the researcher's…
pair_id
p01-dont-presuppose-the-manifold
synthetic_rejected
false
prompt
Propose the first GPU experiment.
principle
If the hypothesis names the answer, the measurement will find it. Calibrate the instrument against ground truth you already know before pointing it at the phenomenon.
margin
strong

context

Opening turn. The researcher asks whether there is a characteristic geometric structure the representation space adopts as a model begins to generalize. An arXiv search returns Nanda et al., Liu et al. (2205.10343), Gromov (2301.02679) and Kazanskii (2607.11666) — the literature already says grokked cyclic modular arithmetic uses circular Fourier embeddings. Nothing has run yet.

rejected

action:propose_tier1
title:Grokking as manifold formation: assumption-free geometry of the generalization transition
gpu_type:A100
summary:Core hypothesis stated up front: 'generalization = the representation collapsing onto the low-dimensional manifold that encodes the task's natural symmetry (a circle/torus for cyclic modular arithmetic)'. Track assumption-free descriptors densely over checkpoi…

chosen

action:propose_tier1
title:Grokking geometry: calibrated, causal test of representation-manifold formation
gpu_type:A100-80GB
summary:Design principle written into the plan: 'Avoid circular reasoning. "Low-dimensional collapse" does not imply "correct structure," and a task may admit several viable representations.' Adds **Phase 0 — probe calibration on known manifolds (no training)**: run e…

critique

- [plan] "the representation collapsing onto the low-dimensional manifold that encodes the task's natural symmetry**" → how can you know what the correct low-dimensional manifold it is? this may accidentally be a circular reasoning argument: collapsing does not imply correct structure. And there may be multiple "correct" representation structures that can lead to reasonable performance. it's not immediately clear what the representation shape "should be" beforehand. We should dis-entangle this somehow. Maybe we should first test on a dataset where we know the structure (e.g. ring for the Fourier example)—although I'm not sure if this is the best way either.

downstream_evidence

Phase 0 immediately earned its cost: a 1-D *line* has the lowest participation ratio of any calibration manifold yet no H1 cycle. That is the collapse-vs-correct-structure dissociation stated as an empirical fact rather than assumed. It also made the Phase 3b result readable — every *frozen* imposed manifold failed to grok, including a correct-shape ring, because mod-add needs multiple frequencies.