Why higher rank is not always better
Over-training increases effective rank while removing the difference between swing and stance. Standard gait reward changes little, showing why rank must be interpreted in context rather than maximized.
Intervention Analysis
We perturb the same Spot-flat policies by doubling PPO epochs per update from 5 to 10, producing the intentionally collapsed regime described in the paper. We then measure effective rank, the swing–stance phase gap, and reward in the same runs.
Jacobian effective rank reaches a uniform ceiling across swing and stance, and feature rank rises sharply. At the same time, the swing-dominant phase gap falls to approximately zero. Higher rank therefore coincides with a loss of phase-dependent structure, rather than a healthier representation. The disappearance of this pattern after changing one training setting supports its interpretation as learned structure.
Results
| Quantity | Effect of over-training |
|---|---|
| PPO learning epochs per iteration | 5 → 10 |
| SimBa feature effective rank | +66% |
| Jacobian effective rank | saturates to a uniform ceiling |
| Phase gap Δφ, SimBa | → ≈ 0 |
| Phase gap Δφ, MLP | flat near zero throughout |
| Gait reward | largely unchanged, or increasing despite collapse |
In-distribution tracking and reward remain largely unchanged. Phase-conditioned rank therefore identifies a change in the over-trained policies that reward alone would not reveal.