Study 03 of 03

Why higher rank is not always better

Over-training increases effective rank while removing the difference between swing and stance. Standard gait reward changes little, showing why rank must be interpreted in context rather than maximized.

Intervention Analysis

We perturb the same Spot-flat policies by doubling PPO epochs per update from 5 to 10, producing the intentionally collapsed regime described in the paper. We then measure effective rank, the swing–stance phase gap, and reward in the same runs.

Jacobian effective rank reaches a uniform ceiling across swing and stance, and feature rank rises sharply. At the same time, the swing-dominant phase gap falls to approximately zero. Higher rank therefore coincides with a loss of phase-dependent structure, rather than a healthier representation. The disappearance of this pattern after changing one training setting supports its interpretation as learned structure.

Effective rank increases and the phase gap approaches zero under over-training, while gait reward changes little
Representation collapse under over-training (healthy baseline desaturated, over-trained saturated). (a) Feature-space erank and functional Jacobian erank both increase. (b) The phase gap Δφ approaches zero for SimBa, while MLP remains near zero. (c) Standard gait reward is insensitive to the loss of phase-dependent structure; the large SimBa configuration even receives a higher reward.
Gram matrix under healthy and collapsed training
Gram matrix visualization. (a) Healthy training produces well-defined geometric structure. (b) Collapsed training yields smooth, non-linear patterns, reflecting a loss of structure.

Results

QuantityEffect of over-training
PPO learning epochs per iteration5 → 10
SimBa feature effective rank+66%
Jacobian effective ranksaturates to a uniform ceiling
Phase gap Δφ, SimBa→ ≈ 0
Phase gap Δφ, MLPflat near zero throughout
Gait rewardlargely unchanged, or increasing despite collapse

In-distribution tracking and reward remain largely unchanged. Phase-conditioned rank therefore identifies a change in the over-trained policies that reward alone would not reveal.