Mind the Phase
Effective Rank and Representation Health in Legged Locomotion
Mobile Robotics Group, University of São Paulo, São Carlos, SP, Brazil
We study what locomotion policies learn and how their representations relate to movement on a real robot. Measuring effective rank separately during swing and stance reveals differences between architectures that a whole-gait average misses. These phase-dependent patterns accompany smoother control, but higher rank alone does not indicate a healthier policy.
Abstract
Reinforcement learning is widely used in legged locomotion, enabling behaviors such as backflips and parkour through massively parallel simulation. Shallow networks remain common under the non-stationarity of PPO training, supported by carefully staged curricula and environments. However, their learned representations remain poorly understood, limiting the signals available during training to assess how a policy may behave on hardware.
We empirically study locomotion policies using the effective rank of the policy Jacobian. Conditioning this measure on gait phase reveals architectural differences that are obscured by global rank. Networks with layer normalization and residual connections have higher effective rank during swing than during stance; this swing-dominant pattern is absent in vanilla MLPs.
We use these representational signatures to guide an architecture and training recipe for smoother sim-to-real transfer. The reduction in joint jitter observed in simulation also holds on a physical Spot. These results suggest that phase-conditioned rank can help assess representation health during training, without treating rank itself as an optimization objective.
Mechanistic studies of policy representations for healthy legged locomotion.
Explore the studies
How rank differs across gait phases
Separating swing from stance reveals differences between policy architectures that are hidden by a global average.
View study
From simulation to smoother robot motion
We compare joint jitter and velocity tracking in simulation and on a physical Spot, alongside representation measurements across quadrupeds and a humanoid.
View study
Why higher rank is not always better
Over-training increases rank while removing the swing–stance distinction, even when gait reward gives little indication of the change.
View studyVideo
Resources
The arXiv version is the partial submission, not the final paper. Camera-ready updates soon.
Cite this work
Felipe Tommaselli, , , R. V. Godoy, . Mind the Phase: Effective Rank and Representation Health in Legged Locomotion. CoRL 2026. Read partial submission (arXiv:2609.06958) ↗