Research released from 1–5 October 2026 focused on a shared weakness in motion systems: a result can look coherent while losing the detail, timing, or physical consequence that made the source movement useful. Four new arXiv preprints address that problem through source-preserving editing, multi-frequency tokens, physics-aware capture, and compliant partnered dance.

SuperMotion keeps compatible source movement during an edit

SuperMotion, released 1 October, edits human motion from text while explicitly reusing the source at every reverse-denoising step. A learned preservation gate decides which frames and feature dimensions should retain source information; a temporal high-frequency loss discourages smoothing away rapid changes. Fa-Ting Hong and Peter Wonka report 33.20% full-pool R@1 on MotionFix, alongside better preservation of motion dynamics.

The important distinction is between copying and preserving. The model does not freeze an entire input clip. It learns which source details remain compatible with the requested change.

FreqMo gives fast and slow movement different temporal scales

Rethinking Fixed Temporal Grids, released 2 October, argues that uniform motion tokens waste capacity on slowly changing trajectories while blurring fast events such as foot contacts. FreqMo separates wavelet frequency bands, preserves their temporal location, and encodes them through a shared codebook. The authors report threefold token-sequence compression while retaining high-frequency detail.

This is not simply a faster tokenizer. It makes the representation acknowledge that a turn across several seconds and a brief heel strike do not need the same temporal unit.

FlowHMR rewards motions a physics controller can actually track

FlowHMR, also released 2 October, treats monocular motion capture as conditional generation rather than direct regression. It first produces multiple 3D motion candidates, then post-trains with one reward for fidelity to the video and another for successful physics-based tracking. On the authors’ Wild-4K evaluation, physical tracking success rises to 82.47%, compared with 62.82% for the strongest reported baseline.

Because monocular depth is ambiguous, diversity is useful only if the selection rule remains visible. Here, “plausible” means both visually faithful and executable by the tested controller.

CoDance learns compliance from a two-person video

CoDance, released 4 October, retargets a single partnered-dance video into a humanoid reference and a moving partner. Multi-link compliance augmentation modifies the reference under forces at both hands, producing force-aware training data. The physical humanoid sustains two-hand dancing with repeated forward and backward transitions; in simulation, policies reproduce about 80% of the augmented demonstrations’ wrist displacement.

Together these papers move beyond smooth output as the default proof of quality. They ask what must survive transformation: unedited source dynamics, brief temporal events, physical trackability, or a partner’s force. That continues Somatic-AI Lab’s analysis of action-relative motion representations and the earlier guide to keeping evidence provenance visible.

References