Four preprints released from 25–27 September 2026 show motion control moving deeper into model design. Rather than asking a prompt to carry every instruction, the new systems place style intensity, partial observations, social affect, or transformation history inside the representation itself.
A continuous slider for motion style
Motion Style Slider, released 25 September, constructs a direction between content and style in a learned embedding and exposes a scalar intensity control. Chen-Chieh Liao and colleagues train only on endpoints, not labelled intermediate intensities, while testing monotonicity, content preservation, realism, and extrapolation. The practical claim is deliberately narrower than a universal style scale: artists need a reliable local control axis, not an objective unit of “more style.”
One masked representation for several motion tasks
MotionMaestro, released 26 September, treats text guidance, pose conditioning, completion, interpolation, trajectory control, and continuation as different observation masks over the same motion sequence. Its masked tokenizer, observation map, and observation loss preserve the supplied conditions. The authors report state-of-the-art results across tasks on RoMo and MotionMillion, but the larger significance is architectural: task boundaries become patterns of known and unknown movement rather than separate models.
Real-time affect reaches a physical humanoid
SocialHumanoid, released 27 September, generates a full-body co-speech motion window in one forward pass, conditions the next window on motion history, and retargets the result to a robot controller. The accompanying AffectMoCap dataset contains four hours from two professional actors with speech, body and hand motion, and emotion labels. The paper reports roughly six-times faster inference than GestureLSM under its protocol and stable long-horizon robot execution. The small actor pool remains an important limit on claims about affective diversity.
Sign-language transfer becomes traceable
Traceable Human-to-Humanoid Sign Language Benchmarking, also released 27 September, contributes 20,648 Chinese Sign Language sequences with four aligned stages: source, repaired human motion, direct robot reference, and geometry-repaired robot reference. Separate scores for handshape, location, palm orientation, and inter-hand relation make it possible to locate errors introduced by repair, retargeting, geometry, or control.
The shared shift is from output inspection to controllable and auditable transformations. That extends Somatic-AI Lab’s earlier account of motion tokenizers as dictionaries and its review of evidence provenance in embodied sensing. A representation is useful not only when it compresses motion, but when it exposes what a practitioner can change and what an evaluator can trace.
References
- Liao, C.-C., et al. (2026). Motion Style Slider: Endpoint-supervised continuous style control for human motion diffusion. arXiv. https://arxiv.org/abs/2609.30795
- Chen, Y., Kim, M., & Do, J. (2026). MotionMaestro: Masked tokenization for unified motion generation. arXiv. https://arxiv.org/abs/2609.37495
- Yang, C., et al. (2026). SocialHumanoid: Towards expressive humanoid behavior via one-step co-speech motion generation. arXiv. https://arxiv.org/abs/2609.33311
- Liu, A., et al. (2026). Traceable human-to-humanoid sign language benchmarking. arXiv. https://arxiv.org/abs/2609.33354