Four releases from 2–8 September 2026 challenge the assumption that motion must come from a visible skeleton. SoundMHPE estimates several people’s 3D poses from acoustic reflections; OVMAN tests navigation after objects have moved or disappeared; ArmPoser estimates an arm from one smartwatch without a calibration pose; and MotionBlind asks whether video-language models can distinguish speed, magnitude and direction in otherwise near-identical clips. Together they widen both the available signals and the tests needed to trust them.
SoundMHPE separates overlapping acoustic bodies
Sound-based Multi-Person 3D Pose Estimation — Oumi, Shibata, Irie, Kimura, Aoki and Isogawa; arXiv:2609.04902; 4 September 2026; accepted at ECCV 2026.
SoundMHPE uses multi-scale acoustic features and a temporal pose decoder to separate reflections associated with different people. The accompanying AMP dataset contains six hours and 432,000 synchronised acoustic–pose frames. The authors report improvement over their baselines. This is an estimate from reflected sound, not an audio recording that directly contains joint coordinates; room geometry and participant configuration are central deployment questions.
→ https://arxiv.org/abs/2609.04902
OVMAN asks robots to navigate through remembered change
OVMAN: A Task and Benchmark for Open-Vocabulary Motion-Aware Navigation — Ghosh; arXiv:2609.06424; 6 September 2026.
An agent visits a scene, returns after a scripted change and must follow instructions such as going to the chair that moved or the place where a vase used to be. The benchmark contains 219 two-visit episodes. The paper reports a 45.2% success rate for its reference agent versus 99.5% for an embodied oracle, with vacated locations especially difficult. Motion here is not a body trajectory but a change that must remain in memory.
→ https://arxiv.org/abs/2609.06424
ArmPoser works in the smartwatch’s native coordinates
ArmPoser: Real-Time, Calibration-Free Arm Pose Estimation from Smartwatch IMU — Dev, Xu, Loh, Gao, Hoffmann and Ahuja; arXiv:2609.08806; 8 September 2026.
ArmPoser trains directly on the coordinate axes produced by consumer smartwatches, avoiding explicit alignment and bone-offset calibration. Physically grounded augmentation varies watch placement and arm morphology, while a configuration module infers crown orientation and forearm placement. The authors report results on public data and a 10-participant, 30-activity study. “Calibration-free” removes a user step; it does not remove drift, placement variability or population limits.
→ https://arxiv.org/abs/2609.08806
MotionBlind tests whether video models read movement at all
MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs — Bhatia, Galoaa, Fritsche, et al.; arXiv:2609.09528; 8 September 2026.
MotionBlind pairs near-identical self-recorded clips that differ only in speed, magnitude or direction. A model earns an instance correct only if it answers four complementary questions consistently, producing a 6.25% chance floor. In the authors’ controlled study, open models remain near that floor; more frames and smarter sampling do not close the gap. The result cautions against using video-language models as motion evaluators simply because they describe scenes fluently.
→ https://arxiv.org/abs/2609.09528
What to watch next
The week’s projects trade setup burden for model dependence: fewer cameras, less calibration and more semantic flexibility all require stronger tests of generalisation. The archive’s guide to choosing movement signals and explanation of frame-sampling limits provide practical context. Better access is meaningful only when the system continues to distinguish observation, inference and generation.
References
Bhatia, D., Galoaa, B., Fritsche, O., et al. (2026). MotionBlind: Probing the illusion of motion understanding in Video-LLMs. arXiv:2609.09528. https://arxiv.org/abs/2609.09528
Dev, B., Xu, V., Loh, X.-A., Gao, C., Hoffmann, H., & Ahuja, K. (2026). ArmPoser: Real-time, calibration-free arm pose estimation from smartwatch IMU. arXiv:2609.08806. https://arxiv.org/abs/2609.08806
Ghosh, D. (2026). OVMAN: A task and benchmark for open-vocabulary motion-aware navigation. arXiv:2609.06424. https://arxiv.org/abs/2609.06424
Oumi, Y., Shibata, Y., Irie, G., Kimura, A., Aoki, Y., & Isogawa, M. (2026). Sound-based multi-person 3D pose estimation. arXiv:2609.04902. https://arxiv.org/abs/2609.04902