A force-aware benchmark makes headlines by showing what vision models can't feel, and dance-conditioned music generation opens a new direction
From the arXiv
ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations (arXiv:2607.06052) Yu et al. (TU Munich) introduce a benchmark built from real human demonstrations with synchronized motion and force measurements across six physical interaction tasks. Motions are retargeted to humanoid references while recorded hand forces are replayed synchronously in simulation. The accompanying Force-Aware Tracking Score (FATS), plus tracking RMSE, survival rate, robustness ratio, and power overhead, evaluate policies under matched force vs. no-force conditions. Headline finding: leading vision-language-action models fall short of human-level reliability when humans physically guide robot arms — because VLA architectures are vision-centric and receive no direct information about applied force. → https://arxiv.org/pdf/2607.06052
Dance to Music Generation Leveraging Pre-training with Unpaired Data and Contrastive Alignment (arXiv:2607.10537) Kimura, Park, Polouliakh & Akama (Sony CSL / Keio / Georgia Tech) reverse the usual dance-music arrow, generating music conditioned on dance. Beat-guided contrastive pretraining aligns pretrained motion and music encoders so the system can exploit abundant unpaired data alongside scarce paired data; a ControlNet-style module conditions a pretrained text-to-audio diffusion model. Improvements on AIST++ in both dance-music alignment and audio quality. Notable for treating the moving body as the source rather than the follower. → https://arxiv.org/abs/2607.10537
In the Press
ThorArena drew unusually wide coverage for a robotics benchmark, appearing in general technology press under framings such as "humanoid robots can't handle human touch" and "new test measures how well humanoid robots handle real-world forces." The coverage centred on a structural point rather than a performance number: that vision-language-action models process what a robot sees rather than what it feels, and that contact state, friction, and slip onset are simply not encoded in camera frames. Commentary noted the implication for safety certification — that current evaluation regimes may not capture failure modes that appear only under physical human contact.
For the somatic-AI community, this is a mainstream-press articulation of the sensing-modality argument: some phenomena are not available to vision at any resolution.
From Import AI
Import AI (Jack Clark, mid-to-late July 2026) — Recent issues continue their tracking of capability acceleration and the institutional questions it raises. (Specific issue number and contents for this week could not be verified at the time of writing; check importai.substack.com directly rather than relying on this summary.) → https://importai.substack.com/
From X / Social
Robotics and embodied-AI community — ThorArena generated substantial discussion, with the "sees vs. feels" framing circulating widely. Several researchers connected it to the broader question of whether the current VLA paradigm — vision and language in, actions out — is missing a modality tier entirely, and whether force/tactile sensing needs to be a first-class input rather than an add-on. A recurring counterpoint: adding force sensing is hardware-expensive and data-scarce, which is precisely why vision-centric approaches dominate.
Music-AI and dance communities — the Sony CSL dance-to-music work drew interest from choreographers and live-performance technologists for its accompaniment framing: a system that generates music from a dancer's movement is directly usable in improvisation and rehearsal contexts where a live musician is unavailable.
Conference & Community Notes
Theme of the week — the modality question. ThorArena makes measurable what has been a conceptual argument: sensing modality determines which phenomena a system can engage at all. This is now an empirical result in robotics, not only a philosophical claim in somatics.
Venues:
- MOCO'26 — correction: The 10th International Conference on Movement and Computing has already taken place (23–25 April 2026, Cité des Arts, Montpellier, France), organised by EuroMov Digital Health in Motion with a special focus on health applications. Earlier entries in this digest listed it as a venue with dates to confirm; that was an error. Proceedings are published in two tracks: ACM Digital Library (doi:10.1145/3802842) and an Open track book of abstracts on HAL/Zenodo (hal-05665489, doi:10.5281/zenodo.20794033, CC BY 4.0).
- ECogS 2026 (OIST, Nov 9–13) — "Embodied cognition and AI."
- NeurIPS 2026 — notifications September.
- SIGGRAPH Asia 2026 — real-time and interactive motion work; check submission window.
Continue through the archive
For connected context, read This Week in Motion AI: The Force Blind Spot and The Robot That Can See Your Hand But Cannot Feel It.
References
Kimura, R., Park, S., Polouliakh, N., & Akama, T. (2026). Dance to music generation leveraging pre-training with unpaired data and contrastive alignment. arXiv:2607.10537. https://arxiv.org/abs/2607.10537
Yu, C., et al. (2026). ThorArena: Benchmarking humanoid physical interaction with human motion-force demonstrations. arXiv:2607.06052. https://arxiv.org/pdf/2607.06052