The strongest community signal between 29 July and 4 August 2026 is a shift from isolated motion toward complete perception–action systems. Google DeepMind released Gemini Robotics 2 for whole-body control, long-horizon planning and multi-robot coordination; ACE-Data-0 recorded human action as synchronized vision, motion, sound and touch; and MPIE-Bench exposed how highly rated image editors still fail when two bodies touch. Together, these releases make the same point from three directions: plausible movement is not enough when bodies must sense, coordinate and physically relate.
Scope note: this scan covers primary papers and official releases dated 29 July–4 August 2026. Hugging Face popularity is treated as evidence of platform activity, not peer review. Items without a traceable date or primary source were omitted.
Google DeepMind expands from robot arms to whole bodies
Gemini Robotics 2 — Google DeepMind, 30 July 2026. The release separates physical AI into three related models: a vision-language-action model that converts vision and instructions into motor control; an embodied-reasoning model that plans multi-step tasks and coordinates robots; and an on-device model designed to adapt to new hardware with fewer than 200 examples and a few hours of data.
The important change is whole-body coordination. DeepMind reports control of humanoid movement from feet to fingertips, including walking, crouching, reaching and manipulating objects within one task. Its own results also preserve a useful limit: multi-finger manipulation remains inconsistent, with reported success varying sharply by task. This is an official lab release rather than a peer-reviewed paper, so the capability claims should be read as company-reported results.
For movement research, the release extends the body-agnostic direction covered in the Lab's August frontier report: one intelligence layer is being adapted across different robot forms instead of being rebuilt around a single body.
→ https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
ACE-Data-0 records the action loop, not only the pose
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine — Cao and colleagues, arXiv:2607.28625, 30 July 2026; featured on Hugging Face Daily Papers on 31 July.
ACE turns furnished home environments into synchronized capture spaces. Its first dataset contains 150 hours and 17 million video frames across 75,000 interaction episodes, combining first- and third-person video, full-body and hand motion, object trajectories, audio and tactile signals. The authors report that current methods remain weak under contact, occlusion, egocentric movement and long time horizons.
That combination matters more than the headline scale. Most motion datasets isolate a skeleton or a video stream. ACE instead records parts of the loop by which a person sees, moves, touches and changes the environment. It still does not capture lived bodily experience, but it gives models a less fragmented exterior account of action.
→ https://arxiv.org/abs/2607.28625
MPIE-Bench catches errors that visual judges miss
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing — Lin, Du, Zhou, Xu and Xie, arXiv:2607.27616, 30 July 2026; featured on Hugging Face Daily Papers on 31 July.
The benchmark targets edited images of contact actions such as embracing, carrying and grappling. Across ten editors, the authors found fused limbs, invented extremities and bodies passing through one another even when vision-language-model checklists scored the same images above 0.95. Their mesh-based measures separately test anatomical completeness and whether body surfaces meet in the way the requested interaction requires.
This is evaluation infrastructure rather than a new generator, and that is precisely why it is useful. It demonstrates that a system can satisfy a semantic prompt — “two people embrace” — while failing the geometry of two bodies sharing space. It supplies a concrete counterpart to the Lab's argument that trained perception must shape what a system is asked to notice.
→ https://arxiv.org/abs/2607.27616
A fourth signal: motion as reusable dynamics
ShadowDancer — Cao, Meng and Zhang, arXiv:2607.28362, 30 July 2026; featured on Hugging Face Daily Papers on 31 July. The preprint learns an action from paired videos that preserve the same dynamics while changing appearance, then reuses that action in new environments. The claimed contribution is a frame-level control representation learned without action labels, motion estimators or task-specific fine-tuning.
This is early preprint evidence, but the direction is relevant: movement is being treated less as an object tied to one visible character and more as transferable dynamics. The unresolved question is what survives that transfer beyond visible timing and displacement.
→ https://arxiv.org/abs/2607.28362
References
Cao, J., Meng, Z., & Zhang, K. (2026). ShadowDancer: Teaching video world models any action by learning unified dynamics representations from a video and its shadow. arXiv:2607.28362. https://arxiv.org/abs/2607.28362
Cao, Y., Xie, H., Wen, B., Yao, R., Liu, Y., Huang, Y., Liao, Z., Wang, Y., Liu, H., Tian, X., Su, D., Zhuo, L., Tao, D., Wang, X., Pan, L., & Liu, Z. (2026). ACE-Data-0: Human-centric ambient capture as embodied data engine. arXiv:2607.28625. https://arxiv.org/abs/2607.28625
Lin, J., Du, M., Zhou, T., Xu, B., & Xie, H. (2026). MPIE-Bench: Benchmarking anatomically plausible multi-person interaction editing. arXiv:2607.27616. https://arxiv.org/abs/2607.27616
Parada, C. (2026, July 30). Gemini Robotics 2 brings whole body intelligence to robots. Google DeepMind. https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/