Two developments close out July 2026. Two2Four, from DisneyResearch|Studios and ETH Zürich (arXiv:2607.26108, 28 July), transfers ordinary human motion onto quadruped animals using a diffusion model trained purely on animal motion — a transfer with no biomechanical correspondence, which makes it a useful test of what survives when body plans do not match. Separately, work on 3D human-human interaction generation (arXiv:2606.24255) reformulates two-person generation as social structure modelling rather than as two solo motions combined. For partnered movement practice, the second is the more consequential shift, because it makes the relationship between movers the primary object rather than a constraint applied afterwards.

This digest covers two papers with public arXiv records; publication dates and affiliations are given for each.

Two2Four: human movement, quadruped body

Paper: Zargarbashi, F., Qiu, Z., Agrawal, D., Coros, S., Sumner, R. W., Guay, M., & Buhmann, J. (2026, 28 July). Two2Four: Generative quadruped puppeteering from human motion. arXiv:2607.26108. DisneyResearch|Studios Switzerland and ETH Zürich.

The system converts ordinary human motion data into quadruped motion. Its architecture is a two-stage diffusion model trained purely on quadruped motion data — not on paired human-animal examples — combined with a structured conditioning and inpainting strategy. It covers walking, running, jumping, sitting and lying, and supports fine-grained control including head movement and puppeteering of individual limbs. The authors report improved motion realism and controllability relative to existing retargeting approaches, aimed at animation and virtual production.

Training only on quadruped data is the notable design choice. The model is not learning a mapping from human to animal movement; it is learning what plausible quadruped movement looks like, then being steered toward configurations that correspond to a human performer's input. Plausibility comes from the animal data; direction comes from the human.

Why this matters for somatic AI. Two limbs to four is transfer across body plans with no biomechanical correspondence — a limit case that isolates the question of what actually crosses. It cannot be the movement itself, since the mechanics are incompatible. What crosses is closer to expressive intention: timing, dynamic quality, directional emphasis, the shape of an impulse. That is a real form of translation and it is worth naming precisely, because such systems are easily described as an animal "performing" a human's movement. It is not that. It is a quadruped generator steered by a human's expressive input, and the felt experience of the performer does not transfer. The Lab examined the analogous question for human-to-human transfer in its explainer on what is preserved and lost when movement moves between bodies.

Social structure in two-person interaction generation

Paper: Wang, Z., Wang, B., Bian, Y., Wang, P., Wang, Z., Dong, D., Li, H., Mo, H., & Sun, Z. Social structure matters in 3D human-human interaction generation. arXiv:2606.24255.

This work formulates text-driven 3D human-human interaction generation as social structure modelling and grounding. Rather than generating two individual motions and reconciling them, it models the structure of the interaction itself as the object being generated. The current framework covers two-person interactions, with the authors noting applications across animation, virtual avatars, embodied AI, social robotics and human-computer interaction.

Why this matters for somatic AI. This is the closest existing formulation of a premise central to partnered movement practice: that in a duet the relationship is the material, not a constraint imposed on two independent performances. In Contact Improvisation, partnering and most ensemble forms, what a mover does is not separable from what is happening between the movers. A generation architecture that treats the interaction structure as primary is structurally aligned with that reality in a way that combine-two-solos approaches are not.

Two qualifications. First, the modelling here is of social structure grounded in text description — roles, relations, interaction type — rather than of physical contact dynamics, which is a different problem addressed by contact-aware synthesis work. Second, and as with everything in this month's research, the signals involved are positional. The mutual physical attunement that partnering practitioners work with involves force, weight and anticipation that positional data does not carry; the Lab's proposal for an anticipatory rather than reactive movement partner addresses that gap directly.

References

Wang, Z., Wang, B., Bian, Y., Wang, P., Wang, Z., Dong, D., Li, H., Mo, H., & Sun, Z. (2026). Social structure matters in 3D human-human interaction generation. arXiv:2606.24255. https://arxiv.org/abs/2606.24255

Zargarbashi, F., Qiu, Z., Agrawal, D., Coros, S., Sumner, R. W., Guay, M., & Buhmann, J. (2026). Two2Four: Generative quadruped puppeteering from human motion. arXiv:2607.26108. https://arxiv.org/abs/2607.26108