Three date-verifiable arXiv releases from 20–26 August 2026 point toward representations that are easier to deploy and inspect: PoseOFF anchors optical flow around human joints for early action anticipation; UCAG-P maps different robot bodies into a camera-centric action geometry; and DeMoDiff factorises human motion across joints and time for local editing. The community signal is not a new sensor. It is a move toward representations that preserve the part of motion a downstream user actually needs.

PoseOFF spends computation where the body moves

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming — Grundy, McCarthy and Fluke; arXiv:2608.25495; 26 August 2026.

PoseOFF conditions optical-flow features on detected pose, capturing local dynamics around semantically meaningful joints instead of processing every pixel equally. The authors report improved early action anticipation across benchmark datasets and backbones, with gains at lower observation ratios. The practical appeal is latency: a robot may need to infer a reaching or turning action before it is complete.

https://arxiv.org/abs/2608.25495

UCAG-P makes the camera the shared action space

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation — Xiaomi Embodied Intelligence Team and collaborators; arXiv:2608.26058; 26 August 2026.

UCAG-P represents manipulation through anchor motion observable in image and camera coordinates, then translates the shared prediction into embodiment-specific controls. The authors train on robot, simulation and human-demonstration data and report zero-shot results across several benchmarks. A camera-centred representation can reduce the burden of translating among robot morphologies, but it also makes depth and unobserved force explicit limitations.

https://arxiv.org/abs/2608.26058

DeMoDiff keeps joints available for editing

Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation — Yang et al.; arXiv:2608.23279; 24 August 2026.

Its joint-and-time factorisation supports spatial and temporal masks during generation. The contribution is relevant to creative tools because it makes “change this part during this interval” a representational operation rather than an instruction the model must guess from a whole-body latent.

https://arxiv.org/abs/2608.23279

What to watch next

All three projects make an observed or requested substructure more explicit: a joint neighbourhood, a camera-visible anchor, or a body-time slice. None should be mistaken for a complete account of movement. Somatic use still requires testing what disappears when a representation privileges observability, kinematics or edit masks. The archive’s guide to choosing body-sensing signals is a useful reminder that every interface begins by deciding what can count as evidence.

References

Grundy, L. de Z., McCarthy, C., & Fluke, C. (2026). Pose-anchored optical flow for low-latency human action anticipation in human-robot teaming. arXiv:2608.25495. https://arxiv.org/abs/2608.25495

Xiaomi Embodied Intelligence Team, Xu, S., Li, F., Zhan, G., et al. (2026). One policy, many embodiments: Unified camera-centric action geometry pre-training for heterogeneous embodied manipulation. arXiv:2608.26058. https://arxiv.org/abs/2608.26058

Yang, C., Xu, L., Li, Y., Liu, F., Gao, J., Zeng, W., & Yan, Y. (2026). Spatiotemporally decoupled autoregressive diffusion model for human motion generation. arXiv:2608.23279. https://arxiv.org/abs/2608.23279