Four releases from 5–11 August 2026 widen what counts as motion data. VSMP-IMU turns video into controllable synthetic wearable signals; HOPE estimates changing hand pressure from ordinary monocular video; a contextual-auditing paper asks how motion-capture skeletons should be judged when “ground truth” is contested; and a soccer model learns by representing several plausible futures rather than one. The common signal is methodological: movement research is becoming less satisfied with a single visible skeleton as either input, output or truth.

Scope note: this scan uses public arXiv submission records in the stated window. Three items are preprints; the auditing paper is listed for AIES 2026. Reported performance is author-reported and should not be read as independent replication.

VSMP-IMU generates wearable signals from structured movement

VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation — Ray, Rey, Liu, Lukowicz and Zhou; arXiv:2608.05782; 6 August 2026; under review.

Labelled inertial-measurement-unit data are expensive to collect across many people, activities and sensor placements. VSMP-IMU extracts a structured Semantic Motion Program from video, varies the program without changing the activity label, synthesises movement, converts that motion to virtual IMU readings and then adapts the signals to a target wearable domain.

Across five public activity-recognition datasets, the authors report an average Macro-F1 of 78.33%, 9.77 percentage points above real-only training and 4.04 points above their strongest synthetic baseline. The larger conceptual contribution is the separation between activity-defining semantics and label-preserving variation. It offers a way to generate diversity without treating every difference between bodies as a different action.

For somatic research, synthetic IMU data could reduce capture burden, but only after calibration against real wearers. A virtual signal derived from visible motion may reproduce accelerations while missing soft-tissue effects, device placement and the practitioner’s lived effort. The Lab’s guide to what movement data actually capture remains the relevant caution.

https://arxiv.org/abs/2608.05782

HOPE estimates where a hand presses and how pressure changes

HOPE: Hand-Object Pressure Estimation from Monocular Videos — Jeon, Kim and Joo; arXiv:2608.06192; 6 August 2026.

HOPE treats pressure estimation as a video problem centred on the hand. It maps pressure and contact annotations from tactile gloves, planar sensors and distance-based hand–object datasets into a shared hand-mesh space. A video transformer then tracks persistent hand vertices over time and predicts contact and normal pressure, enforcing zero pressure where there is no contact.

The paper reports transfer from metric pressure supervision collected mainly with gloved hands to bare-hand, first-person and in-the-wild videos. That does not turn a camera into a tactile sensor: the model infers pressure from learned visual and pose patterns. Even so, it expands video analysis from “where did contact occur?” toward “how was force distributed over time?”

This is directly relevant to the archive’s account of the force blind spot in motion AI. HOPE addresses a narrow but important part of that blind spot at the hand–object interface. It does not estimate the mover’s global effort, pain, intention or subjective intensity.

https://arxiv.org/abs/2608.06192

Contextual auditing challenges the idea of a neutral skeleton

Context and Symmetry in Auditing: A Case Study of Skeleton Inference in Motion Capture — Harvey, Moss, Sandhaus, Jacobs and Sloane; arXiv:2608.10194; 10 August 2026; listed for the AAAI/ACM Conference on AI, Ethics, and Society 2026.

The authors argue that an audit should examine measurements within the practices that produce them. This matters when the expected output is not self-evident: a motion-capture system infers a skeleton, but the choice of markers, model, calibration, body assumptions and intended use all shape what counts as correct. Their contextual audit makes those choices explicit. They also draw on “symmetry” from science and technology studies to compare human and technical accounts when ground truth is unknown or contested.

This is not another accuracy leaderboard. It is a method for asking who gets to define the body against which accuracy is measured. For practitioners, it creates a route to document a mismatch that may be meaningful in use even when a generic benchmark says the skeleton is acceptable.

https://arxiv.org/abs/2608.10194

Soccer representation learning keeps more than one future alive

Capturing Uncertainty in Human Motion for Representation Learning in Soccer — Xu, Bretzner, Wang and Maki; arXiv:2608.11203; 11 August 2026.

The model learns 3D skeleton representations by predicting future player motion. Instead of requiring one future trajectory, its conditioning module represents a probability distribution over discretised future motions. The authors argue that human movement is inherently uncertain and that multiple plausible futures are necessary to capture its dynamics. They report improved prediction and transfer to several downstream soccer tasks on large-scale tracking data.

The domain is constrained and tactical, not somatic practice. Yet the design choice is broadly important. At the instant a player shifts weight, more than one continuation may remain possible; collapsing that branch too early can erase information about responsiveness. A probability distribution is not the same as human possibility or choice, but it is a more faithful technical object than a single supposedly correct continuation.

https://arxiv.org/abs/2608.11203

What to watch next

The four projects distribute uncertainty across different layers: sensor variation, inferred pressure, contested measurement and future action. The opportunity is to connect those layers without pretending they are interchangeable. A future movement system might combine visible trajectory, inferred contact, wearable signals and a set of possible continuations. The corresponding risk is that every added channel creates another model-dependent inference that can be mistaken for direct access to the body. Better instrumentation increases responsibility for contextual validation; it does not remove it.

References

Harvey, E., Moss, E., Sandhaus, H., Jacobs, A. Z., & Sloane, M. (2026). Context and symmetry in auditing: A case study of skeleton inference in motion capture. arXiv:2608.10194. https://arxiv.org/abs/2608.10194

Jeon, S., Kim, B., & Joo, H. (2026). HOPE: Hand-object pressure estimation from monocular videos. arXiv:2608.06192. https://arxiv.org/abs/2608.06192

Ray, L. S. S., Rey, V. F., Liu, M., Lukowicz, P., & Zhou, B. (2026). VSMP-IMU: Video-grounded semantic motion programs for sensor-aware synthetic IMU generation. arXiv:2608.05782. https://arxiv.org/abs/2608.05782

Xu, Y., Bretzner, L., Wang, T., & Maki, A. (2026). Capturing uncertainty in human motion for representation learning in soccer. arXiv:2608.11203. https://arxiv.org/abs/2608.11203