Publication note (2026-08-03). This page previously contained only an internal planning summary rather than an article — an editorial error that went uncorrected for several months. The full article has now been written and published here, preserving the original publication date and URL. The Lab regrets the lapse.
When an AI system learns human movement, it converts that movement into coordinates: a list of numbers describing where each joint sits in space, frame by frame. This works remarkably well for reproducing how a movement looks. What it cannot capture is the dimension somatic practice is built on — the felt, first-person experience of moving, which is not a property of joint positions at all. This article explains what that conversion keeps and what it discards, why the gap is structural rather than a temporary limitation, and what it means for anyone using movement AI in therapy, education or research.
What a movement becomes inside a model
Start with what actually happens to a movement when a system learns it.
A motion capture system records a body as a set of tracked points — typically somewhere between 17 and 50 joints. At each moment, each joint gets three numbers: its position along three spatial axes. A one-minute sequence at 30 frames per second becomes a table of roughly a hundred thousand numbers.
Generative models compress this further. Systems such as the Human Motion Diffusion Model (Tevet et al., 2022) learn to represent movement in what researchers call a latent space — a compact mathematical space in which each point corresponds to a possible movement and similar movements sit near one another. The model builds this space by processing large quantities of recorded motion and finding the compressed representation that best allows it to reconstruct what it was shown.
This is a genuine achievement. A well-trained latent space captures real structure: it encodes that walking and running are related, that a reach has a beginning and a completion, that bodies move continuously. Movements can be generated, blended and interpolated in ways that look convincing.
But notice what defined the whole process. The model was optimised to reconstruct the recorded coordinates. Everything it learned, it learned because doing so helped reproduce where joints were. Anything that does not affect joint position was, from the system's point of view, never part of the problem.
The gesture and its double
Here is the concrete case that makes the gap visible.
Take a dancer with twenty years of training performing a gesture they have refined across that entire period — say, a slow lift of the arm that they execute with a specific quality of attention, weight and internal support. Record it with high-quality motion capture. Feed it to a generative model. Ask the model to reproduce it on a synthetic skeleton.
The result will be, in coordinate terms, close to perfect. The arm follows the same path at the same speed. A side-by-side comparison of the two trajectories would show minimal error.
And yet what made the original gesture what it was is missing. The organisation of effort through the torso that supported the arm. The relationship between the movement and the breath. The quality of attention directed into the reach. The twenty years of correction, injury, adaptation and refinement sedimented into how that particular body solves that particular problem.
None of this was in the recording, so none of it could be in the model. The capture system measured position. The dancer's expertise lives in how position was produced — and two very different productions can yield nearly identical positions.
This is what the phenomenological tradition has argued for decades. Merleau-Ponty (1945/2012) described the body not as an object we command but as the medium through which we have a world at all. Varela, Thompson and Rosch (1991) built a cognitive-science research programme on the claim that mind is enacted through a living body rather than computed over representations of one. Gibson (1979) made a related point from the perception side: what an organism perceives is not raw geometry but affordances — possibilities for action defined relative to that organism's own body. In each case the argument is the same. A third-person description of a body's configuration is not the same kind of thing as the first-person activity of being that body.
Why this is structural, not temporary
The natural response is to assume better technology will close the gap. More markers, higher frame rates, larger models.
It will not, for a reason worth stating precisely: the information was never in the input. A camera or marker system measures where the body is. The felt quality of a movement is not a fact about where the body is. Increasing the resolution of a position measurement yields a more precise position measurement, not a different kind of measurement.
This is the difference between something being hard to detect and something being absent from the channel. Improvements in the first do not touch the second. Any system whose only input is positional will remain, in principle, unable to represent what position does not encode.
Three places this matters in practice
Rehabilitation and therapy feedback. Movement AI is increasingly used to assess whether a patient is performing an exercise correctly. Positional accuracy is a real and useful signal, but it is not the same as neurological integration or safe movement organisation. A patient can trace the prescribed path while recruiting the wrong muscles, bracing through a joint, or compensating in ways that cause problems later — and a system evaluating position will register success. Research on virtual-reality neurorehabilitation has argued that embodiment and body perception need to be treated as first-class ingredients of treatment design rather than as by-products of correct trajectories (Perez-Marcos et al., 2018).
Somatic education. The core skill a somatic teacher develops is perceiving how a student organises a movement, not merely whether the shape is right. A tool that evaluates shape can tell a student they achieved the position. It cannot tell them they achieved it through unnecessary effort — which is frequently the entire content of the lesson.
Research datasets. This is the compounding one. What is not encoded cannot be represented, and what is not represented cannot be studied. If the datasets underpinning movement AI record position only, the research programme built on them inherits that boundary, and the qualities somatic practice attends to become institutionally invisible rather than merely unmeasured.
Three directions, honestly assessed
Multimodal physiological training. Recording signals closer to how movement is produced — muscle activity, force, inertial dynamics — alongside position. This is the most direct response and genuinely narrows the gap, because these signals carry information about effort and organisation that position does not. It does not close the gap: measuring a muscle's electrical activity is not measuring what the movement felt like.
Practitioner annotation. Having trained practitioners label movement qualities, creating a training signal for distinctions machines cannot currently make. Promising, and epistemologically complicated: experts disagree, somatic traditions use different vocabularies, and some trained discriminations resist explicit specification. Annotation makes perception transferable only to the extent that perception can be articulated.
Augment rather than replace. The most modest framing, and possibly the most durable: build tools that extend what a practitioner can record, compare and revisit, while leaving perceptual judgement with the human. This asks less of the technology and more of the working relationship, and it does not pretend the gap is closed.
What to hold onto
None of this makes movement AI useless. Position data supports genuinely valuable things — archiving, comparison across time, analysis at scales no single observer could manage.
The honest position is narrower and more useful than either enthusiasm or dismissal. These systems capture a real and partial aspect of movement: its outward geometry. The aspect somatic practice is organised around — how movement is produced and how it feels from within — is not in that channel, and no refinement of the channel will put it there. Knowing precisely which part you have is what makes the part you have usable.
Continue through the archive
For connected context, read Frontier Scout Report | 2026-03-31 → 2026-04-06 and Community News — Embodied AI & Dance Technology (2026-04-07).
References
Gibson, J. J. (1979). The ecological approach to visual perception. Houghton Mifflin.
Merleau-Ponty, M. (2012). Phenomenology of perception (D. A. Landes, Trans.). Routledge. (Original work published 1945)
Perez-Marcos, D., Bieler-Aeschlimann, M., & Serino, A. (2018). Virtual reality as a vehicle to empower motor-cognitive neurorehabilitation. Frontiers in Psychology, 9, 2120. https://doi.org/10.3389/fpsyg.2018.02120
Tevet, G., Raab, S., Gordon, B., Shafir, Y., Cohen-Or, D., & Bermano, A. H. (2022). Human motion diffusion model. arXiv:2209.14916. https://arxiv.org/abs/2209.14916
Varela, F. J., Thompson, E., & Rosch, E. (1991). The embodied mind: Cognitive science and human experience. MIT Press.