Motion AI represents bodies as coordinates, rotations, trajectories and latent vectors. These are powerful descriptions, but they are not neutral translations of movement. Maxine Sheets-Johnstone’s phenomenology offers a precise challenge: movement is not first a shape that later acquires quality. Its qualitative dynamics—tension, timing, force, expansion and contraction—are how movement is lived and recognised in the first place.

Kinesthetic consciousness is primary

Sheets-Johnstone argues that kinaesthetic experience is not an interpretation added to a finished motor event. We encounter ourselves through moving. A reaching gesture is not a sequence of joint positions plus a subjective label; its sense unfolds through the changing dynamics of effort and direction. This claim does not make kinematics useless. It identifies what kinematics leaves unspecified.

Current datasets usually preserve where markers travelled and when. They may include action labels, captions or style tags, but those labels are external summaries. Two trajectories can be geometrically close while differing in hesitation, commitment or yield. Conversely, two bodies can realise a similar qualitative dynamic through different joint paths.

The architectural contact point

A motion manifold is often described as a lower-dimensional space of plausible poses or sequences. Diffusion and flow models learn a field that guides noisy samples toward regions supported by data. This is an effective statistical geometry. The phenomenological question is different: what coordinates would let a system distinguish a movement’s qualitative organisation without reducing it to a new decorative label?

Three gaps follow.

First, temporal thickness. Framewise pose and short clips privilege visible states. Yet a movement’s quality can reside in how preparation gathers before an event and how release continues after it. A trajectory token may preserve order while losing the felt gradient of anticipation.

Second, relational effort. Force is not a scalar attached to one joint. It is organised through support, breath, gravity, object contact and the mover’s history. A model can infer some correlates from video or wearables, but an inferred pressure value is not a direct account of effort.

Third, situated meaning. The same velocity profile can be playful, defensive or ceremonial depending on context. Language annotations help, but they inherit the vocabulary and authority of whoever wrote them. More samples of the same labels do not automatically add the missing perspective.

What each field gains

Phenomenology gains a technical partner for testing its distinctions. Instead of claiming that quality is absent from coordinates, researchers can ask whether adding force, contact, breath or practitioner descriptions improves discrimination under controlled perturbations. The model becomes an instrument for operationalising a philosophical distinction, not proof that the distinction has been solved.

AI gains a reason to separate representation from experience. A latent direction that reliably correlates with “soft” or “hesitant” movement is still a learned regularity, not softness or hesitation felt by the system. Keeping that asymmetry visible prevents anthropomorphic claims while enabling better tools.

A phenomenologically adequate benchmark

An adequate benchmark would not replace kinematic metrics. It would layer them. Start with a recorded phrase and create variations that preserve visible action while changing timing, support or effort. Ask practitioners to discriminate the variations, describe what changed and select which one preserves a stated intention. Then test whether a model’s representation predicts those judgements across bodies and contexts.

The benchmark must preserve disagreement. If practitioners diverge, that is evidence about situated meaning, not annotation noise to average away. The system should return alternatives and provenance: what was observed, what was inferred and what was generated. This echoes the archive’s account of multiple valid edits and critique of neutral skeletons.

Open questions

Can a representation encode qualitative dynamics without fixing one culture’s vocabulary as universal? Can a model learn personal movement signatures without turning a performer into a biometric profile? Can real-time feedback preserve the temporal thickness that offline evaluation sees? These are architectural, ethical and methodological questions at once.

Sheets-Johnstone does not provide a recipe for a better encoder. She provides a test for conceptual honesty. If a system claims to understand movement, we should ask whether it can account for how movement organises experience—not only whether its joints land near the recorded points.

References

Gallagher, S. (2005). How the body shapes the mind. Oxford University Press. https://doi.org/10.1093/0199271941.001.0001

Holden, D., Saito, J., & Komura, T. (2016). A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics, 35(4), Article 138. https://doi.org/10.1145/2897824.2925975

Sheets-Johnstone, M. (2011). The primacy of movement (2nd ed.). John Benjamins. https://doi.org/10.1075/aicr.82

Yang, C., Xu, L., Li, Y., et al. (2026). Spatiotemporally decoupled autoregressive diffusion model for human motion generation. arXiv:2608.23279. https://arxiv.org/abs/2608.23279