Richard Shusterman's philosophical project of somaesthetics rests on a claim that cuts against an assumption widespread in motion AI: that bodily perception is not fixed equipment but a trainable capacity, and that improving it is a distinct undertaking from improving what is measured. This analysis argues that the distinction between measurement and perception is real and consequential, that motion AI has invested almost entirely in the first while treating the second as solved, and that this explains a specific pattern — systems that sense more while noticing less. The argument does not conclude that sensing is unimportant. It concludes that sensing and perceiving are different problems, and that only one of them is receiving attention.

Status. This is an interpretive analysis by Somatic-AI Lab. Shusterman's somaesthetics and Höök's soma design are established published positions, cited below; the application to motion AI architecture is the Lab's own argument.

Somaesthetics: the position

Richard Shusterman introduced somaesthetics as a philosophical discipline concerned with the body as a site of sensory appreciation and creative self-fashioning. Its distinguishing move is to treat the body not merely as an object of study but as the instrument through which experience is had — and, critically, as an instrument that can be improved.

Shusterman divides the field into three branches. Analytic somaesthetics studies bodily perception and practice descriptively. Pragmatic somaesthetics concerns methods for improving bodily awareness and function — the disciplines of somatic practice fall here. Practical somaesthetics is the actual doing: the training itself, not writing about it.

The pragmatic branch carries the philosophical weight relevant to this analysis. It asserts that bodily perception is not a fixed biological given. Two people in the same room with the same sensory apparatus do not perceive the same things, because perception is shaped by trained attention. A practitioner with years of somatic training perceives distinctions in another person's movement — where an impulse initiates, whether effort is held or released, whether the breath supports the phrase — that an untrained observer does not perceive at all. Not perceives-but-cannot-name: does not perceive.

This is an empirical claim about human capacity, and it is well supported across expertise research in music, medicine, sport and the somatic disciplines themselves. Expert perception is trained perception.

The distinction that matters

If perception is trainable, then two separable variables determine what any observational system can establish:

  1. What the instrument can detect — determined by sensing modality and resolution.
  2. What the observer is trained to notice — determined by the perceiver's cultivated discrimination.

Improving the first without the second yields more data and no more insight. This is not a hypothetical failure mode; it is the ordinary condition of a novice with an expensive instrument.

Motion AI has invested overwhelmingly in the first variable. The trajectory documented on this platform through 2026 is one of steadily improving measurement — markerless capture reaching marker-grade accuracy, muscle-level simulation validated against real electromyography, consumer-device reconstruction, force-aware benchmarking. Each is a genuine advance in what can be detected.

The second variable has no equivalent programme. There is no research effort devoted to the question of what a motion AI system should be trained to notice in the data it receives, in the sense of cultivated discrimination rather than task performance. The implicit assumption is that with sufficient data and a suitable objective, relevant distinctions will be learned automatically.

Why the assumption fails

The assumption is not absurd — it describes how deep learning has succeeded in many domains. But it fails in a specific and diagnosable way for movement quality.

A learned system discriminates the distinctions its training signal rewards. If the objective is reconstruction accuracy, it learns to discriminate what affects reconstruction accuracy. If the objective is human judgements of realism, it learns to discriminate what affects untrained human judgements of realism. In neither case does it learn distinctions that trained perceivers make but that are absent from the training signal.

Movement quality is precisely such a case. Two performances with near-identical joint trajectories can differ in initiation, whole-body connectivity, effort organisation and intention legibility — differences a trained observer reliably detects. Because those differences barely register in reconstruction error, and because untrained raters largely miss them, neither standard objective rewards learning them. The system is not failing to learn something available in its training signal; the distinction was never in the signal to learn.

This is why the field's most recent evaluation work is significant beyond its immediate contribution. Physical-perceptual fidelity metrics combine physical plausibility with modelled human perception — the two available axes. Both are exterior. Neither encodes what a trained perceiver notices, which this platform has argued constitutes a missing third axis of movement evaluation.

The design response

The clearest existing response to this problem comes not from AI research but from interaction design. Kristina Höök's soma design, set out in Designing with the Body (MIT Press, 2018), holds that designers must cultivate their own bodily sensitivity as a precondition for designing well — on the grounds that one cannot design for qualities of experience one has not learned to notice. The Lab's source-verified profile of Höök sets out the position in more detail.

Read alongside Shusterman, the implication for AI systems is specific. If the distinctions a system can learn are bounded by what its training signal encodes, and if the training signal is constructed by people, then the perceptual training of those people is a hard constraint on the system's ceiling. A team that cannot perceive a distinction will not encode it in an objective, will not annotate for it, and will not build an evaluation that measures it. The system inherits the perceptual limits of its designers, invisibly.

This reframes what somatic expertise contributes to somatic AI. The usual framing casts practitioners as domain users to be consulted about requirements. The somaesthetic framing casts them as the source of the perceptual discriminations the system's objectives must encode — a role upstream of requirements, at the level of what is measurable in principle.

Three limits of this argument

Intellectual honesty requires marking where the argument stops.

Trained perception is not infallible. Expert perception is real, but experts also disagree, and somatic traditions differ in their vocabularies and emphases. Treating trained perception as ground truth without inter-rater validation would import those disagreements as noise. Any evaluation built on trained perceivers needs the same methodological care as any other human-judgement protocol.

Not everything perceived is transferable. Some trained discriminations may be reliably made by experts while resisting explicit specification. If a distinction cannot be articulated or annotated consistently, it cannot enter a training signal regardless of how real it is. The argument establishes that trained perception exceeds current objectives; it does not establish that the excess is fully formalisable.

This is an argument about ceilings, not a method. Establishing that the designer's perception constrains the system does not yield a procedure for encoding cultivated perception into an objective function. That is unsolved, and this analysis does not solve it. The Lab's five-dimension assessment framework is a first attempt at making the relevant discriminations explicit enough to be applied consistently; it is a rubric, not yet a metric.

Conclusion

The agnosticism turn documented in this month's frontier report shows motion AI removing structural assumptions at speed — about skeletons, sensors, body plans. What remains untouched is the assumption that perception is a solved problem, that relevant distinctions will emerge from sufficient data and reasonable objectives.

Somaesthetics gives a reason to doubt this. If perception is trainable, then it is also trained — someone's perceptual capacities determined what the system was asked to distinguish. In motion AI, those capacities have mostly belonged to people trained to model bodies from outside, not to people trained to perceive movement from within. The resulting systems perceive what their designers perceive, which is exactly what one would expect, and exactly what the somaesthetic tradition would predict.

The corrective is not better sensors. It is the presence, in the design process itself, of perception trained on the phenomenon.

References

Höök, K. (2018). Designing with the body: Somaesthetic interaction design. MIT Press. https://direct.mit.edu/books/monograph/4131/Designing-with-the-BodySomaesthetic-Interaction

Shusterman, R. (2008). Body consciousness: A philosophy of mindfulness and somaesthetics. Cambridge University Press. https://doi.org/10.1017/CBO9780511802829

Shusterman, R. (2012). Thinking through the body: Essays in somaesthetics. Cambridge University Press. https://doi.org/10.1017/CBO9781139094030

Zhao, S., et al. (2025). PP-Motion: Physical-perceptual fidelity evaluation for human motion generation. arXiv:2508.08179. https://arxiv.org/abs/2508.08179