Provocation

What if the most useful output from a movement model were not its best answer, but three defensible alternatives that force a practitioner to discover what they meant?

This brief proposes a concrete experiment: a Somatic Counterfactual Editor. A practitioner records a short movement phrase, requests one change and identifies one or two relationships that must remain. The system produces three distinct edits satisfying those constraints. The practitioner then performs each version, not merely watches it, and chooses which one preserves the phrase’s felt logic.

The proposal is speculative. Current motion editors can alter body part, amplitude, timing, action and style; current interactive dance systems can present selectable phrases. No cited work demonstrates that re-embodied comparison can elicit a practitioner’s latent preservation constraints or improve subsequent generation. That is the claim this experiment would test.

Argument

Instruction-driven movement editing is underdetermined. “Make the reach softer” names a desired difference but not a complete trajectory. Softness might be produced by slowing the approach, redistributing support, reducing abrupt acceleration, changing the hand’s contact, allowing more sequential articulation, or some combination. Several outcomes may satisfy the words, while only some preserve what the mover considers essential.

A single-output interface conceals this ambiguity. It encourages the user either to accept the model’s interpretation or to keep rewriting the prompt. A choice set makes the ambiguity inspectable. By holding named constraints constant while varying the solution, it turns generation into a question: Is this what you meant to preserve?

The practitioner’s body is part of the evaluation instrument. Watching can reveal shape, timing and gross dynamics. Re-performing can reveal an altered support pattern, an unexpected effort concentration or a transition that no longer prepares the next action. The experiment does not treat felt judgement as universal truth. It treats it as context-specific evidence about whether an edit serves the person and practice for which it was requested.

Evidence

Three current technical developments make the experiment feasible.

UniMoFlow edits 3D human motion across body-part, amplitude, temporal, action and style dimensions while anchoring unmodified content to the source. Its semantics-aware evaluation explicitly recognises that a valid edit can differ from one reference answer. CustomDance places human selection inside a coarse-to-fine workflow: it identifies temporal anchors, retrieves candidate phrases and generates transitions around chosen options. MRBench shows that motion–language systems change performance when descriptions move from concise to fine-grained, supporting the need to record precisely how practitioners describe the same movement at different levels.

Motor-control evidence supplies the experimental logic. The principle of motor abundance distinguishes variation that changes a task-relevant result from variation that leaves that result stable. Reviews of motor learning further distinguish unwanted noise from variability that supports exploration. This evidence does not prove that model diversity maps to useful bodily variation; that uncertainty determines the study design.

Phase 1 — Elicit the edit and the invariants

Recruit 12–20 experienced movement practitioners across at least two practices. Each selects a 10–20-second phrase they can repeat comfortably. Record optical motion plus video; add lightweight inertial sensors only if they do not materially change the phrase.

For each trial, the practitioner provides:

  • one edit instruction, such as “delay the turn,” “make the reach less effortful” or “keep the gesture but make the pathway more indirect”;
  • one outcome invariant, such as final orientation, contact point or musical arrival;
  • one relational invariant, such as pelvis initiating before the arm, continuous connection to the floor, or the turn preparing the descent.

The facilitator asks for observable clarification without forcing every experiential term into a joint-angle rule. The record retains both the practitioner’s language and the measurable proxy chosen for generation.

Phase 2 — Generate a calibrated choice set

Use an instruction-driven motion editor as the base. For each source phrase, generate a larger candidate pool, then select three outputs that pass physical and task constraints while maximising meaningful difference from one another.

The three options should occupy controlled diversity bands: conservative, intermediate and exploratory. All must satisfy the requested edit and measurable invariant thresholds. Reject candidates with foot sliding, implausible contacts, balance failures or changes to protected temporal anchors. Present the options without ranking and randomise their order.

A single “best-score” edit serves as the control condition. This makes the central comparison testable: does a structured choice set elicit information that iterative work with one optimised answer does not?

Phase 3 — Re-embody, choose and regenerate

The practitioner watches, learns and performs each option in counterbalanced order. After every attempt, collect:

  • perceived preservation of the original phrase;
  • satisfaction of the requested change;
  • effort, comfort and continuity ratings;
  • a short spoken explanation of what was preserved or lost.

The practitioner selects one option or rejects all three. Convert the explanation into a revised constraint only after the practitioner confirms the interpretation—for example, “keep the pelvis–arm delay within this range,” not the broader and unsupported claim “this number measures softness.” Generate a second choice set using the confirmed constraint and repeat the assessment.

One week later, repeat a subset of trials to test choice stability. Include unseen phrases to test whether constraints transfer or merely fit one recording.

What Success Would Mean

The experiment succeeds if the choice-set condition, compared with the single-output control:

  1. produces more specific, reusable preservation constraints;
  2. improves second-round ratings for both requested change and phrase preservation;
  3. yields above-chance within-practitioner choice consistency on repeated trials;
  4. maintains physical validity and genuine diversity rather than offering cosmetic variants; and
  5. allows “reject all” often enough to show that the interface does not force agreement.

The falsifiable prediction is that re-embodied comparison will improve the next generation only when the practitioner can confirm a constraint that distinguishes the chosen option. If ratings improve equally without re-embodiment, or if choices are unstable and explanations do not transfer, the proposed somatic feedback loop has not added evidence beyond ordinary visual selection.

Success would not show that the model understands softness, intention or possibility as a person does. It would show something narrower and useful: bodily comparison can expose task-relevant constraints that a one-shot instruction leaves unspecified.

Invitation

The field has spent years making motion generators produce a convincing answer. This experiment asks whether they can instead support a disciplined conversation about what counts as the same movement after change.

For choreographers, somatic educators and movement researchers, the invitation is practical: bring a phrase for which several visible edits would be acceptable but only some would remain coherent in performance. The crucial data are not simply which clip wins. They are the distinctions that become articulable only after the body tries the alternatives.

This extends the Lab’s proposal for kinaesthetic calibration from personalising a system to an individual body toward personalising what an edit must preserve. It also operationalises the evaluation problem raised in how to grade a movement: the criterion is not generic realism alone, but successful change within a practitioner-confirmed field of invariants.

References

Dhawale, A. K., Smith, M. A., & Ölveczky, B. P. (2017). The role of variability in motor learning. Annual Review of Neuroscience, 40, 479–498. https://doi.org/10.1146/annurev-neuro-072116-031548

Hua, Y., Jing, B., Zheng, C., Zhou, H., Luo, Y., & Yang, W. (2026). UniMoFlow: Grounding instruction-driven 3D human motion editing in generation. arXiv:2608.09143. https://arxiv.org/abs/2608.09143

Latash, M. L. (2012). The bliss (not the problem) of motor abundance (not redundancy). Experimental Brain Research, 217(1), 1–5. https://doi.org/10.1007/s00221-012-3000-4

Liu, F., Xu, L., Yang, C., Zhang, Y., Yan, Y., & Yang, X. (2026). MRBench: A comprehensive benchmark for human motion-text retrieval. arXiv:2608.07993. https://arxiv.org/abs/2608.07993

Tang, X., Yang, K., Guo, X., Balakrishnan, P., & Alghofaili, R. (2026). CustomDance: Customized 3D dance generation with coarse-to-fine human-centered interactive control. arXiv:2608.06722. https://arxiv.org/abs/2608.06722