A new benchmark just measured something movement practitioners have always known: there are things about touch that no camera will ever capture.
Put your hand on someone's shoulder and guide them gently to turn. Almost nothing about what just happened is visible.
A camera would record two bodies and a hand making contact. What it would not record is the thing that actually did the work: the pressure you applied, whether it was firm or suggestive, whether their body yielded or resisted, the small negotiation of force between you that told each of you what the other intended. All of that was transmitted through touch — through force, friction, and the felt resistance of another body — and none of it left a visual trace worth speaking of.
This month, a team at TU Munich built a benchmark that measures exactly this gap, and the results are unusually clear. It turns out that the most advanced robot control systems in the world — the ones that read camera images and language instructions and produce movement — fail systematically the moment a human physically touches them. Not because the systems are badly built, but because of something more fundamental: they process what the robot sees, not what the robot feels.
What the Benchmark Did
The benchmark, called ThorArena, took an unusual approach. Rather than testing robots on how accurately they could reproduce a movement in isolation, the researchers recorded humans performing physical interaction tasks while measuring both their motion and the forces they exerted through their hands — two synchronized streams, position and force together.
They then replayed those recorded forces onto robots in simulation while asking the robots to track the corresponding motions. The question was simple: when a real, measured human force is applied to you, can you still do what you were doing?
For the leading systems, the answer was frequently no — or not reliably enough. The researchers introduced a new score to measure it (a "Force-Aware Tracking Score"), because the existing measures, which only looked at how accurately a robot followed a path, could not see the failure at all. A robot could score beautifully on the old metrics and still fall apart the moment a hand pressed against it.
Why This Isn't Just a Bug
The natural reaction is to assume this is a fixable shortcoming — better training, more data, a smarter model. But the diagnosis the researchers and commentators converge on is more structural than that, and it is worth stating carefully.
The dominant architecture in robotics right now takes camera images and language instructions as input and produces motor commands as output. It is, by design, vision-centric. And here is the problem: the things that matter in physical contact — how much force is being applied, whether the surfaces are gripping or beginning to slip, the state of the contact itself — are simply not present in a camera frame. They aren't hard to extract. They aren't there.
So when a human hand pushes against a robot arm, a vision-based system receives no direct information about that push whatsoever. It can try to infer something from how the scene changes afterward, but it is inferring backwards from consequences, not perceiving the force itself. The information was never in the input.
This is the difference between a measurement being difficult and a measurement being absent. No improvement in camera resolution, model size, or training data will put force information into a channel that does not carry it. The gap is about which channel you are listening to, not how well you are listening.
What Movement Practitioners Already Knew
If you have spent time in any practice built on physical contact — Contact Improvisation, partnering, martial arts, bodywork, or simply years of dancing with other people — this finding will sound less like news than confirmation.
Anyone who works with touch knows that watching a physical interaction and being in one are entirely different epistemic situations. You can observe two dancers in contact and describe what you see with great precision, and still have almost no access to what they are actually doing: where the weight is traveling, who is supporting whom, whether the contact is inviting or bracing, the constant micro-negotiation of pressure that constitutes the exchange. That information is available through the hands, through weight, through the felt sense of another body — and it is largely invisible.
This is why contact practices are taught through touch rather than demonstration. A teacher can show you what a movement looks like, but they cannot show you what supporting someone's weight feels like. They have to put their hands on you, or put you in contact with a partner, because the knowledge lives in a channel that watching does not access.
What ThorArena adds is not the insight but the measurement. The somatic claim — that touch carries information vision cannot — is now a quantified robotics result with a metric attached. That matters, because arguments that were previously philosophical can now point at a benchmark.
The Wider Lesson
The reason this generalises beyond robotics is that the same question applies to every system built to engage with human movement: which channel is it listening to, and does the phenomenon actually live there?
For force and contact, ThorArena has now shown the answer for vision: no. For the qualities this series has tracked all year — the felt effort behind a gesture, the intention forming before movement becomes visible, the interior organisation that distinguishes an integrated movement from an assembled one — the answer is likely the same, for the same structural reason. These live in channels that cameras do not carry: in muscle activity, in force, in the body's own interior signals.
None of which makes vision useless. Cameras are extraordinary at what they capture, and most of what AI systems know about human movement they learned by watching. But the honest conclusion from this month's result is that watching has a boundary, and the boundary is not a matter of resolution. Some things about bodies are only available to systems that touch them, sense them, or read the signals they generate from within.
The robot can see your hand perfectly. Feeling it is a different problem — and it requires a different sense.
Continue through the archive
For connected context, read This Week in Motion AI: The Force Blind Spot and Community Digest: Week of 14–20 July 2026.
References
Yu, C., et al. (2026). ThorArena: Benchmarking humanoid physical interaction with human motion-force demonstrations. arXiv:2607.06052. https://arxiv.org/pdf/2607.06052
Paxton, S. (1975). Contact improvisation. The Drama Review, 19(1), 40–42.