Robots Are Learning to Feel
Dexterous manipulation remains one of the biggest barriers keeping robots from successfully tackling a wide range of everyday tasks. A sense of touch could be the key, but a lack of quality data has h...
Researched and edited by Kiran Ch and the WhatIsFuture editorial team. Reviewed for factual accuracy before publication.
Over the past three years, I've spent almost every waking hour tracking artificial intelligence as it conquers the digital world at a terrifyingly fast pace. At WhatIsFuture.com, my inbox is flooded daily with press releases and research papers announcing massive multimodal transformers that analyze complex medical scans in milliseconds, write intricate software pipelines, and generate photorealistic video on command. We are living through an unprecedented explosion of digital cognitive power that continually defies my expectations. Yet, whenever I watch a $150,000 state-of-the-art humanoid robot try to perform a simple everyday task—like folding a freshly washed towel, picking up a soft ripe peach, or twisting open a childproof medicine bottle—I am reminded of a stark reality: digital intelligence and physical intelligence are two completely different beasts.
Right now, our most sophisticated robotic platforms are physically numb. They possess supercomputer brains fed by high-resolution visual sensors, yet they walk through the world like an amputee operating with thick, frozen leather gloves. In my view, we have hit a temporary ceiling in robotics not because our AI models aren't smart enough, but because we forgot how fundamentally crucial the sense of touch is to human intelligence. If we want machines to genuinely operate alongside us in our kitchens, hospitals, and factories, we must teach them how to feel.
Join Our Tech Community
Get instant alerts on the most critical AI breakthroughs on our WhatsApp channel. No spam, just signal.
The Vision-Only Trap: Why Modern Humanoids Are So Clumsy
If you have followed humanoid robotics over the last two years, you’ve likely seen breathtaking videos of machines doing backflips, walking over rough terrain, or sorting colored blocks. But as someone who closely analyzes these systems, I can tell you that many of these demos rely heavily on smoke and mirrors or hyper-controlled environments. The moment you place those same robots in an unpredictable human setting, their spatial competence deteriorates rapidly.
The primary reason for this failure is our obsession with computer vision. We have tried to solve physical interaction almost entirely through cameras. We feed camera frames into large vision-language-action (VLA) models and expect the robot to infer everything it needs to know about the physical world from photons striking a CMOS sensor. But vision only tells a robot where an object is; it tells the robot almost nothing about how that object behaves when handled.
Vision allows a robot to locate a glass of water on a table, but only tactile feedback tells the robot that the glass is slipping from its grip by two millimeters per second and requires an instant three percent increase in squeeze force.
When you or I reach for an egg, we don't calculate its exact structural integrity using our eyes. We touch it. Our nervous system receives real-time micro-feedback about pressure, friction, shear forces, and weight distribution. We make sub-millisecond micro-adjustments in our muscle tone before our conscious brain even realizes what happened. When a robot relies purely on visual loops operating at 30 or 60 frames per second, it suffers from a latency gap that turns soft interactions into catastrophic failures. The robot either crushes the delicate object because it didn't feel the resistance, or drops it because it was afraid to grip tightly enough. This phenomenon is a stark modern manifestation of Moravec's paradox: tasks that are computationally hard for humans (like playing grandmaster chess) are easy for AI, but tasks that are effortless for a human toddler (like holding a slipper without squishing it) are monstrously difficult for machines.
The Hardware Revolution: From Rigid Metal to Electronic Skin
Fortunately, a quiet revolution is happening in hardware laboratories across the globe. Researchers are realizing that to solve physical AI, we must engineer synthetic mechanoreceptors. In my recent conversations with materials scientists and robotics engineers, I’ve seen a dramatic shift from rigid metallic end-effectors toward bio-inspired soft robotics and electronic skin (e-skin).
Human skin is an engineering masterpiece. It contains roughly 17,000 tactile units in the palm alone, capable of sensing light touch, deep pressure, lateral friction, vibration, and thermal conductivity. Replicating this in silicon and polymer has been one of the toughest challenges in robotics, but several promising technologies are finally reaching commercial maturity:
- Optical Tactile Sensors (GelSight and DIGIT): These sensors place a camera inside a hollow finger pointing at a flexible rubber gel coated with reflective paint. When the finger presses against an object, the interior camera measures the exact micro-deformation of the gel. This effectively turns a tactile signal into an ultra-high-resolution height map, allowing a robot to "see" textures and contact geometry directly at the point of impact.
- Piezoelectric and Piezoresistive E-Skins: Arrays of flexible polymer films change their electrical resistance or generate charge when compressed or stretched. These can be wrapped around complex, multi-jointed robotic hands to give them continuous, multi-point pressure sensing across their entire surface area, mimicking human tactile nerves.
- Micro-Fluidic Tactile Arrays: Channels filled with conductive liquid embedded inside soft silicone hands detect minute variations in pressure by tracking changes in fluid resistance. These systems are extraordinarily resilient to mechanical wear and tear, solving the historical fragility of electronic skin.
When you combine these tactile materials with modern deep learning, something extraordinary happens. The robot stops treating the physical world as a collection of static 3D meshes and starts perceiving it as a dynamic dynamic field of forces, textures, and physical constraints.
The Data Bottleneck and Closed-Loop Haptic AI
Installing flexible sensors onto a robot's palm is only half the battle. The far bigger challenge—and the one that keeps AI researchers up at night—is processing tactile data at scale. The current revolution in generative AI was made possible because we had the entire public internet to train Large Language Models on billions of pages of text. But there is no "internet of touch." You cannot scrape the web for the physical sensation of holding a wet bar of soap or pulling a stuck key out of a rusted lock.
To overcome this data bottleneck, leading research teams are turning to novel neural network architectures designed specifically for closed-loop sensory-motor control. Rather than feeding tactile signals into a slow, high-level reasoning model, the best systems operate on a dual-rate architecture:
1. High-Frequency Tactile Feedback (Sub-Millisecond Loop)
A low-level neural controller operates directly on the robot's hardware at rates between 500 Hz and 1,000 Hz. This loop doesn't worry about long-term task planning; its sole job is to manage stability, detect microscopic slip events, and adjust motor torque instantaneously based on tactile sensor streams. This acts like the autonomic reflexes in the human spinal cord.
2. High-Level Semantic Planner (Low-Frequency Loop)
A larger multimodal transformer running at 10 Hz to 30 Hz processes vision, natural language commands, and aggregated tactile summaries. It dictates overall goals—such as "peel this banana" or "plug in this USB cable"—while relying on the low-level reflexive loop to execute the fine-grained physical contact safely.
In my opinion, this hybrid approach is the breakthrough that will finally allow robots to transcend controlled factory floors. I spent time observing research on embodied AI simulators, and the progress in bridging the "sim-to-real gap" for touch is astounding. By simulating soft-body physics, micro-friction, and material elasticity in GPU-accelerated environments like NVIDIA Isaac Gym, we can now train a robot end-effector across millions of virtual grasp attempts in a matter of hours before deploying the learned weights onto physical hardware.
Transforming
This analysis was inspired by a story originally reported by IEEE Spectrum. Read the original report →
Supercharge Your Workflow with Claude AI
The AI assistant used by professionals worldwide. Write, code, analyse — all in one place.


