From Ray-Ban Meta, PLAUD, and Amazon Bee, to Rabbit R1, robot dogs, and humanoid robots: AI is systematically giving itself “organs,” one by one.
If the dominant AI products of the past two years were ChatGPT, Claude, and Cursor, then the most interesting shift in 2026 is that AI has started “moving out” of the screen.
It has moved into glasses, pins, earbuds, small boxes, robot dogs, and humanoid robots.
Some give AI a pair of eyes, some give it ears, some give it memory, some give it a display, and some are even starting to give it hands and legs.
So I’m increasingly convinced that “AI is looking for a physical shell” is not a metaphor.
But “shell” here doesn’t necessarily mean a humanoid robot. It could be a pair of glasses, a pin, a wristband, a robot dog, or ultimately a humanoid robot truly capable of acting in the real world. They are all trying to solve the same problem: how does AI go from “understanding the world” to “existing in the world”?
I’m an engineer working on world models, spending my days with MuJoCo and DreamerV3. But over the past six months I’ve been paying closer and closer attention to AI hardware — because the world models, simulation training, and sim-to-real work I do ultimately face a core question: how do we get AI from “understanding the world” to “acting upon the world”? And that AI will inevitably need some kind of physical embodiment.
This article is not a technical survey. It’s more like a “walk-through-the-exhibition-hall observation of AI hardware.” I want to approach it from an interesting angle: AI is acquiring physical form organ by organ — from perception to interaction to action, progressively building a complete body. Of course, “organ” is a metaphor here; more precisely, AI is acquiring different physical interfaces.
AI’s Eyes: Ray-Ban Meta

The first product I want to discuss isn’t a robot — it’s a pair of glasses.
Released in September 2025, the Ray-Ban Meta Gen 2 has, by 2026, become the most mature and most consumer-ready product form factor in the AI glasses space. What makes it interesting is precisely that it doesn’t look like an AI product. No large screen, no sci-fi styling. It’s just a normal pair of glasses — but packed with a camera, microphone, speakers, and AI.
Its truly interesting capability isn’t “you can ask ChatGPT things.” It’s this:
AI can see what you’re looking at.
“What plant is this?” “Help me translate this menu.” “What’s this part called?” — AI has gained a very special input: the wearer’s first-person perspective.
I think one of the keys to its success is that it’s first and foremost a normal pair of glasses, and only then an AI device.
AI’s Eyes + Output: Meta Ray-Ban Display

If Ray-Ban Meta gave AI a pair of eyes, then Meta Ray-Ban Display starts adding a true visual output.
This is the first product in the Ray-Ban AI glasses line to include a display, starting at $799. AI’s responses no longer reach you only through earbuds — they appear directly in front of your eyes.
It brings captions, messages, translations, and select AI information into your field of view. Before, it was AI → earbuds → you listen. Now it becomes AI → glasses → you see.
Even cooler is the accompanying Neural Band: a wristband that captures subtle hand movements through EMG signals on the wrist, letting you control the glasses silently — no speech needed, just finger gestures. This is probably the most “futuristic” interaction method of 2026.

AI’s Eyes + Display: RayNeo X3 Pro

Beyond Meta, there’s a Chinese company worth watching. The RayNeo X3 Pro is a true binocular full-color MicroLED AI+AR device: 640x480 MicroLED display, 12MP camera, Snapdragon AR1 chip, and integrated Gemini AI.
It represents a different, more aggressive path: turning AI glasses directly into AR devices with visual output. AI no longer has only visual input — it’s also beginning to have visual output directed at the wearer.
What interests me about products like this isn’t “watching movies on AR glasses.” It’s this: when AI has both visual input and visual output, I think it stops being just a Q&A tool — it starts becoming a layer of “intelligent interface” within your field of view.
AI’s Ears: PLAUD NotePin S

If I had to pick a product that “isn’t as cool, but might be more useful than a lot of robots,” I’d pick PLAUD.
The PLAUD NotePin S is an AI recording device. It doesn’t sound sexy at all: record → transcribe → summarize.
But this is actually a very sensible form factor for AI hardware. Because the problem with phones is: you need to actively take them out. And during meetings, interviews, walks, or drives, you really don’t want to keep holding your phone.
The PLAUD NotePin S turns the recording device directly into something wearable, then hands the work off to AI for transcription, summarization, and information organization.
If Ray-Ban Meta gave AI a pair of eyes, then the NotePin S is more like giving AI a pair of “ears”:
AI doesn’t need hands yet. Just having a pair of ears is enough to start.
AI’s Memory: Amazon Bee

Take one step further, and an even more interesting product form factor emerges: not just recording, but pushing toward “continuous context.”
After acquiring Bee, Amazon continued developing this AI wearable into 2026. Bee can be used as a clip-on device or a wristband. Its core capability is recording, transcribing, and summarizing daily conversations, and then further leveraging context to provide reminders and suggestions.
What’s interesting about Bee isn’t just that it’s “always listening” — it’s that it tries to turn individual conversations into a long-term personal context.
PLAUD is more like an “AI voice recorder.” Bee is closer to an “AI life memory.”
And this is also where AI wearables get most sensitive: when a device starts perceiving continuously, the boundary between “personal AI” and “ambient surveillance device” becomes very blurry. Privacy concerns shift from “what data does it know about me” to “what data does it know about other people around me.”
AI Trying to Get “Hands in the Digital World”: Rabbit R1

Rabbit R1 has continued iterating into 2026, so I’d rather think of it as an ongoing Agent hardware experiment.
What’s truly worth paying attention to isn’t whether it became a “phone killer,” but the question it raised:
If an AI Agent can operate the digital world on your behalf, do users still need traditional apps?
Traditional interaction: human → open app → find feature → click → type. Agent interaction: human → tell AI what I want → AI operates on its own.
The R1 is more of a product form factor experiment than a concluded attempt. Its most valuable contribution may not be whether it works well today, but that it keeps testing: does a dedicated AI Agent entry point actually need to be a standalone device?
From Perception to Action: The Real Watershed
Most of the AI hardware discussed so far hasn’t truly “entered the physical world.” Glasses only see. The Pendant only listens. The Display only shows. The R1 primarily operates in the digital world.
The real watershed is when AI’s actions start changing the physical world. That’s where robots enter the picture.
AI’s Legs: Unitree Go2

If the products above were adding eyes, ears, and memory to AI, then robot dogs start adding: legs.
The Unitree Go2 officially positions itself as an “Embodied AI” platform, equipped with cameras, IMU, 4D LiDAR, and other perception capabilities. It packs perception, localization, motion control, and compute into a robot body that actually moves.
At this point, AI shifts from “answering questions” to: I take an action, and then the world changes.
This is precisely what I find most interesting after working on world models. Because the real question reinforcement learning faces isn’t “what is this thing?” but rather: “If I do this right now, what will happen in the next moment?”
AI’s Hands: Dexterous Manipulation

After legs, the natural next step is giving AI a pair of hands.
In 2026, dexterous manipulation is one of the most active directions in robotics. CMU’s LEAP Hand, the Shadow Dexterous Hand, and dexterous hands from Unitree and Zhiyuan all share the same goal: enabling AI not just to move, but to manipulate.
A robotic arm can only handle certain objects, but a dexterous hand tries to let AI grasp, rotate, and assemble things the way humans do. This is precisely the most critical step between “mobile robot” and “truly useful robot.”
And this also explains why humanoid robots have suddenly become important — because they may be the best form factor for having both legs and hands simultaneously.
AI’s Complete Body: Humanoid Robots

Take one more step forward, and you reach humanoid robots. Figure, Unitree, Zhiyuan, and others are all building them.
What’s truly interesting about humanoid robots isn’t that they “look like humans,” but that they’re trying to become a general-purpose body. A robotic arm can only handle certain objects. A robot dog is good at locomotion. But a humanoid robot can, in theory, directly enter factories, warehouses, homes, and tool environments designed for humans. Its greatest potential advantage isn’t necessarily athletic capability — it’s that the human world has already built massive infrastructure around the human body form: doors, stairs, tools, workbenches, shelves, cars, living spaces.
NVIDIA’s GR00T, Isaac, Cosmos, and other infrastructure efforts are trying to stitch together models, simulation, data, and edge computing into a full Physical AI technology stack. Ultimately, the model’s predictions and policies need to become real actions through motors and mechanical structures. The entire system enters a closed loop of: perceive → think → act → receive feedback → learn again.
Failed Experiments Worth Studying
An AI hardware article would be boring if it only covered successes. There are two cases worth discussing separately.
The Humane AI Pin. In February 2025, Humane announced it was shutting down consumer AI Pin services; operations officially ceased on February 28. Some of its AI capabilities, talent, and intellectual property were subsequently acquired by HP for approximately $116 million. It tried to pull AI out of the phone: stop staring at screens. It sounded wonderful. But the problem was: the phone is already an extremely mature computing platform. Asking a small device with only a camera, microphone, projector, and voice interaction to replace a phone meant going head-to-head with decades of ecosystem accumulation. The lesson it left behind: just because AI can do something doesn’t mean users need to buy a new device for it.

The Limitless Pendant is a case very worth studying. It explored a genuinely interesting direction: AI no longer waits for you to actively open it — it persistently exists within your life context. In December 2025, Meta acquired Limitless, which subsequently stopped selling new hardware. A similar “ambient AI / continuous context” approach also appears in products like Amazon Bee.

The Full Picture at a Glance
If you lay all the products discussed in this article side by side, an interesting progression emerges:
| Capability AI Gains | Product | What It Really Adds |
|---|---|---|
| Visual Perception | Ray-Ban Meta Gen 2 | First-person camera |
| Auditory Perception | PLAUD NotePin S | Continuous audio input |
| Long-term Context | Amazon Bee | Personal life data |
| Visual Output | Meta Ray-Ban Display / RayNeo X3 Pro | AR display |
| Novel Input | Meta Neural Band | EMG control |
| Digital Action | Rabbit R1 | Agent entry point |
| Physical Movement | Unitree Go2 | Legs + environmental interaction |
| Physical Manipulation | Dexterous Hands / Humanoid Robots | Hands + legs + body |
“Looking for a physical shell” isn’t a sudden jump from phones to humanoid robots — it’s a gradual process of increasing perceptual and action capabilities.
Closing Thoughts
Perhaps “AI hardware” won’t end up being a single product category.
Glasses, earbuds, pins, robots, even cars — they may all just be different interfaces connecting AI to the real world. Some AI only needs a pair of eyes, because its task is to understand the world you see. Some AI needs a pair of ears, because its task is to remember what you’ve said. Some AI needs a pair of legs and a pair of hands, because its task has finally become changing the world.
So what’s truly interesting about “AI is looking for a physical shell” isn’t what AI will ultimately look like — it’s that we’re witnessing for the first time: what kind of body does an AI system actually need to truly enter the real world?
And the world models, simulation training, and sim-to-real work I do sits right at the other end of this path: once AI truly has eyes, ears, hands, and legs, it still needs to know “what will happen next.”
That may be what’s truly interesting about world models. Once AI has a body, the question is no longer just “can you answer me” — it becomes: “If I do this right now, what will happen to the world?”
The physical world is vast. AI has only just reached the doorstep.
Comments