By Jim Shimabukuro (assisted by Claude)
Editor
[Related: “The Increasingly Alien World of Embodied AI Agents” and “A Field Guide to Generative AI, Agentic AI, AGI, ASI, and The Singularity]
Summary: The next office suite may not live in a window at all—it may walk, listen, and act in the physical world, turning software from a tool you use into a worker that does the work for you. –Perplexity
The strangest thing about the past four years is how little the furniture moved. Generative AI arrived and knocked down a wall of assumptions we had leaned on since the first computer science department opened its doors — that machines could not write, could not draw, could not hold a conversation that felt like one. All of that fell. And yet the room you are sitting in looks exactly as it did in 2021. The desk is the same desk. The coffee cup still has to be carried to the sink by a human arm. The revolution happened entirely behind glass. But that is beginning to change, and the first evidence is hiding somewhere thoroughly unglamorous: in the toolbar of a word processor.
The cursor learns to click
In April 2026, Microsoft announced that Copilot’s “agentic capabilities” in Word, Excel, and PowerPoint had become generally available — the default experience, not an experiment. Sumit Chauhan, who runs the Office product group, described the shift in terms that are worth reading slowly. When Copilot first shipped, she wrote, the underlying models “were not powerful enough to use Copilot to command the applications. This meant Copilot was a passive partner in documents: it could answer questions but missed the mark when it was asked to take action on the canvas directly.” Now, she said, it can “take on high-level tasks and execute them like a human would.” Her first lesson learned, listed above all the others: “Taking action matters.” [1]
Read that again with an eye on the verbs. Command. Take action. Execute. Something that used to sit beside you and offer opinions has been handed the mouse.
The numbers Microsoft published alongside the announcement are the sort of thing that gets skimmed past in a press release and looks quite different in hindsight. Weekly attempts per user rose 52 percent in Word, 67 percent in Excel, 11 percent in PowerPoint. [1] People, it turns out, ask for a great deal more once the answer arrives as finished work rather than as advice.
A month later, in May 2026, Microsoft pushed the idea one step further and made “computer-using agents” generally available in Copilot Studio, rolling them out across all its commercial regions. The company’s own framing is the clearest description of the moment anyone has written: a computer-using agent “gets the same tools a human employee has: a browser, a screen, a keyboard, and the ability to read what’s on the page and take the next step.” It navigates live interfaces by looking at them, the way you do, rather than through the brittle plumbing of an API. [2]
This is the hinge. Not intelligence, exactly — competence. For decades, software could only reach the things that had been wired up for it in advance. An agent that can look at a screen and act is not limited to the connections an engineer anticipated. It can use any application built for a person, including applications nobody thought to integrate, including applications that will be written next year.
And once a system can operate anything designed for human hands and human eyes, the interesting question stops being what it can do on a screen. The interesting question becomes what else was designed for human hands and human eyes.
The answer, as it happens, is nearly everything.
The big bang, as advertised
Silicon Valley has a habit of naming its own weather. At NVIDIA’s GTC conference in March 2026, the company declared that “the big bang of physical AI has started” — a phrase that arrived, as these phrases do, in a wash of green stage lighting. But the partner list behind the slogan was less theatrical and more telling: ABB, Agility, FANUC, Figure, KUKA, Skild AI, Universal Robots, Yaskawa. Old-line industrial arm manufacturers standing next to humanoid startups, all of them building on the same simulation and model stack. [3]
Jensen Huang, NVIDIA’s chief executive, put it in the register he favors: “The dawn of a new industrial revolution has arrived, where physical AI and autonomous AI agents are fundamentally reinventing how the world designs, engineers and manufactures.” [3] Discount that by whatever factor you reserve for men selling chips. What survives the discount is a genuine technical bottleneck and a genuine attempt to break it.
The bottleneck is data. Large language models worked partly by accident of history: humans have been generating text for several thousand years, and we obligingly put a great deal of it online. There is no equivalent archive for reaching for a mug. Nobody logged the ten thousand small corrections your wrist made this morning. NVIDIA’s answer at GTC was an architecture for manufacturing that missing experience — gathering real robot data, generating synthetic data, curating and evaluating both. Rev Lebaredian, who runs the company’s Omniverse group, summed up the logic in four words that would look at home on a monument: “In this new era, compute is data.” [3]
That is the industrial version. The research version is more interesting, because it has already produced results that are hard to explain away.
Think before acting
In October 2025, Google DeepMind’s robotics team published a report on Gemini Robotics 1.5, a model family built to close the gap between a system that understands a scene and a system that can do something about it. The technical claim that matters to a non-specialist is this: the model “interleaves actions with a multi-level internal reasoning process in natural language,” which lets a robot “think before acting” and improves its ability to break down long, multi-step tasks. The authors describe the result, without much hedging, as “an era of physical agents — enabling robots to perceive, think and then act so they can solve complex multi-step tasks.” [4]
There is a second innovation, a technique the team calls Motion Transfer, that allows a single model to learn from data collected on physically different robots. [4] It is a quiet detail with loud consequences. It means the awkward hours logged by a one-armed research arm in a lab in London can teach a two-armed machine in a warehouse in Ohio. Experience becomes portable in a way that human experience never has been.
The most vivid demonstration of what that portability buys came a few months earlier, from Physical Intelligence, a San Francisco startup. In April 2025 the company released π0.5 (“pi zero point five”), a model that controls a mobile robot cleaning kitchens and bedrooms in homes it has never seen. The company is careful about the claim — the goal, they wrote, “is not to accomplish new skills or exhibit high dexterity, but to generalize to new settings.” The robot fails often. But the description of how it fails is what stays with you: “It does not always succeed on the first try, but it often exhibits a hint of the flexibility and resourcefulness with which a person might approach a new challenge.” [5]
Their scaling result is the part to remember. As the researchers added more distinct homes to the training data, performance in unseen homes climbed steadily; after roughly a hundred training environments, the model performed nearly as well in a strange house as a model trained directly on that house. [5] Generalization, in other words, is not magic. It is a curve, and someone has found the dial.
Cut and paste, in sheet metal
Which brings us back to the office suite, and to a plant in Spartanburg, South Carolina.
In June 2026, BMW announced that it was deploying Figure AI’s Figure 03 humanoid at Spartanburg, following an eleven-month run with the previous model in the body shop. The earlier robot had inserted sheet-metal parts for welding during the production of more than 30,000 BMW X3 vehicles. “Our 11-month deployment of Figure 02 proved that humanoids are no longer lab experiments,” said Figure’s founder, Brett Adcock. “They can be a valuable asset in establishing a flexible, reliable manufacturing workforce.” [6]
The new job is worth describing precisely, because it is where the metaphor stops being a metaphor. Components arrive at the plant in large containers, unsorted. The robot picks them out and arranges them into a sequencing trolley, in the order the assembly workers will need them. The trolley is then wheeled to the line and delivered “just in sequence.” [6]
That is a sort operation. It is Data > Sort performed on objects with mass. A spreadsheet function that has been running invisibly inside a computer since VisiCalc has climbed out and acquired hands, tactile sensors in its fingertips, and palm cameras. [6]
Once you notice the pattern you start seeing it everywhere. Agility Robotics’ Digit, a bipedal machine that moves totes between conveyors and mobile robots, had logged more than 65,000 cumulative operating hours across nine customer facilities by mid-2026, including a GXO Logistics warehouse in Georgia where it has moved over 100,000 totes. [7] That is a batch job. That is Find and Replace, executed on a pallet.
Amazon’s Blue Jay system runs several robotic arms in coordinated concert at a single station to pick, stow, and consolidate — a multi-track workflow, of the kind you would once have built in a project planner. Alongside it, the company introduced Project Eluna, an agentic AI that reads historical and live data across an entire building to anticipate bottlenecks and recommend actions to human operators. [8] That is the assistant in the sidebar, except its document is a warehouse the size of a small town.
None of this is arriving into an empty world. The International Federation of Robotics counted 4,664,000 industrial robots in operation worldwide in 2024, up nine percent in a year, with 542,000 new units installed. [9] The machines were already here. What is new is that they are being handed something that resembles judgment.
The window, relocated
The other half of this story is not about robots at all. It is about where the interface goes.
In October 2025, Amazon unveiled prototype smart glasses for its delivery drivers, a system it calls Amelia. When a driver reaches a stop, the glasses wake, help locate the right package in the van, give walking directions to the door, and capture proof of delivery — all without the driver looking down at a handheld scanner. Amazon estimates the setup could save drivers up to thirty minutes a shift. [10]
Strip away the logistics and look at the shape of the thing. The document window has left the desk and attached itself to a person’s field of view. The application is no longer something you visit; it is a layer over the street. The driver walks up a path in Ohio while a set of instructions renders itself onto the hedges.
Meanwhile the most widely used embodied agent in America is not a humanoid at all; it is a car with nobody in the driver’s seat. Waymo set itself a public target of one million paid rides a week in the United States by the end of 2026 [11], and in July 2026 it opened fully driverless service in Las Vegas with Denver, San Diego, and Tampa named as next. [12] Several hundred thousand people a week are now routinely handing their physical bodies to a piece of software, and have largely stopped finding it remarkable. This is how these transitions actually go. Not with a gasp, but with a shrug and a phone check in the back seat.
Home directory
The place where all of this becomes genuinely strange is, predictably, the house.
In October 2025, the Norwegian company 1X opened consumer pre-orders for NEO, a humanoid built specifically for domestic use — $20,000 outright or $499 a month, with first deliveries promised in 2026, primarily in the United States. It weighs 66 pounds, wears a soft polymer body designed to be leaned on without injury, and can lift considerably more than it weighs. Bernt Børnich, the founder, framed the launch in a line that will either look prescient or hilarious in a decade: “Humanoids were long a thing of sci-fi… then they were a thing of research, but today — with the launch of NEO — humanoid robots become a product. Something that you and I can reach out and touch.” [13]
NEO has a feature list that reads like an office suite that has wandered into a kitchen. There is “Chores,” where you hand it a task list and a schedule. There is memory, so it carries context between conversations — grocery lists, birthdays, where you left off in your Spanish lessons. There is visual awareness, so it can look at what is on the counter and suggest what you might cook. [13] Task list, calendar, notes, a recommendation engine. All of the little productivity boxes, unboxed.
And then there is the detail that tips the whole picture into the surreal. For any chore NEO does not yet know how to do, the owner can schedule a human teleoperator at 1X to take remote control and guide the machine through it — completing the job and, in the same motion, generating the training data that will let the robot do it alone next time. [13]
Sit with that. A machine is folding your laundry in your bedroom. Behind its eyes, some of the time, is a person in another country. Every fumbled sock is a lesson, and the lesson is the product. The company has promised privacy controls — no-go zones, the ability to blur faces [13] — which is both reassuring and a reminder of what would need reassuring.
This is the texture of the world that is arriving: not obviously artificial, not entirely real. A house that is also a data collection site. A helper that is sometimes autonomous, sometimes a stranger, and gives no outward sign of which. Familiar rooms, subtly re-lit.
The industrial version has the same doubled quality. Before a single physical part reaches the line at Spartanburg, BMW runs the process through virtual 3D simulations, and uses what it calls the Virtual Factory to model human movement and refine ergonomics. [6] The plant exists twice — once in concrete and once in a rendering — and the two versions correct each other continuously. Which one is the real factory is not a question the software finds meaningful.
The part that has not happened yet
It would be dishonest to end on the swell of the music, because the most interesting skeptic in this argument is also one of the most credentialed people in the field.
In September 2025, Rodney Brooks — an MIT roboticist, a founder of iRobot, the man who put a Roomba in half the living rooms in America — published an essay arguing that today’s humanoids will not learn dexterity, and that the money being poured into the attempt is being poured into a hole. The plan, as he describes it, is for humanoid robots to be “plug compatible” with humans, stepping into existing jobs without anyone having to redesign the work. His verdict: “In my opinion, believing that this will happen any time within decades is pure fantasy thinking.” [14]
His reason is not sentimental, and it is not about consciousness. It is about touch. Human hands are dense with sensors, and dexterity depends on force and contact feedback that vision alone cannot supply. We have never built the equivalent for machines, and — his sharpest point — we have never even built a way to record touch the way we record images and sound. “To think we can teach dexterity to a machine without understanding what components make up touch, without being able to measure touch sensations, and without being able to store and replay touch is probably dumb,” he wrote. “And an expensive mistake.” [14]
He may be right. Notice, though, that Figure 03 shipped with tactile sensors in improved hands and cameras in its palms [6], and that Physical Intelligence’s robots are demonstrably better in a stranger’s kitchen than anything from three years ago [5]. The disagreement is not really about whether this happens. It is about the exponent, and about how many billions get burned finding out.
The other reservation is one nobody has a technical fix for. Software has an undo button. The physical world does not. A word processor that misreads your intent produces a bad paragraph; an agent that misreads your intent while holding a knife produces something else. Every company in this story leads with safety — BMW points to soft components and speech interaction [6], 1X to compliant tendon-driven joints and a soft body [13], Microsoft to review-and-approve controls in the document [1] — and every one of them is describing a problem that is genuinely unsolved rather than a problem they have solved.
The toolbar, walking
Still, stand back far enough and the arc is legible.
For forty years the arrangement was fixed. The machine held the memory and did the arithmetic; you supplied the hands and the going-and-getting. Every application ever written accepted that division. Word could format your letter but could not fetch the paper. Excel could model the warehouse but could not lift anything in it.
What is dissolving now is the division itself. The agent that learned to drive your screen is learning to drive a chassis, and it is the same agent — the same reasoning, the same instruction-following, the same habit of thinking in language before it moves. Gemini Robotics 1.5 talks itself through a task before it acts. [4] π0.5 tells itself “pick up the pillow” and then picks up the pillow. [5] The narration you have watched scroll past in a chat window is the narration now running inside a machine that is walking across your bedroom.
Which means the useful way to picture the next decade is probably not “the robots are coming.” It is closer to this: the menu bar is unfolding into the room. Sort, filter, find, replace, copy, schedule, remind, revise — all of it stepping out of the rectangle and taking up residence among the furniture, wearing bodies, carrying things, occasionally getting it wrong in ways a screen never could.
The furniture is finally about to move. Some of it will move by itself.
References
1. Chauhan, Sumit. “Copilot’s agentic capabilities in Word, Excel, and PowerPoint are generally available.” Microsoft 365 Blog, April 22, 2026. https://www.microsoft.com/en-us/microsoft-365/blog/2026/04/22/copilots-agentic-capabilities-in-word-excel-and-powerpoint-are-generally-available/
2. Lazrek, Mustapha. “Computer-using agents in Microsoft Copilot Studio are now generally available.” Microsoft Community Hub, May 13, 2026. https://techcommunity.microsoft.com/blog/copilot-studio-blog/computer-using-agents-in-microsoft-copilot-studio-are-now-generally-available/4519427
3. Heater, Brian. “NVIDIA Declares ‘Big Bang of Physical AI’ at GTC 2026.” Association for Advancing Automation (A3), March 16, 2026. https://www.automate.org/ai/industry-insights/nvidia-declares-big-bang-of-physical-ai-at-gtc-2026
4. Gemini Robotics Team. “Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer.” arXiv:2510.03342, October 2, 2025. https://arxiv.org/abs/2510.03342
5. Physical Intelligence. “π0.5: a VLA with Open-World Generalization.” April 22, 2025. https://www.pi.website/blog/pi05
6. BMW Group. “BMW Group advances the use of Physical AI in production with Figure 03 project in Spartanburg.” BMW Group PressClub, June 25, 2026. https://www.press.bmwgroup.com/global/article/detail/T0458778EN/bmw-group-advances-the-use-of-physical-ai-in-production-with-figure-03-project-in-spartanburg?language=en
7. Ramachandran, Gowtham. “Agility Robotics goes public at $2.5B valuation.” The Conveyor, June 25, 2026. https://www.theconveyor.co/p/agility-robotics-ipo-digit-humanoid
8. Amazon. “Introducing Blue Jay and Project Eluna, Amazon’s latest robotics and AI technology for its operations.” About Amazon, October 22, 2025. https://www.aboutamazon.com/news/operations/new-robots-amazon-fulfillment-agentic-ai
9. International Federation of Robotics. “World Robotics 2025 report — Industrial Robots.” September 2025. https://ifr.org/ifr-press-releases/news/global-robot-demand-in-factories-doubles-over-10-years
10. Amazon. “Amazon is developing smart glasses allowing delivery drivers to work hands-free.” About Amazon, October 22, 2025. https://www.aboutamazon.com/news/transportation/smart-glasses-amazon-delivery-drivers
11. Ohnsman, Alan. “Waymo Targets 1 Million Robotaxi Rides A Week.” Forbes, December 10, 2025. https://www.forbes.com/sites/alanohnsman/2025/12/10/waymo-targets-1-million-robotaxi-rides-a-week/
12. “Waymo goes driverless in Las Vegas, with Denver, San Diego, Tampa next.” Electrek, July 8, 2026. https://electrek.co/2026/07/08/waymo-driverless-las-vegas-four-new-cities/
13. Oitzman, Mike. “NEO humanoid designed for household use, available for preorder.” The Robot Report, October 30, 2025. https://www.therobotreport.com/1x-announces-pre-order-launch-neo-humanoid-robot/
14. Brooks, Rodney. “Why Today’s Humanoids Won’t Learn Dexterity.” September 26, 2025. https://rodneybrooks.com/why-todays-humanoids-wont-learn-dexterity/
###
Filed under: Uncategorized |






























































































































































































































































































































































































































































































































Leave a comment