Robot-Use Agents (RUA)

Research Notes › RSI · CLI · RUA

Notes on agents that operate a robot the way they operate a computer. A robot is a software system — drivers, middleware, planners, controllers — so a coding agent that can drive a codebase and a terminal can, in principle, drive an arm. One paper per entry.

Start here → Robot-Use Agents (Phillip Isola, Sep 2026), which names the paradigm, and CaP-X (Fu et al., Mar 2026), the first benchmark that measures it. The premise: a frontier coding agent already knows geometry, physics and programming, and a robot already exposes a programmable interface — so the policy can be written rather than learned, and then inspected, edited and debugged instead of retrained.
The full reading list is at the bottom of this page.

Why this is the same question as the CLI series

The CLI series asks where an agent's action contract comes from when it operates an application. Robotics asks the identical question with harder consequences: the contract is a motion interface, the feedback is physical, and a bad action breaks hardware rather than a database row. And the answers line up almost one to one — hand-lifted primitive libraries are L1, frozen skills replayed from demonstrations are L2, and the robot's own middleware (ROS 2) is the L0 case where the contract already ships with the platform. The interesting claim, in both fields, is that the pre-existing native interface beats the bespoke one, because nobody has to maintain it and it is identical across embodiments and between simulation and hardware.

The map

Four families, ordered by how much of the control loop the agent itself owns — plus an evaluation axis that runs across all of them.

FamilyWhat the agent acts throughWho closes the loopPapers
A · Agent over a learned policySemantic tool calls: pick(obj), place(obj, tray)The VLA's own networkSayCan; Harness VLA; Zetta; RoboHarness; Emerge-Policy; VOLO
B · Agent writes code over primitivesA curated library: grasp sampler, IK solver, joint/gripper movesThe agent, between primitive callsCode as Policies; CaP-X; RHO; ASPIRE; GTA-2; Playful Agentic Robot Learning
C · Agent emits actions directlyEnd-effector poses, step by step, often from a visual interfaceThe agent, every stepVIA; GUAVA; Show-Harness; Agent-as-Policy; GPT-as-Policy
D · Agent on the robot's native stackROS 2 topics, services, actions — ros2cli and rclpyThe agent, in code it writes itselfROSA; ROS-LLM; rosaOS; ROSClaw
Eval axisHow do you score a policy nobody trained?—LIBERO-Plus → LIBERO-PRO → CaP-Bench → RoboCasa365

A and B pay for their contract up front: someone collects demonstrations, or someone writes and maintains the primitive set. C pays per step, in agent turns. D is the only family where the contract is free, because the robot already shipped it — but it is also the family with the least evidence, since the existing ROS-agent frameworks predate agentic coding and are evaluated on in-house demos rather than benchmarks.

Five claims the field currently rests on

  1. Learned visuomotor policies are bounded by their demonstrations. OpenVLA and π0 score 0% on LIBERO-PRO — the same tasks they were trained on, with the scenes perturbed. The headline numbers on LIBERO were memorisation.
  2. The tool interface hides the failure. When a learned policy is wrapped as pick(obj), the orchestrating agent cannot see why the grasp missed, so it cannot diagnose or recover — it can only retry. The agent's world knowledge never reaches the layer where the error happened.
  3. A fixed primitive set caps what any amount of experience can buy. On contact-rich tasks (nut insertion), agents writing programs over the same curated primitives stay near zero however much task-specific experience they accumulate: nothing in the set senses contact, so no program built from it can correct a pose mid-motion.
  4. Most reported gains are not zero-shot. Rehearsing each evaluation task beforehand, carrying a task memory, or reading the benchmark's success predicate at run time are all common — and rarely tabulated. Any comparison in this area needs explicit train / experience / oracle columns.
  5. Privileged state and bespoke APIs do not transfer. Systems that read ground-truth object poses from the simulator, or call a hand-written sample_grasp_pose(obj), have to be re-engineered for every new embodiment — and the first category cannot run on hardware at all.

The open question this leaves: if the agent is handed the robot's native interface and nothing else, is a frontier coding agent already a zero-shot visuomotor policy? That is the question I am writing a paper about; the numbers stay off this page until it is out of review.

Reading list

In writing order. The tag is the family from the table above.

Adjacent work — for related-work only. π0.7 and GR00T N1 (the generalist policies the agent systems sit on top of) · World Action Models are Zero-Shot Policies (the other route to zero-shot: predict observations and actions jointly) · Act-Observe-Rewrite (multimodal coding agents as in-context policy learners) · Demonstration-Free Robotic Control via LLM Agents · VLCP (closed-loop code replanning) · Revisiting Push-T with Agentic Robotics · minimal-interface zero-shot agents for VLN (the navigation counterpart of the same claim) · Claude Plays Robotics and GPT-6 Astra on robotic manipulation (vendor reports; no protocol, but the first public evidence).