Robot-Use Agents (RUA)
Research Notes › RSI · CLI · RUA
Notes on agents that operate a robot the way they operate a computer. A robot is a software system — drivers, middleware, planners, controllers — so a coding agent that can drive a codebase and a terminal can, in principle, drive an arm. One paper per entry.
Start here → Robot-Use Agents (Phillip Isola, Sep 2026), which names the paradigm, and CaP-X (Fu et al., Mar 2026), the first benchmark that measures it. The premise: a frontier coding agent already knows geometry, physics and programming, and a robot already exposes a programmable interface — so the policy can be written rather than learned, and then inspected, edited and debugged instead of retrained.
The full reading list is at the bottom of this page.
Why this is the same question as the CLI series
The CLI series asks where an agent's action contract comes from when it operates an application. Robotics asks the identical question with harder consequences: the contract is a motion interface, the feedback is physical, and a bad action breaks hardware rather than a database row. And the answers line up almost one to one — hand-lifted primitive libraries are L1, frozen skills replayed from demonstrations are L2, and the robot's own middleware (ROS 2) is the L0 case where the contract already ships with the platform. The interesting claim, in both fields, is that the pre-existing native interface beats the bespoke one, because nobody has to maintain it and it is identical across embodiments and between simulation and hardware.
The map
Four families, ordered by how much of the control loop the agent itself owns — plus an evaluation axis that runs across all of them.
| Family | What the agent acts through | Who closes the loop | Papers |
|---|---|---|---|
| A · Agent over a learned policy | Semantic tool calls: pick(obj), place(obj, tray) | The VLA's own network | SayCan; Harness VLA; Zetta; RoboHarness; Emerge-Policy; VOLO |
| B · Agent writes code over primitives | A curated library: grasp sampler, IK solver, joint/gripper moves | The agent, between primitive calls | Code as Policies; CaP-X; RHO; ASPIRE; GTA-2; Playful Agentic Robot Learning |
| C · Agent emits actions directly | End-effector poses, step by step, often from a visual interface | The agent, every step | VIA; GUAVA; Show-Harness; Agent-as-Policy; GPT-as-Policy |
| D · Agent on the robot's native stack | ROS 2 topics, services, actions — ros2cli and rclpy | The agent, in code it writes itself | ROSA; ROS-LLM; rosaOS; ROSClaw |
| Eval axis | How do you score a policy nobody trained? | — | LIBERO-Plus → LIBERO-PRO → CaP-Bench → RoboCasa365 |
A and B pay for their contract up front: someone collects demonstrations, or someone writes and maintains the primitive set. C pays per step, in agent turns. D is the only family where the contract is free, because the robot already shipped it — but it is also the family with the least evidence, since the existing ROS-agent frameworks predate agentic coding and are evaluated on in-house demos rather than benchmarks.
Five claims the field currently rests on
- Learned visuomotor policies are bounded by their demonstrations. OpenVLA and π0 score 0% on LIBERO-PRO — the same tasks they were trained on, with the scenes perturbed. The headline numbers on LIBERO were memorisation.
- The tool interface hides the failure. When a learned policy is wrapped as
pick(obj), the orchestrating agent cannot see why the grasp missed, so it cannot diagnose or recover — it can only retry. The agent's world knowledge never reaches the layer where the error happened. - A fixed primitive set caps what any amount of experience can buy. On contact-rich tasks (nut insertion), agents writing programs over the same curated primitives stay near zero however much task-specific experience they accumulate: nothing in the set senses contact, so no program built from it can correct a pose mid-motion.
- Most reported gains are not zero-shot. Rehearsing each evaluation task beforehand, carrying a task memory, or reading the benchmark's success predicate at run time are all common — and rarely tabulated. Any comparison in this area needs explicit train / experience / oracle columns.
- Privileged state and bespoke APIs do not transfer. Systems that read ground-truth object poses from the simulator, or call a hand-written
sample_grasp_pose(obj), have to be re-engineered for every new embodiment — and the first category cannot run on hardware at all.
The open question this leaves: if the agent is handed the robot's native interface and nothing else, is a frontier coding agent already a zero-shot visuomotor policy? That is the question I am writing a paper about; the numbers stay off this page until it is out of review.
Reading list
In writing order. The tag is the family from the table above.
- framing Robot-Use Agents — Phillip Isola (MIT), Sep 2026 · essay · names the paradigm: an agent operating a robot the way it operates a computer. The series intro goes here — why robot control is a software problem, and what carries over from computer use and what does not (contact, irreversibility, real time).
- B Code as Policies — Liang, Huang, Xia, Xu, Hausman, Ichter, Florence, Zeng (Google), ICRA 2023 · paper · the origin. An LLM writes executable policy code over perception and control primitives; the policy becomes an artifact you can read. Everything in family B is a descendant, and every limitation of family B is already visible here: the primitives are the ceiling.
- B CaP-X: Benchmarking and Improving Coding Agents for Robot Manipulation — Fu, Yu, El-Refai, Kou, Xue, Huang, Xiao, Wang, Li, Shi, Wu, Sastry, Zhu, Goldberg, Fan (Berkeley + NVIDIA + Stanford), Mar 2026 · arxiv · introduces CaP-Bench (lift, stack, wipe, nut, restack, bimanual lift, handover) with human-expert programs as a reference point (82.6%). The single most useful number in the series: the human ceiling, per task.
- B RHO: Your Coding Agent Is Secretly a Roboticist — Elmaaroufi, Svegliato, Kalade, Schelle, Seshia, Zaharia (Berkeley), Jun 2026 · arxiv · a coding agent driving a manipulation stack end to end. Read next to ASPIRE: both accumulate experience on the evaluation tasks, which is where most of the gain over plain CaP comes from.
- B ASPIRE: Agentic Skill Discovery for Robotics — Lu, Wu, Kou, Fu, Xiao, Mandlekar, Xu, Shi, Goldberg, Chen, Chowdhury, Zhu, Fan, Wang, Jul 2026 · arxiv · the L2 analogue: discover and freeze a skill library, then call it. Same compile-once-reuse-forever bargain as SkillDroid in the CLI series, with the same failure mode — the library cannot cover what it never saw.
- A SayCan — Ichter, Brohan, Chebotar et al. (Google), CoRL 2022 · paper · the hierarchical template every family-A system still follows: LLM proposes, affordance-scored skills dispose. Worth re-reading to see exactly which of its assumptions the coding-agent era removes.
- A Harness VLA: Steering Frozen VLAs via Memory-Guided Agents — Zhang, Zhang, Gao, Li, Liu et al., Jul 2026 · arxiv · the strongest agent-over-policy result, and the cleanest illustration of claim 4: it rehearses each evaluation task once and reads the success predicate during execution. On held-out composite tasks it falls back to the underlying policy's ceiling, which is the whole argument against the architecture.
- A Zetta ζ — Ding, Mi, Huang, Wang, Zhang et al. (Microsoft Research Asia), Aug 2026 · arxiv · closed-loop embodied harness with self-evolution; pair with RoboHarness (memory-driven orchestration of heterogeneous policies) and Skills in Weights, Memory in Code — three takes on where to put state when the motor skill is frozen.
- C VIA: Visual Interface Agent for Robot Control — Hu, Sundaresan, Gao, Sadigh (Stanford), Jul 2026 · arxiv · the agent acts through a visual interface, issuing poses. The direct test of whether a multimodal model can ground a pixel to a centimetre — and the evidence that it cannot, which is what pushes family D toward self-written geometry.
- C GUAVA — Liu, Li, Yao, Shi, Zhou, Huang, Huang, Mao (Maryland + MIT), Jun 2026 · arxiv · "universal harness for embodied manipulation"; read with Show-Harness ("just a VLM agent can play robots") and Agent as Policy. The three of them define what the minimal per-step action interface looks like.
- D ROSA / ROS-LLM / rosaOS / ROSClaw — Royce et al. (NASA JPL), Oct 2024 · arxiv · Mower et al. (Huawei Noah's Ark), Jun 2024 · arxiv · Ge et al., ACL 2026 demo · Cardenas et al., Mar 2026 · arxiv · the pre-agentic ROS-LLM lineage: natural language onto ROS 2 topics and services, no code generation, evaluated on in-house demos. They establish the plumbing and leave the question open — the gap this series is about.
- eval LIBERO-PRO — Zhou, Xu, Tie, Chen, Zhang, Chu, Zhou, Sun, Oct 2025 · arxiv · perturbs objects, spatial layouts, goals and instructions on LIBERO and watches OpenVLA and π0 go to zero. Companion: LIBERO-Plus (Fei et al.) and LIBERO-PARA (paraphrase robustness). Together, the case that VLA leaderboard numbers measure memorisation.
- eval Zero-shot has no distribution to fall out of — my own angle, written last. If a policy was never trained, "seen" versus "unseen" splits stop being a property of the policy and become a property of the baselines. What should replace them: task length, contact richness, and cost per success — and a mandatory disclosure of training, prior experience, and oracle access.
Adjacent work — for related-work only. π0.7 and GR00T N1 (the generalist policies the agent systems sit on top of) · World Action Models are Zero-Shot Policies (the other route to zero-shot: predict observations and actions jointly) · Act-Observe-Rewrite (multimodal coding agents as in-context policy learners) · Demonstration-Free Robotic Control via LLM Agents · VLCP (closed-loop code replanning) · Revisiting Push-T with Agentic Robotics · minimal-interface zero-shot agents for VLN (the navigation counterpart of the same claim) · Claude Plays Robotics and GPT-6 Astra on robotic manipulation (vendor reports; no protocol, but the first public evidence).



