Papers
arxiv:2609.35432

Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

Published on Sep 28
· Submitted by
Hongcheng Gao
on Sep 29
Authors:
,
,
,
,
,
,
,
,
,
,
,

Abstract

Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, progress, and failure recovery are implicitly encoded in action sequences, making them difficult to inspect or revise. Digital coding agents offer a precedent: LLMs call tools, verify results, and revise from feedback as executable code. The same working pattern of explicit state, manageable execution, and revisable procedures underlies generalization and long-horizon execution in the physical world, letting physical experience return as reusable programs, memory, or evidence. We propose Physical Coding, representing task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes planning, verification, recovery, and execution. We build HexaAnything, which calls perception, planning, and control tools, including VLA/WAM policies, and makes in-the-loop decisions from external feedback. Verified traces become data and memory, enabling evolution from tools and Harness to model weights, architectures, and ultimately hardware and task design. On RoboCasa365, HexaAnything improves Composite-Unseen and overall success over XR-1 VLA, and its Harness-trained HexaModel beats the base on every split, indicating code traces internalize physical execution. On PhyBench and a dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results. We observe data, model, and tool self-evolution; future work targets weight internalization, autonomous redesign of architectures, languages, representations, and tasks, and deployment in manufacturing and science.

Community

Paper author Paper submitter

Physical Coding represents task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes actions, verification, and recovery. HexaAnything implements this interface through a Harness connecting models, tools, and external evaluation. On RoboCasa365, it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 60.8% over the native XR-1 VLA. A 27B HexModel trained on Harness-collected data improves over its base model on every split. Evaluations also cover tool revision, simulated scientific experiments, and real-robot execution. These provide initial evidence for Harness, data, model, and tool improvement; broader co-evolution remains future work.

that's really interesting! but your project pages seems kind of slow that some images and videos cannot be seen.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35432
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35432 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.35432 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35432 in a Space README.md to link it from this page.

Collections including this paper 2