EyeRobot 2.0:
Active Gaze for Precise Manipulation
without Wrist Cameras

Kush Hari∗1, Justin Kerr∗1,2, Nidhya Shivakumar1, Samarth Mahapatra1, Carmelo Sferrazza2, Jiahui Lei1, Jitendra Malik1,2, C. Karen Liu2,3, Ken Goldberg†1, Angjoo Kanazawa†1,2

1UC Berkeley 2Amazon FAR 3Stanford University ∗Equal contribution, random order †Equal advising

Decoupling physical attention from wrist cameras with active gaze

How should robots see the world?

Precise manipulation requires high acuity vision. But how do we achieve this? With only ego-centric views, this is non-trivial (a single HD image requires 7500 tokens to represent!). If we use more tractable image sizes, it's difficult to see what's going on in the scene.

Too many tokens!

Visual tokens 16 × 16 pixel patches

One token at low res!

The most popular approach is wrist cameras, which give high-res inputs right where manipulation is happening.

Left wrist view of the wrench and gripper, with a visual-token grid
Right wrist view of the wrench and gripper, with a visual-token grid

But wrist cameras have their own issues

  1. Manipulation necessarily occludes wrist cameras, especially when using tools, manipulating large objects, or doing whole-body manipulation.

  2. It prohibits decoupling vision and action. Many times, you want to look somewhere different from your grippers to achieve a task, wrist cams don't allow this.

  3. It restricts gripper designs. Introducing wrist cameras makes manipulators bulkier and prone to motion blur and collisions.

  4. It complicates transfer from human ego data, the easiest form of manipulation data to scale. Imagine picking up something with a camera on your hand!

So how could we get the benefits of wrist cameras purely from an ego view, while preserving the ability to search the scene independently of manipulation?1

Robots should look around.

Inspired by human vision, EyeRobot 2.0 uses active gaze to look at what matters, carrying out precise, multi-step tasks from an ego-centric stereo view.

It matches wrist cameras and beats them under tool occlusions.

See experiments and full results →

Average real-world success rate: No Gaze 27%, Ego plus Wrist 52%, EyeRobot 2.0 67%, over 640 trials

See what EyeRobot 2.0 sees

When most people think of active vision, they think of information seeking behaviors like moving your neck to get a better viewpoint. In this work we study a type of active vision more akin to information filtering. Given an already well-observable workspace, can active vision improve manipulation performance by focusing compute, attention, and representation capacity towards important parts of the scene? We borrow the biological term for this type of active vision: fixation. Concentric rectangles illustrate the model's foveated2 image processing, which allocates high resolution image tokens towards the center.

The robot learns to coordinate its gaze with a two-stage reinforcement learning (RL) pipeline, with no gaze demonstrations required. We first train a gaze controller which can look at each per-task object, then learn a gaze sequencing policy which figures out when to look at each object. Like in our prior EyeRobot work, the gripper policy co-trains with the gaze sequencing policy in a behavior cloning (BC) - RL loop. The gripper is trained to mimic demonstrations while the gaze sequencing policy optimizes the gripper policy's prediction accuracy as its reward.

How it works

EyeRobot 2.0 separates where to look from how to act. A target selector chooses an object, a gaze policy looks at it, and a gripper policy predicts actions from the resulting fixated stereo views. The target selector and gripper policy learn together, favoring gaze choices that make actions easier to predict.

We learn where to look in two stages: first a gaze policy that fixates on a requested object, then a target selector that chooses what to look at during manipulation. Both learn through RL without gaze demonstrations.

Synthesizing gaze from real-world data

Original left stereo image showing yellow and grey tapeL
Original right stereo image showing yellow and grey tapeR
sampled view fixation
L
R

Loading the recorded stereo pair…

To train gaze with RL, we need a massively parallel environment to do it in. EyeRobot 1.0 used panoramic 360 images to make an RL gym, but here we need high-resolution stereo. We can do this from real-world stereo video by synthesizing gaze using calibrated image geometry, treating each eye as a partial 360 image to look around in. This reproduces how the scene would look as the eyes rotate, within the original cameras’ field of view.

Goal-conditioned gaze

Goal-conditioned gaze transformer Two foveated stereo views, eye proprioception, and a target prompt feed a gaze transformer through MLP projection. The transformer outputs delta theta, delta phi, and delta distance to control two eyes and a 3D fixation point. The moving fixation point is labeled Depth during radial motion and Direction during angular motion. Foveated stereo Eye Proprio (θ, ϕ, d) Prompt “a yellow tape” MLP projection Gaze Transformer Δθ Δϕ Δd Direction Depth Direction · depth

Our gaze controller continuously servos the 2 eyes to fixate on a target in 3D. It takes in foveated input views, the current eye proprioception, and a goal to look at, and outputs eye motion. We train it with PPO with a dense reward based on the detected positions of objects in 3D. In a sense, you could consider this step distilling slow foundation-model object detection into a fast, realtime servoing policy.

Sequencing gaze

We train a tiny RL agent target selector to choose what object to look at during each stage of the task, then pass this choice to the goal-conditioned gaze policy described above. It learns in parallel with the BC gripper policy, using the BC-RL framework we proposed in EyeRobot 1.0 where gaze gets rewarded with the BC policy's prediction accuracy.

The target selector chooses an object. The selected target and the same fixated views feed parallel gripper and gaze policies. The gripper predicts an action chunk, whose error against the demonstration rewards the selector. The gaze policy updates fixation for the next observation.

Take a look at the output probabilities from the target selector during execution here!

Fixation-centric gripper policy

The gripper policy combines stereo image views with the robot's state to predict a chunk of bimanual actions. It uses an ACT-style transformer decoder, except all input observations are expressed relative to fixation. First, images are passed in as fixation-centric multi-resolution crops, preserving detail at the center and providing larger context at the periphery. Second, we represent end effector proprioception and actions relative to a moving gaze reference frame, which canonicalizes them into a more compact input space.

Fixated stereo views

Left and right eye views, each represented by a pyramid of centered crops with more detail near fixation.
Frozen DINOv3 Multi-scale image tokens
Gripper & gaze state Proprioception, gaze depth and direction
Target object “yellow tape”

Gripper policy

Transformer decoder

Image, state & goal tokens Cross-attention Action queries

Fixation-relative action chunk

The physical stereo camera stays fixed while a gaze point moves gently around the tape. The fixation frame turns from an origin between the eyes, keeping its optical +z axis aimed at the gaze point. The world frame and predicted gripper poses remain stationary. Bimanual SE(3) poses + gripper positions

Compacting the data manifold with fixation

Loosely speaking, there's a few main ways to make a neural net better: 1) improve the architecture or optimization dynamics to better interpolate your data, 2) increase data manifold coverage by using more data, or 3) increase data manifold coverage by making the manifold itself more compact. Much of EyeRobot 2.0's design follows the third approach, reducing the complexity of the learning problem by representing manipulation relative to fixation3. This shares a motivation with equivariant learning: reducing variation by expressing observations and actions in a common frame, here defined by gaze.

Fixation-centric observations

By gazing at a point, EyeRobot 2.0 compacts the input image manifold by centering views on regions of interest. This is particularly noticable in pick and place tasks shown below.

Fixation-centric action frame

The same idea applies to the action space. In world coordinates, reaching for the same object at different locations produces different gripper trajectories. Expressing proprioception and actions relative to fixation brings these reaches into a more compact manifold, reducing the variation the policy needs to learn. This lets the same data more densely cover the relevant motions, rather than spreading it across different object locations. See action-frame ablation →

World frame

Relative to the robot

Fixation frame

Relative to the gaze direction

Point to a frame to locate the pose above

Experiments

We compare against two baselines that use the same policy architecture as EyeRobot 2.0 and are trained on the exact same demonstrations4. The first baseline uses the RGB stereo inputs directly ("No Gaze"), and the second uses one stereo image along with two wrist-mounted cameras ("Ego + Wrist"). Note that both baselines have a significant observability advantage, since EyeRobot 2.0 cannot see the entire stereo image at once. Both baselines receive a full view of the table, while the wrist-camera baseline also receives two additional high-resolution views of the gripper.

1.5x speed, autonomous

In total, we evaluated over 1,000 real-world rollouts across all baselines and ablations for EyeRobot 2.0. We controlled evaluation conditions as tightly as possible, running 25 trials per policy-task pair with the same initial object positions across policies (shown below).

Evaluation spatial distribution for the marker task Marker
Evaluation spatial distribution for the toaster task Toaster
Evaluation spatial distribution for the wrench task Wrench
Evaluation spatial distribution for the boba task Boba
Evaluation spatial distribution for the tea task Tea
Evaluation spatial distribution for the tape task Tape
Evaluation spatial distribution for the pot lid task Pot
Visualization of the 25 evaluation locations used for all policies. Objects are aligned with image overlays for consistent comparison.

How does it compare to SOTA camera configs?

Active stereo >> passive stereo

Both policies start from the exact same source stereo images and are trained on the same demonstrations. By learning where to look within those images, EyeRobot 2.0 raises average success from 27% to 67%. Passive stereo struggles with spatial understanding; learning about 3D servoing is harder from a static perspective than from a fixation-centered perspective.

Average task success
Average task success
Passive stereoEyeRobot 2.0

Passive stereo also lacks the localized resolution necessary for precise tasks, particularly noticeable in its near-zero performance on the marker task. Matching the resolution of EyeRobot 2.0’s fovea with uniform resolution would require over 10× the number of tokens.

Token grids comparing dense stereo images with EyeRobot 2.0's three-level foveated stereo pyramid
Achieving the same resolution as EyeRobot's foveated inputs with passive stereo would require a cumbersome amount of input tokens. This is even more problematic at higher resolutions; foveation requires roughly linear scaling, dense tokenization requires quadratic.

Active stereo > wrist cameras under tool occlusion

When wrist cameras have a clear view, EyeRobot 2.0 matches the ego + wrist baseline. When they don’t, the gap widens. We intentionally collected tasks that are hard to see from the YAM’s wrist cameras because grasped objects block their views. In boba and pot lid, the wrist-camera policy must rely on its ego-centric view for useful information, and performance drops to roughly that of passive stereo.

Success under tool occlusion
Success under tool occlusion
Passive stereoEgo + WristEyeRobot 2.0
Ego + Wrist
EyeRobot 2.0

Different wrist camera designs, such as dual cameras or fisheye views, can alleviate these problems, but cannot always maintain good observability. Teleoperators may also have to use unintuitive manipulation strategies to keep objects visible, which can harm data quality.

Active stereo = wrist cameras in unoccluded tasks

Using the same input images as passive stereo, active fixation closes the performance gap between wrist-free and wrist-camera policies. We evaluate marker, tea, toaster, and wrench with unobstructed wrist camera views. These tasks require precise grasping and tolerances, making clear wrist views particularly helpful. Even on these tasks, EyeRobot 2.0 matches or exceeds the performance of the wrist camera baseline.

Success with unobstructed wrist views
Success with unobstructed wrist views
Ego + WristEyeRobot 2.0
Ego + Wrist
EyeRobot 2.0

Robustness to visual clutter

We evaluated 3 tasks with unseen distractor objects scattered throughout the workspace. All comparisons use the exact same object positions to ensure comparability. Notice how in these examples the wrist camera often reaches for the largest or most distinct object in its view. Once it begins reaching the policy cannot recover since the wrist camera loses sight of the target. In contrast, EyeRobot 2.0 reaches for the correct object when its gaze is correctly directed at it, and mostly misses grasps when it is looking in the wrong place. Decoupling vision and action could improve robustness in cluttered scenes by literally ignoring irrelevant information, as foveation does.

EyeRobot 2.0 has higher success and progression rates, as it can ignore distractor objects and focus on the correct objects.

Average performance with distractors
Average performance with distractors
Passive stereoEgo + WristEyeRobot 2.0

See more evaluation successes and failures for each task here.

Simulated Eval

We found it extremely helpful for project development to build a MuJoCo twin of the YAM arms for high-volume, low-noise evaluations. Our goal here was not sim-to-real transfer, but rather a clean simulation-only benchmark that replicates as closely as possible the real world pipeline. We built a VR teleoperation platform using the same GELLO leader arms as real teleop, and designed 6 tasks by importing assets from SAM-3D. Task stage success is labeled based on programatic physics checks. We will open-source the full simulation setup, including the task environments and teleoperation, training, and evaluation code.

We train models on these simulated-only tasks with the same method, hyperparameters, camera parameters, and robot positions as in our real-world experiments. The method uses no information that would be unavailable to the real robot.

Simulation task success
Simulation task success
Passive stereoEgo + WristEyeRobot 2.0

Ablations

Stereo and foveated vision

Ablating either stereo inputs or foveal inputs harms policy performance. Take a look at the sankey diagrams for a stage-by-stage breakdown; the failure cases of these ablations tend to cluster around stages requiring precision like capping the marker, or pulling out the toaster oven tray. Interestingly, boba performance was unaffected which we think was because the most precise phase of inserting the straw was performed near the stereo camera, making any gains from foveation less beneficial.

Ablation table: per-task success rates for EyeRobot 2.0, EyeRobot 2.0 without foveation, and EyeRobot 2.0 without stereo (mono)
EyeRobot 2.017/25 succeeded · 68%
EyeRobot 2.0, Marker: stage-by-stage trial progression; 17 of 25 trials succeeded
Without foveation5/25 succeeded · 20%
Without foveation, Marker: stage-by-stage trial progression; 5 of 25 trials succeeded
Without stereo8/25 succeeded · 32%
Without stereo, Marker: stage-by-stage trial progression; 8 of 25 trials succeeded

Gaze-centric actions

Instead of fixation-relative actions, we train and test models in simulation using world-frame SE(3) inputs and outputs. These also degrade in performance and interestingly roughly match the performance of the entirely non-fixating policy. This seems to suggest that without fixation-frame canonicalization, the model can learn proprioceptive shortcuts without looking at the image. One interpretation is that data coverage of the task in world-frame is sparser, leading to a worse policy.

Fixation-relative action ablation
Fixation-relative action ablation
World-relative actionsEyeRobot 2.0

Failures

EyeRobot 2.0's failures most often come from narrowly missing a stage and entering an unrecoverable state. It also experiences the common failure cases of low-data policies trained without DAgger data, namely very minimal recovery capabilities, and a tendency to sometimes get stuck if it enters an odd state (see toaster example).

Citation

If you use this work or find it helpful, please consider citing: (bibtex)

@article{2026eyerobot2,
  title={{EyeRobot 2.0}: Active Gaze for Precise Manipulation without Wrist Cameras},
  author={Hari, Kush* and Kerr, Justin* and Shivakumar, Nidhya and Mahapatra, Samarth and Sferrazza, Carmelo and Lei, Jiahui and Malik, Jitendra and Liu, C. Karen and Goldberg, Ken and Kanazawa, Angjoo},
  journal={arXiv preprint},
  year={2026}
}