BeyondCSe
arXiv 2026
Zero-shot event-referential grasping

Beyond the Current Scene Event-referential grasping with active view selection

*Equal contribution  ·  †Corresponding author

1Seoul National University 2NVIDIA
Wrist camera only

The target is defined by a past event and may be hidden by the time the robot acts. BeyondCSe resolves the reference from the event history, then moves the camera until it can grasp.

Grasp success 150 real-robot trials · zero-shot
Best baseline · visible: 40% (20/50) 40% BeyondCSe · visible: 76% (38/50) 76% Visible target Best baseline · occluded: 55% (55/100) 55% BeyondCSe · occluded: 77% (77/100) 77% Occluded target
Strongest baselineBeyondCSeFull results
01Demonstrations

Same history, different question.

One event history, several instructions. Switch the instruction to change the target. During execution the robot sees only its wrist camera; the third-person footage is for you. Executions at 2× speed.

Occluded target

Objects hidden behind a box and inside a crate

A baseball, a red bell pepper, and a Rubik's cube are handled in sequence and end up out of the wrist camera's initial view. The three instructions differ only in event order. For each, the system recovers the target's past track and selects a view that reveals it before grasping.

Event history video
Rec · wrist camera
Robot execution3rd-person view for visualization
Wrist camera only
Visible target

Tools referenced by the action they were used for

The instruction names a tool by what it did and in which order, so the system must link “drive the screw” to the right object in the event record among visually similar distractors.

Event history video
Rec · wrist camera
Robot execution3rd-person view for visualization
Wrist camera only
Simulation

“Pick up the object moved last” against VLA/WAM baselines

The can ends up behind a blue crate. VLM Orchestrator (VoLo) correctly identifies the can, but neither π0.5 nor Cosmos 3 can locate it in the current view. BeyondCSe moves the camera to see behind the crate, then grasps. Baselines follow the DROID setup with wrist and third-person cameras; ours uses the wrist camera only.

Event history video
Rec · front camera
“Pick up the object moved last”
Baseline comparisonOurs: wrist only · Baselines: wrist + 3rd-person
VoLo + π0.5VLA · wrist + 3rd-person · Fails
VoLo + Cosmos 3WAM · wrist + 3rd-person · Fails
BeyondCSeActive view · wrist only · Success
02The problem

The current scene alone cannot tell a robot what to grasp, or where to look.

An event-referential instruction names a target by its role in a past event, not by how it looks now. BeyondCSe uses the event history twice: to identify the object or part, and to build an event prior over where it went, so the camera can move to a view that reveals it.

“Pick up the handle of the cup holding the pear.”  ·  “Pick up the tool I used to drive the screw at the beginning.”
Comparison: VoLo identifies the can but the downstream VLA or world action model fails to grasp the occluded target; BeyondCSe initializes a spatial belief from past observations, selects a revealing viewpoint, and succeeds.Comparison: VoLo identifies the can but the downstream VLA or world action model fails to grasp the occluded target; BeyondCSe initializes a spatial belief from past observations, selects a revealing viewpoint, and succeeds.

(A) VoLo identifies the target and hands a text description to a VLA (π0.5) or a world action model (Cosmos 3). Both fail on the occluded can. (B) BeyondCSe initializes a spatial belief from the can's past 3D observations, selects a revealing viewpoint, and grasps it.

(a) Current-scene grounding

Blind to history

LERF-TOGO, GraspSplats, Point2Act, GraspMolmo resolve names and parts in the current scene. An event-defined target is not identifiable from the current view.

(b) Memory-conditioned policies

Identity without evidence

Memory VLAs and orchestrators like VoLo can recover which object is meant, but give no current visual evidence once it is hidden.

(c) BeyondCSe

Identify, then go look

Resolve the reference, recover the target's past 3D track, and drive viewpoint selection with an event-conditioned belief until the target is graspable.

03Method

Reason over the video, then search the scene.

Two modules linked by the target's description and 3D history. Visible target: grasp directly. Hidden target: its past observations seed an active search. Expand a module for details.

System overview: video reasoning produces a target description and event prior; event-conditioned active perception selects the next view and updates from new observations until grasp validation succeeds.System overview: video reasoning produces a target description and event prior; event-conditioned active perception selects the next view and updates from new observations until grasp validation succeeds.
Module A

Video reasoning: instruction → target + cue → action point.

Four stages of one off-the-shelf MLLM turn the video and instruction into a pixel. If the target is not visible now, its past frames are tracked and lifted to 3D: the event track 𝒫 that seeds the search.

01 · Record

Event record

Video → time-ordered events, written without seeing the instruction.

02 · Select

Target & cue

Events + instruction → what the target looks like, and which event singles it out.

03 · Locate

Region proposal

Cue → a box in the initial image that narrows the search.

04 · Point

Action pixel

Crop + target → a pixel, lifted to 3D with depth and pose.

Video reasoning pipeline (Record, Select, Locate, Point) and a pointing example where crop-based pointing succeeds on a small flower stem while direct and full-image pointing fail.Video reasoning pipeline (Record, Select, Locate, Point) and a pointing example where crop-based pointing succeeds on a small flower stem while direct and full-image pointing fail.

Cropping to the located region is what lets Point hit small parts such as a stem.

Module B

Event-conditioned active perception: a belief that searches.

Event-conditioned active perception: event track and current observation initialize the belief and TSDF map; belief-guided view search selects a view; target confirmation triggers grasp validation.Event-conditioned active perception: event track and current observation initialize the belief and TSDF map; belief-guided view search selects a view; target confirmation triggers grasp validation.

A volumetric belief over where the target is, initialized from its event track, is updated after every view. Candidate views are scored by how much belief they would clear, discounting sight lines through space no one has observed yet.

Stage 1

Event prior

Gaussian centred on the last past observation, shaped by the tail of the track, mixed with a uniform floor.

Stage 2

View search

Feasible views ranked by expected belief cleared, weighted by transmittance through unobserved space.

Stage 3

Bayesian update

New depth and a target-miss likelihood downweight, but never exclude, hypotheses.

Stage 4

Confirm & grasp

Once pointed, validate AnyGrasp candidates; one refinement view if geometry is incomplete.

04Results

150 real-robot trials, and fewer views when it matters.

ROBOTIS OMY-F3M arm, one wrist-mounted RealSense D435i. All methods share the robot, calibration, collision checks, and grasp pipeline.

Zero-shot grasping comparison

50 visible and 100 occluded trials per method. Reasoned instructions are grounding queries generated by Qwen3-VL-8B from the history video, shared across baselines. Segments show success or the first failed stage. GraspMolmo (GM) uses only the initial view and is not evaluated on occluded targets.

Stacked bars of localization failure, plan failure, grasp failure, and success for LERF-TOGO, Point2Act, GraspSplats, GraspMolmo, Point2Act with RGB-D, and ours, under original and reasoned instructions, visible and occluded.Stacked bars of localization failure, plan failure, grasp failure, and success for LERF-TOGO, Point2Act, GraspSplats, GraspMolmo, Point2Act with RGB-D, and ours, under original and reasoned instructions, visible and occluded.

LERF*, P2A, GS, GM denote LERF-TOGO, Point2Act, GraspSplats, GraspMolmo. P2A† uses RGB-D reconstruction. Multi-view baselines receive 30 predefined views; ours starts from the same initial view and adds views only as needed.

Grasp success over 150 trials

Pooled across visible and occluded conditions.

InstructionLERF*P2AGSP2A†Ours
Original0.09.30.022.776.7
Reasoned18.719.334.750.0

Grasp success in %. Ours uses its own video reasoning in both conditions.

Active perception & belief ablations

Four heavily occluded scenes, five trials each. Breyer et al. is given the ground-truth 3D target box; ours finds the target from the event history alone.

MethodS1S2S3S4Views ↓Grasp ↑
Breyer et al. (GT box)4.24.22.22.83.3515/20 (75%)
Ours w/o belief weights4.04.05.65.04.6516/20 (80%)
Ours w/o transmittance2.82.05.22.43.1018/20 (90%)
Ours w/o event prior3.84.04.24.04.0019/20 (95%)
Ours (full)2.42.02.42.02.2019/20 (95%)
Finding 1

The event prior does not change whether the target is found, but how many views it takes.

With a uniform initial belief instead of the history-derived one, grasp success stays at 95% but the mean view count rises from 2.20 to 4.00. In the qualitative comparison in the paper, ours reaches a revealing viewpoint in one move; Breyer et al., despite a privileged target box, explores several first.

Finding 2

Treating unobserved space as free makes the search optimistic and slow.

Without transmittance attenuation, the mean view count rises to 3.10, with one scene needing 5.2 views. Without belief weights, all remaining hypotheses count equally: 4.65 views and 80% success.

Finding 3

Reasoned instructions help every baseline, yet identification alone is not enough.

A video-derived grounding query improves every baseline, but the best reaches 50.0% overall against 76.7% for ours. Localization is the dominant failure: we localize 43/50 visible and 88/100 occluded targets, 32 and 20 points above the strongest baselines.

Key takeaway

An event history should guide both target identification and spatial search. Turning the target's own past observations into a belief, and scoring views with transmittance-aware visibility, lets a zero-shot system grasp what it cannot yet see with fewer camera moves than a baseline handed the answer's location.

↳Cite this work

BibTeX

@article{lee2026beyondcse,
  title   = {Beyond the Current Scene: Event-Referential Grasping
             with Active View Selection},
  author  = {Lee, Hyunjoon and Jung, Haebeom and Cha, Eunsung and
             Lee, Daeun and Choe, Jaesung and Wang, Yu-Chiang Frank and
             Park, Jaesik},
  journal = {arXiv preprint arXiv:2609.39375},
  year    = {2026}
}