*Equal contribution · †Corresponding author
The target is defined by a past event and may be hidden by the time the robot acts. BeyondCSe resolves the reference from the event history, then moves the camera until it can grasp.
One event history, several instructions. Switch the instruction to change the target. During execution the robot sees only its wrist camera; the third-person footage is for you. Executions at 2× speed.
A baseball, a red bell pepper, and a Rubik's cube are handled in sequence and end up out of the wrist camera's initial view. The three instructions differ only in event order. For each, the system recovers the target's past track and selects a view that reveals it before grasping.
The instruction names a tool by what it did and in which order, so the system must link “drive the screw” to the right object in the event record among visually similar distractors.
The can ends up behind a blue crate. VLM Orchestrator (VoLo) correctly identifies the can, but neither π0.5 nor Cosmos 3 can locate it in the current view. BeyondCSe moves the camera to see behind the crate, then grasps. Baselines follow the DROID setup with wrist and third-person cameras; ours uses the wrist camera only.
An event-referential instruction names a target by its role in a past event, not by how it looks now. BeyondCSe uses the event history twice: to identify the object or part, and to build an event prior over where it went, so the camera can move to a view that reveals it.

(A) VoLo identifies the target and hands a text description to a VLA (π0.5) or a world action model (Cosmos 3). Both fail on the occluded can. (B) BeyondCSe initializes a spatial belief from the can's past 3D observations, selects a revealing viewpoint, and grasps it.
LERF-TOGO, GraspSplats, Point2Act, GraspMolmo resolve names and parts in the current scene. An event-defined target is not identifiable from the current view.
Memory VLAs and orchestrators like VoLo can recover which object is meant, but give no current visual evidence once it is hidden.
Resolve the reference, recover the target's past 3D track, and drive viewpoint selection with an event-conditioned belief until the target is graspable.
Two modules linked by the target's description and 3D history. Visible target: grasp directly. Hidden target: its past observations seed an active search. Expand a module for details.

Four stages of one off-the-shelf MLLM turn the video and instruction into a pixel. If the target is not visible now, its past frames are tracked and lifted to 3D: the event track 𝒫 that seeds the search.
Video → time-ordered events, written without seeing the instruction.
Events + instruction → what the target looks like, and which event singles it out.
Cue → a box in the initial image that narrows the search.
Crop + target → a pixel, lifted to 3D with depth and pose.

Cropping to the located region is what lets Point hit small parts such as a stem.

A volumetric belief over where the target is, initialized from its event track, is updated after every view. Candidate views are scored by how much belief they would clear, discounting sight lines through space no one has observed yet.
Gaussian centred on the last past observation, shaped by the tail of the track, mixed with a uniform floor.
Feasible views ranked by expected belief cleared, weighted by transmittance through unobserved space.
New depth and a target-miss likelihood downweight, but never exclude, hypotheses.
Once pointed, validate AnyGrasp candidates; one refinement view if geometry is incomplete.
ROBOTIS OMY-F3M arm, one wrist-mounted RealSense D435i. All methods share the robot, calibration, collision checks, and grasp pipeline.
50 visible and 100 occluded trials per method. Reasoned instructions are grounding queries generated by Qwen3-VL-8B from the history video, shared across baselines. Segments show success or the first failed stage. GraspMolmo (GM) uses only the initial view and is not evaluated on occluded targets.

LERF*, P2A, GS, GM denote LERF-TOGO, Point2Act, GraspSplats, GraspMolmo. P2A† uses RGB-D reconstruction. Multi-view baselines receive 30 predefined views; ours starts from the same initial view and adds views only as needed.
Pooled across visible and occluded conditions.
| Instruction | LERF* | P2A | GS | P2A† | Ours |
|---|---|---|---|---|---|
| Original | 0.0 | 9.3 | 0.0 | 22.7 | 76.7 |
| Reasoned | 18.7 | 19.3 | 34.7 | 50.0 |
Grasp success in %. Ours uses its own video reasoning in both conditions.
Four heavily occluded scenes, five trials each. Breyer et al. is given the ground-truth 3D target box; ours finds the target from the event history alone.
| Method | S1 | S2 | S3 | S4 | Views ↓ | Grasp ↑ |
|---|---|---|---|---|---|---|
| Breyer et al. (GT box) | 4.2 | 4.2 | 2.2 | 2.8 | 3.35 | 15/20 (75%) |
| Ours w/o belief weights | 4.0 | 4.0 | 5.6 | 5.0 | 4.65 | 16/20 (80%) |
| Ours w/o transmittance | 2.8 | 2.0 | 5.2 | 2.4 | 3.10 | 18/20 (90%) |
| Ours w/o event prior | 3.8 | 4.0 | 4.2 | 4.0 | 4.00 | 19/20 (95%) |
| Ours (full) | 2.4 | 2.0 | 2.4 | 2.0 | 2.20 | 19/20 (95%) |
With a uniform initial belief instead of the history-derived one, grasp success stays at 95% but the mean view count rises from 2.20 to 4.00. In the qualitative comparison in the paper, ours reaches a revealing viewpoint in one move; Breyer et al., despite a privileged target box, explores several first.
Without transmittance attenuation, the mean view count rises to 3.10, with one scene needing 5.2 views. Without belief weights, all remaining hypotheses count equally: 4.65 views and 80% success.
A video-derived grounding query improves every baseline, but the best reaches 50.0% overall against 76.7% for ours. Localization is the dominant failure: we localize 43/50 visible and 88/100 occluded targets, 32 and 20 points above the strongest baselines.
An event history should guide both target identification and spatial search. Turning the target's own past observations into a belief, and scoring views with transmittance-aware visibility, lets a zero-shot system grasp what it cannot yet see with fewer camera moves than a baseline handed the answer's location.
@article{lee2026beyondcse,
title = {Beyond the Current Scene: Event-Referential Grasping
with Active View Selection},
author = {Lee, Hyunjoon and Jung, Haebeom and Cha, Eunsung and
Lee, Daeun and Choe, Jaesung and Wang, Yu-Chiang Frank and
Park, Jaesik},
journal = {arXiv preprint arXiv:2609.39375},
year = {2026}
}