Bridge3D: Enabling Vision–Language–Action Models to See and Act in 3D

Real-world manipulation. Six tasks on the XMAN-R1 robot, shown at 4× speed.

Explore the Tasks

84.8%

RoboTwin 2.0

+11.7pp

Real-world improvement

6.1×

Fewer demonstrations

Overview

Vision–Language–Action (VLA) models demonstrate strong generalization in robotic manipulation through large-scale multimodal pretraining. However, their reliance on 2D-centric observations limits precise spatial manipulation. Existing approaches introduce implicit spatial priors to improve 3D awareness, but still lack explicit geometry guidance for action generation.

Bridge3D integrates both implicit and explicit 3D geometry guidance into pretrained 2D VLA models, enabling them to “see” and “act” in 3D. Implicit Fusion enriches visual tokens with features from 3D foundation models to improve 3D perception. Explicit Conditioning integrates action denoising with an explicit 3D semantic field to guide precise action generation. Layer-wise linear probing identifies critical blocks for geometric conditioning, improving learning efficiency.

Experiments demonstrate strong performance in high-precision, spatially sensitive manipulation. RDT-Bridge3D achieves the highest average success rate of 84.8% on the 10 selected RoboTwin 2.0 tasks under clean scenes. In real-world experiments, Bridge3D reaches 49.0% average success, outperforming Spatial Forcing by 11.7 percentage points.

Method

Bridge3D architecture: a frozen 3D visual foundation model augments 2D visual tokens through a trainable geometry adapter; the action expert queries a 3D semantic field through spatial cross-attention.
The architecture of Bridge3D. Implicit Fusion augments 2D visual features with geometry-aware representations from a pretrained 3D foundation model. Explicit Conditioning integrates a 3D semantic point field into the action expert for precise action generation.

Implicit Fusion

A lightweight residual adapter fuses VGGT geometry features with 2D visual tokens while preserving the pretrained VLM representation.

Explicit Conditioning

Spatial cross-attention and 3D rotary position embeddings let action hidden states query a semantic point field in a shared coordinate frame.

Selective Conditioning

Linear probing identifies spatial degradation in RDT blocks 7–11, directing geometric guidance to these layers.

Results

RoboTwin 2.0 · Clean scenes

GR00T-N1.5
47.5
GR00T-N1.5-Bridge3D
54.4
RDT
68.7
π₀
71.7
Spatial Forcing
77.1
RDT-Bridge3D
84.8
RoboTwin 2.0 results. RDT-Bridge3D achieves the highest average success rate of 84.8% across 10 selected tasks, surpassing 3D Diffusion Policy (82.8%), Spatial Forcing (77.1%), and PointVLA (76.0%).
Full benchmark results
MethodRoboTwin cleanRoboTwin randomizedReal world
Diffusion Policy49.00.0—
ACT72.85.5—
3D Diffusion Policy82.84.1—
GR00T-N1.547.513.726.3
RDT68.726.134.3
π₀71.732.7—
Spatial Forcing77.135.837.3
PointVLA76.0——
GR00T-N1.5-Bridge3D54.421.2—
RDT-Bridge3D84.833.649.0

Real world experiments

GR00T-N1.5RDTSpatial ForcingRDT-Bridge3D
Real world experimentsSuccess rate (%). Bars start at zero.Success rate (%)0255075100InsertFlowerGR00T-N1.5 · Insert Flower: 12%RDT · Insert Flower: 24%Spatial Forcing · Insert Flower: 28%RDT-Bridge3D · Insert Flower: 46%HangMugGR00T-N1.5 · Hang Mug: 32%RDT · Hang Mug: 28%Spatial Forcing · Hang Mug: 32%RDT-Bridge3D · Hang Mug: 40%PutObjectGR00T-N1.5 · Put Object Cabinet: 36%RDT · Put Object Cabinet: 56%Spatial Forcing · Put Object Cabinet: 52%RDT-Bridge3D · Put Object Cabinet: 68%SweepTrashGR00T-N1.5 · Sweep Trash: 8%RDT · Sweep Trash: 18%Spatial Forcing · Sweep Trash: 30%RDT-Bridge3D · Sweep Trash: 42%PlaceCupGR00T-N1.5 · Place Cup Rack: 42%RDT · Place Cup Rack: 48%Spatial Forcing · Place Cup Rack: 52%RDT-Bridge3D · Place Cup Rack: 54%ClearTrashGR00T-N1.5 · Clear Trash: 28%RDT · Clear Trash: 32%Spatial Forcing · Clear Trash: 30%RDT-Bridge3D · Clear Trash: 44%Avg.GR00T-N1.5 · Average: 26.3%RDT · Average: 34.3%Spatial Forcing · Average: 37.3%RDT-Bridge3D · Average: 49%
Experimental results on real robot. RDT-Bridge3D outperforms all compared methods across the six tasks, with a 42.9% relative gain over RDT on average. Bridge3D demonstrates a substantial advantage in spatial manipulation tasks that require high precision, such as Insert Flower.

Layer-wise linear probe

t = 0.8
Layer-wise linear probeError (×10⁻²). Bars start at zero.Error (×10⁻²)024680t = 0.8 · 0: 1.146 ×10⁻²t = 0.8 · 1: 1.253 ×10⁻²t = 0.8 · 2: 1.506 ×10⁻²t = 0.8 · 3: 1.675 ×10⁻²t = 0.8 · 4: 1.703 ×10⁻²5t = 0.8 · 5: 1.696 ×10⁻²t = 0.8 · 6: 1.686 ×10⁻²t = 0.8 · 7: 2.143 ×10⁻²t = 0.8 · 8: 2.174 ×10⁻²t = 0.8 · 9: 2.557 ×10⁻²10t = 0.8 · 10: 2.397 ×10⁻²t = 0.8 · 11: 3.196 ×10⁻²t = 0.8 · 12: 3.280 ×10⁻²t = 0.8 · 13: 3.183 ×10⁻²t = 0.8 · 14: 3.094 ×10⁻²15t = 0.8 · 15: 3.798 ×10⁻²t = 0.8 · 16: 3.506 ×10⁻²t = 0.8 · 17: 3.502 ×10⁻²t = 0.8 · 18: 3.584 ×10⁻²t = 0.8 · 19: 4.040 ×10⁻²20t = 0.8 · 20: 3.669 ×10⁻²t = 0.8 · 21: 3.839 ×10⁻²t = 0.8 · 22: 3.873 ×10⁻²t = 0.8 · 23: 3.896 ×10⁻²t = 0.8 · 24: 4.425 ×10⁻²25t = 0.8 · 25: 4.207 ×10⁻²t = 0.8 · 26: 3.880 ×10⁻²t = 0.8 · 27: 3.956 ×10⁻²28t = 0.8 · 28: 4.763 ×10⁻²Block index
Layer-wise linear probing. Both spatial localization error and variance rise sharply across blocks 7–11, revealing degraded spatial awareness and motivating layer-selective Explicit Conditioning.

Data efficiency

RDTRDT-Bridge3D
Data efficiencySuccess rate (%). Bars start at zero.Success rate (%)02550751002RDT · 2: 38%38RDT-Bridge3D · 2: 27%275RDT · 5: 42%42RDT-Bridge3D · 5: 58%5810RDT · 10: 53%53RDT-Bridge3D · 10: 83%8325RDT · 25: 70%70RDT-Bridge3D · 25: 95%9550RDT · 50: 74%74RDT-Bridge3D · 50: 98%98Training demonstrations
Data efficiency. Bridge3D improves learning from limited demonstrations, requiring 6.1× fewer training samples than RDT to reach 74% success. With 50 demonstrations, it achieves 98% success, highlighting the benefits of 3D geometric guidance for sample efficiency.

Key component ablation

w/o IF & ECw/o ECw/o IFFull Bridge3D
Key component ablationSuccess rate (%). Bars start at zero.Success rate (%)0255075100w/oIF & ECw/o IF & EC · Average: 78.8%78.8w/o ECw/o EC · Average: 90.4%90.4w/o IFw/o IF · Average: 92%92FullBridge3DFull Bridge3D · Average: 95.6%95.6Average
Ablation on key components in Bridge3D. Implicit Fusion (IF) and Explicit Conditioning (EC) play complementary roles. Removing either module reduces task success, while the full framework achieves the highest average success rate of 95.6% across five tasks.

Point semantics & encoder

Without point semanticsWith PointNet++Raw semantic points
Point semantics & encoderSuccess rate (%). Bars start at zero.Success rate (%)0255075100GeneralgraspWithout point semantics · General grasp: 96%96With PointNet++ · General grasp: 80%80Raw semantic points · General grasp: 98%98FunctionalmanipulationWithout point semantics · Functional manipulation: 45%45With PointNet++ · Functional manipulation: 43%43Raw semantic points · Functional manipulation: 53%53AverageWithout point semantics · Average: 70.5%70.5With PointNet++ · Average: 61.5%61.5Raw semantic points · Average: 75.5%75.5
Ablation on Point Semantics and Point-cloud Encoder. Point-wise semantics improve fine-grained functional manipulation, while PointNet++ encoding reduces success across both tasks. Preserving raw point-wise geometry achieves the highest average success rate of 75.5%.

Spatial generalization

RDTRDT-Bridge3D
Spatial generalizationSuccess rate (%). Bars start at zero.Success rate (%)02550751000RDT · 0: 74%74RDT-Bridge3D · 0: 80%801RDT · 1: 52%52RDT-Bridge3D · 1: 76%763RDT · 3: 38%38RDT-Bridge3D · 3: 68%685RDT · 5: 12%12RDT-Bridge3D · 5: 48%48Height offset (cm)
Spatial Generalization. Trained at a single tabletop height, Bridge3D generalizes to unseen height offsets of 1, 3, and 5 cm, reducing performance degradation by 48.4% relative to RDT.

Visualization of 3D Semantic Fields and Spatial Attention

Three robotic manipulation examples showing input observations in the top row, DINOv3 features in the middle row, and constructed 3D semantic fields in the bottom row.
Qualitative visualization of 3D semantic fields. We visualize the 3D semantic fields across a variety of robotic manipulation scenarios. For each example, we show the input observations (top), DINOv3 (Siméoni et al. 2025) features (middle), and the constructed 3D semantic fields (bottom).
Attention maps for cup and bell manipulation, comparing observations, attention with Implicit Fusion, and attention without Implicit Fusion.
Spatial priors guide action conditioning. With Implicit Fusion, spatial cross-attention focuses on task-relevant geometry, such as the cup’s edge and the bell’s button; without it, attention is scattered across the scene. These spatial priors guide downstream Explicit Conditioning.

Citation

@article{li2026bridge3d,
      title={Bridge3D: Enabling Vision-Language-Action Models to See and Act in 3D},
      author={Li, Haoxuan and Yan, Sixu and Zhu, Lianghui and Tang, Xuanlai and Wang, Shikang and Wang, Xinggang},
      journal={arXiv preprint arXiv:2609.24525},
      year={2026}
    }

Research figure