哈啰 Robotaxi

HelloWorld

Towards Practical Applications of Generative Driving World Models

HelloWorld Team

Technical Report · Coming soon Explore applications

HelloWorld as a Data Generalizer

Edit a driving scene to generate controllable, multi-modal training data.

HelloWorld as a Closed-Loop Simulator

Recorded driving on the left and an Alpamayo 1.5 closed-loop rollout in HelloWorld on the right. The policy follows a slow-moving lead vehicle and adjusts its gap while a stopped vehicle in the right lane limits an immediate merge.

Model Overview

HelloWorld combines controllable multi-view video generation, few-step inference, super-resolution, and RGB-conditioned LiDAR synthesis in a unified driving world model system.

HelloWorld overview: single-view causal learning, synchronized multi-view generation, geometry-aware scene conditioning, and efficient multi-sensor extensions.
Controllable causal generation with few-step distillation, super-resolution, and conditional LiDAR synthesis.
HelloWorld architecture: causal multi-view backbone, few-step distillation, super-resolution, and RGB-to-LiDAR backbone.
Causal multi-view generation, few-step distillation, super-resolution, and RGB-to-LiDAR synthesis.

Pose Control

These demos prioritize adherence to the prescribed pose trajectory, even when it leads to collisions with barriers or off-road motion. Simulating such extreme trajectories supports closed-loop stress testing and the analysis of driving-policy failure modes.

Turn left · 7 views · 81 frames · 10 fps · 8.1 s. Paired with Straight and Turn right using the same scene and initial frame. Green: projected future pose trajectory.

Single-view pose · 29 frames · Table 3
Single-view pose · 29 frames
MethodRot (°) ↓Trans ↓Trans* (m) ↓
HY-WorldPlay8.63051.0669–
Lingbot World 2.03.23100.2853–
HelloWorld2.06200.25700.6830

Trans uses Sim(3) alignment; Trans* is metric-scale translation error, reported only for HelloWorld on real-world on-road cases. These benchmarks are separate from the seven-view demos. Report · Table 3 · p. 18

Single-view pose · 100 frames · Table 4
Single-view pose · 100 frames
MethodRot (°) ↓Trans ↓Trans* (m) ↓
HY-WorldPlay15.08355.4504–
Lingbot World 2.05.41691.4988–
HelloWorld2.70561.39975.5857

Trans uses Sim(3) alignment; Trans* is metric-scale translation error, reported only for HelloWorld on real-world on-road cases. These benchmarks are separate from the seven-view demos. Report · Table 4 · p. 18

Single-view perceptual quality (VBench) · 29 frames · Table 5
Single-view perceptual quality (VBench) · 29 frames
MethodSubjectBackgroundMotionDynamicAestheticImagingI2V subjectI2V backgroundMean
HY-WorldPlay0.96690.95430.99330.28000.47620.57740.98290.98470.7770
Lingbot World 2.00.91500.94240.97600.89000.48190.59330.96230.96760.8411
HelloWorld0.91480.93640.98070.92000.45210.53920.95630.96320.8328

Mean is the unweighted mean of the eight raw scores; it is not the full VBench benchmark score. Report · Table 5 · p. 19

Single-view perceptual quality (VBench) · 100 frames · Table 6
Single-view perceptual quality (VBench) · 100 frames
MethodSubjectBackgroundMotionDynamicAestheticImagingI2V subjectI2V backgroundMean
HY-WorldPlay0.93540.93250.99510.14000.46850.55840.98440.98550.7500
Lingbot World 2.00.88410.92700.98140.99000.44890.58160.95730.96240.8416
HelloWorld0.88760.92880.98420.92000.44530.49760.95710.96270.8229

Mean is the unweighted mean of the eight raw scores; it is not the full VBench benchmark score. Report · Table 6 · p. 19

Multi-view pose control · Table 7
Multi-view pose control
SettingRot (°) ↓Trans ↓Trans* (m) ↓
Cosmos Transfer 2.5 (w/ ff)5.5341.8808.886
Stage 2 (w/ ff)3.3151.5254.486
Stage 2 (w/o ff)5.1632.1897.972
Stage 3 (w/ ff)2.0271.3524.904
Stage 3 (w/o ff)3.3341.8448.485

Stage 2 uses pose conditioning; Stage 3 adds scene controls. w/ ff: with first-frame conditioning; w/o ff: without it. Trans* is in meters. Report · Table 7 · p. 20

Multi-view cross-camera consistency · Table 8
Multi-view cross-camera consistency
SettingCSE (px) ↓ΔCSE (px) ↓
GT video1.7300.000
Cosmos Transfer 2.5 (w/ ff)6.549+4.819
Stage 2 (w/ ff)6.192+4.462
Stage 2 (w/o ff)7.509+5.779
Stage 3 (w/ ff)4.180+2.450
Stage 3 (w/o ff)4.892+3.162

ΔCSE is measured relative to GT video. CSE measures epipolar agreement, not complete 3D consistency. Report · Table 8 · p. 20

Multi-view perceptual quality (VBench) · Table 9
Multi-view perceptual quality (VBench)
SettingSubjectBackgroundMotionDynamicAestheticImagingMean
Cosmos Transfer 2.5 (w/ ff)0.89160.93760.98780.95550.43660.42390.7722
Stage 2 (w/ ff)0.89990.93240.98430.98000.43740.49530.7882
Stage 2 (w/o ff)0.87160.91990.98220.97000.44660.48440.7791
Stage 3 (w/ ff)0.90890.94620.98760.95000.41380.44560.7754
Stage 3 (w/o ff)0.91110.94720.97880.97000.42590.50740.7901

Mean is the unweighted mean of six raw scores, including dynamic degree. w/ ff and w/o ff denote first-frame conditioning. Report · Table 9 · p. 21

nuScenes driving generation · short horizon · Table 10
nuScenes driving generation · short horizon
MethodARMulti-viewVideoDSFFID ↓FVD ↓
DrivingGPT✗✗✓✓12.78142.61
DrivingWorld✗✗✓✓7.4090.90
Vista✗✗✓✓6.9089.40
Epona✗✗✓✓7.5082.80
BEVControl✗✓✗✓24.85–
BEVGen✗✓✗✓25.54–
MagicDrive✗✓✗✓16.20–
UniScene✗✓✓✗6.4571.94
DiST-4D✗✓✓✗7.4025.55
OmniNWM✗✓✓✗5.4523.63
MagicDrive✗✓✓✓18.75218.12
GenAD✗✓✓✓15.40184.00
Panacea✗✓✓✓16.96139.00
Drive-WM✗✓✓✓15.80122.70
DriveDreamer-2✗✓✓✓25.00105.10
FAR-Drive✓✓✓✓11.9282.78
HelloWorld (Ours)✓✓✓✓7.7541.08

All short-horizon rows from Table 10. AR: autoregressive; DSF: dense-supervision-free. Image-only methods have no FVD. The Video column distinguishes the two MagicDrive settings. These are benchmark results, not scores for the displayed demos. Report · Table 10 · p. 21

Layout Editing

Explore object insertion, removal, and trajectory editing through bird’s-eye-view animations and seven-view generated videos with projected lanes and boxes. A recorded-layout following reconstruction is included for reference.

Edit the lead vehicle trajectory so it brakes and stops in the lane, then inspect the seven-view generated result. 4K · 30 fps · 24.7 s.

RGB-to-LiDAR Generation

The model generates LiDAR conditioned on seven-view RGB observations. The left panel previews the front-wide conditioning view, while the right shows the corresponding generated point cloud over time.

Input condition

RGB video

Front wide camera

Generated output

LiDAR point cloud

Distance color · drag to rotate

Loading generated point cloud…

Front-wide RGB preview beside the generated cloud · 63 frames, 10 fps, 6.3 s. Colors encode sensor distance on a 5–28 m scale, with values outside this interval clamped to the endpoint colors. Drag to rotate, scroll to zoom. All finite returns are drawn; point count varies by frame and case. Download sunset sample PLY ↗

Full-panorama RGB-to-LiDAR generation · Table 17
Full-panorama RGB-to-LiDAR generation
MetricGenerator+ Single-frame RRN+ Temporal RRN
Range MAE (m) ↓2.27972.25322.2194
Chamfer Distance (m) ↓1.07381.10381.0920
Edge Pred→GT (m) ↓0.88140.83390.7990
Ray-Align ↓0.43660.32380.2822
Static-Align ↑0.67410.70810.8310

100 clips (2,900 frames), 35 sampling steps. All arms reuse the same saved generator outputs; RRN does not rerun the DiT. Static-Align uses shared reference-ICP transforms. Alignment scores are fractions. The scoring region, MAE validity support, and point filtering differ from Table 15, so values should not be compared directly across the two tables. These are not per-demo scores. Report · Table 17 · p. 30

LiDAR comparison · common three-camera region · Table 15
LiDAR comparison · common three-camera region
MetricCosmos1Ours+ Temporal RRN
Range MAE (m) ↓6.17242.21992.1629
Chamfer Distance (m) ↓2.57091.01341.0426
Edge Pred→GT (m) ↓1.58020.79760.7115
Ray-Align ↓0.37760.42450.2757
Static-Align ↑0.45030.69280.8463

100 clips, frames 0–28. All methods share the GT-derived output scoring region; MAE uses pixels valid in all three predictions (82.7671% of regional valid GT). HelloWorld uses seven-camera video input and Cosmos1 uses three cameras per frame. Equal scoring regions do not equalize input information. Static-Align uses shared reference-ICP transforms. Alignment scores are fractions. Report · Table 15 · p. 29

LiDAR VAE reconstruction · Table 16
LiDAR VAE reconstruction
MetricVAE+ Single-frame RRN+ Temporal RRN
Range MAE (m) ↓0.39810.38390.3775
Chamfer Distance (m) ↓0.39780.38150.3695
Edge Pred→GT (m) ↓0.29970.23310.2106
Ray-Align ↓0.29670.26100.2589
Static-Align ↑0.82150.85100.8974
Warp Chamfer (m) ↓0.46680.44220.3668

100 clips (2,900 frames). Reconstruction of recorded LiDAR, not RGB-conditioned generation. Range MAE uses common-valid returns; temporal metrics use recorded poses. Alignment scores are fractions. Report · Table 16 · p. 29

Few-Step Distillation

Compare the 20-step teacher on the left with the 4-step DMD student on the right. Shared playback controls align the two sequences for visual comparison of the few-step results.

Teacher 20 steps

DMD 4 steps

0.0 / 6.1 s

Left turn · Teacher 20 steps (left) / DMD 4 steps (right) · 7 views each · 10 fps · 6.1 s. Synchronized playback.

Distillation · pose control · Table 11
Distillation · pose control
SettingRot (°) ↓Trans ↓Trans* (m) ↓
Baseline (w/ ff)2.0271.3524.904
Baseline (w/o ff)3.3341.8448.485
dCM (w/ ff)2.5681.2632.862
dCM (w/o ff)4.6192.9075.311
DMD (w/ ff)1.6610.6881.885
DMD (w/o ff)2.3111.0343.193

Baseline is the RF teacher; dCM and DMD are distilled students. w/ ff and w/o ff denote first-frame conditioning. Evaluation uses a frozen 100-clip set of 100-frame videos, separate from the displayed demos. Report · Table 11 · p. 27

Distillation · cross-camera consistency · Table 12
Distillation · cross-camera consistency
SettingCSE (px) ↓ΔCSE (px) ↓
GT video1.7300.000
Baseline (w/ ff)4.180+2.450
Baseline (w/o ff)4.892+3.162
dCM (w/ ff)5.201+3.471
dCM (w/o ff)6.399+4.669
DMD (w/ ff)3.510+1.780
DMD (w/o ff)4.223+2.493

CSE uses all seven cameras. ΔCSE is relative to GT video (1.730 px); lower is better. Report · Table 12 · p. 27

Distillation · perceptual quality (VBench) · Table 13
Distillation · perceptual quality (VBench)
SettingSubjectBackgroundMotionDynamicAestheticImagingMean
Baseline (w/ ff)0.90890.94620.98760.95000.41380.44560.7754
Baseline (w/o ff)0.91110.94720.97880.97000.42590.50740.7901
dCM (w/ ff)0.89610.94090.98550.88000.41760.42810.7580
dCM (w/o ff)0.90510.94690.98660.76000.40860.38240.7316
DMD (w/ ff)0.90290.94300.98290.99000.42100.50390.7906
DMD (w/o ff)0.92090.94880.98261.00000.41710.49350.7938

VBench is evaluated on the front-wide camera. Mean summarizes six raw scores, including dynamic degree; perceptual improvements vary by metric and initialization. Report · Table 13 · p. 27

Distillation · inference cost · Table 14
Distillation · inference cost
PhaseTeacherDMD studentT/S
Text encode (s)2.52.51.0×
VAE encode RGB (s)6.77.10.95×
VAE encode control (s)13.613.61.00×
VAE decode (s)11.911.91.00×
DiT prefix (s)1.11.70.68×
DiT ODE (s)305.330.89.91×
DiT KV refresh (s)12.16.02.02×
Cache flush (s)3.72.3–
Sampling wall (s)357.777.84.60×
Video write (s)68.164.5–
End-to-end (s)426.7143.22.98×
Peak allocated (GiB)61.661.9–
nvidia-smi peak (GiB)73.973.3–

Teacher: 20 steps, CFG=3; DMD: 4 steps, CFG=1. Measured on eleven 720p clips (1280 × 720, 61 frames, 10 fps, with first-frame conditioning) using 8 GPUs and context-parallel size 8; the first warmup clip is excluded from the means. Reported speedups: DiT ODE 9.91×, sampling 4.60×, end-to-end 2.98×. T/S is teacher/student; ratios follow the report, including rounding. These are inference costs, not browser playback latency. Report · Table 14 · p. 28

Environment Control

Weather and lighting conditions guide the appearance of the generated driving scene. Compare seven-view results under rain, snow, fog, and different times of day.

Special-Scenario Generation

Targeted scene descriptions guide the generation of scenarios such as road construction, pedestrian crossings, and dense traffic. Each example presents the scene across seven camera views.

A road construction zone with excavators, exposed ground, and red-and-white barriers around the driving path.

Long-Horizon Generation

Autoregressive generation extends driving sequences using previously generated context. These examples allow inspection of scene consistency and motion continuity over longer rollouts.

Following traffic along an urban road, with changing roadside scenery and vehicles ahead.

nuScenes driving generation · long horizon · Table 10
nuScenes driving generation · long horizon
MethodFID ↓FVD ↓
MagicDrive-V220.9194.84
HorizonDrive13.8292.99
HelloWorld (Ours)17.8188.10

All long-horizon rows from Table 10. Short- and long-horizon scores use separate evaluations and should not be directly ranked across horizons. These scores are not measured on the individual demos above. Report · Table 10 · p. 21