OVER logo
Introducing OverMaps360-50: real-world 360° scenes, ready to simulate

blog

Introducing OverMaps360-50: real-world 360° scenes, ready to simulate

What we released

Today we are releasing OverMaps360-50, a free research dataset of 50 real-world 360° walk-throughs, each delivered as a physics-certified simulation scene for MuJoCo and Isaac Sim. It weighs 1.13 TB, lives on Hugging Face under CC BY-NC 4.0, and is our second open dataset after OverMaps-1K.

Every scene starts as a person walking a street, a park, a temple or a mall with an Insta360 X5 on a pole. It ends as a metric COLMAP reconstruction, a 3D Gaussian Splatting model, a textured mesh with a closed floor, and a simulation bundle that has passed eight physics checks. Download it from Hugging Face and explore the samples on the project page.

What’s inside

Fifty scenes, 1.13 TB, and every modality from the raw 360° recording to the simulation bundle.

Measure Value
Scenes 50 walk-throughs in 6 countries (Thailand, Latvia, United States, Indonesia, Spain, Mexico)
Fisheye frames 118,816 at 3840×3840 per lens, median 2,404 per scene
Perspective views 594,080 pinhole views at 1600×1600, 90° FOV, 5 per fisheye frame
Smartphone photos 5,591 ARKit-tracked photos with metric poses, median 72 per scene
IMU 29.8 M gyroscope and accelerometer samples at 1 kHz, synchronised to the video
Walked distance 16.3 km, median 316 m per scene
Certified walkable area 180,000 m², median 3,548 m² per scene
3D Gaussian Splatting 250 M Gaussians, 5 M per scene
Dense points 381 M, median 7.4 M per scene
Volume 1.13 TB, median 22 GB per scene

Each scene is a self-contained folder, so you can download one and have everything for it:

  • the original Insta360 X5 recording with people and license plates blacked out, plus its 1 kHz IMU and a 1 Hz GPS track
  • fisheye frames for both lenses, with per-frame privacy masks
  • a metric COLMAP reconstruction with five perspective views per fisheye frame, and the smartphone photos registered in the same model
  • a dense point cloud and a textured mesh (glTF and USDZ) whose floor is closed
  • a 3DGS model in .ply and .splat, and three collision-mesh LODs at 1, 3 and 10 cm error
  • a MuJoCo scene, a USD scene for Isaac Sim, a walkable navmesh, and the certification reports
  • a caption plus scene type, weather, lighting, time of day and crowd density

Every asset in a scene shares one metric frame, scaled from the ARKit trajectory of the phone.

From a walk to a simulation scene

Each scene is recorded by one person walking through the environment with two devices. An Insta360 X5 sits on a pole about 2.2 m above the ground, recording dual-fisheye 360° video at 3840×3840 per lens with its 1 kHz IMU embedded in the file. A smartphone running the OverTheReality app tracks the same walk with ARKit visual-inertial odometry, takes tracked photos along the way, and provides the metric scale, GPS at 1 Hz and a compass heading.

Contributors are paid through map2earn only after the capture passes our validation pipeline. That single rule keeps the protocol tight and steers people toward geometrically interesting scenes rather than low-effort captures.

From there, every scene goes through the same automated pipeline:

  1. Frame extraction at about 2 fps, one JPG per lens.
  2. Privacy masking: people and license plates are segmented on both lenses and on the smartphone photos. The masks are applied to everything we publish, including the 360° video, and masked pixels are blacked out rather than inpainted.
  3. Matching and structure from motion: learned image matching over sequential, cross-lens and loop-closure pairs, then a rig-aware SfM with a fisheye model per lens and a fixed two-lens rig. The smartphone photos are registered in the same reconstruction.
  4. Metric alignment: the reconstruction is scaled and gravity-aligned to the ARKit trajectory, which we validated against ground-truth markers.
  5. Perspective views: each fisheye frame is re-projected into five pinhole views (front, left, right, up, down) with the masks carried over. This is what the dense reconstruction and the 3DGS training consume.
  6. Dense reconstruction and mesh: multi-view stereo on the perspective views, then a textured mesh whose floor is closed with a monocular geometry prior, so agents never fall through the gaps between the walked corridor and the walls.
  7. 3D Gaussian Splatting: 5 M Gaussians per scene, trained on the perspective views and initialised from the dense cloud.
  8. Simulation bundle: the mesh becomes a MuJoCo scene with a height-field ground and convex collision parts, a USD scene for Isaac Sim, a navmesh on 0.25 m cells and three mesh LODs. Then the scene is gated, as described next.
  9. Annotation: a vision-language model writes the caption and assigns scene type, weather, lighting, time of day and crowd density.

Certified, not just “simulation-ready”

“Simulation-ready” usually means a mesh that loads. Here it means every scene passed eight physics checks in MuJoCo and a drop test in Isaac Sim before publication, and all 50 scenes reached tier T1a.

The MuJoCo gate runs on the actual simulation bundle, not on the source mesh:

# Check Threshold
1 Up direction, verified against the splat and the camera heights ≥ 90% of cameras have ground below them
2 Scene is plumb (walls vertical) tilt < 2°
3 Floor registration: navmesh vs walked floor, coverage of the walk median < 5 cm, ≥ 20% of cameras on walkable cells
4 Alignment residual between splat and collision geometry < 25 cm
5 Walkable area ≥ 5 m²
6 Probes falling through the collision geometry 0%
7 Probes settled within 3 s ≥ 90%
8 Contact penetration, p95 < 20 mm

The Isaac Sim report is a supporting drop test: a rigid sphere released on a certified floor point, run headless with PhysX. Across the 50 scenes the measured values sit well inside the thresholds: floor registration median 0.1 cm, plumb tilt median under 0.5°, walkable area median 3,548 m², and zero leaked probes. The per-scene numbers are in the JSON reports and in each scene’s datasheet.

All of this runs on ovr-maps-360-simkit, our open-source toolkit that builds the simulation bundle, the LODs and the datasheet from a scene, then gates it. Every report we ship can be reproduced with simkit regate simulation/<id8>.sre, so you do not have to take our word for the certification.

Why we are releasing it

Spatial AI is short of real-world data that is metric, multi-modal and captured on purpose, and most of what exists sits inside a handful of large labs. We built OverMaps360-50 so that anyone working on 3D reconstruction, neural rendering, embodied AI or robot navigation can train and benchmark on real places without a capture team of their own.

The dataset is released under Creative Commons Attribution-NonCommercial 4.0: research use is allowed and encouraged, commercial use is not. Please attribute OverTheReality and link back to the repository when you share derivatives. The certification toolkit, ovr-maps-360-simkit, is open source.

If you publish with it, train on it or break it in an interesting way, tell us. We want to feature what the community builds on top of it.

Our second open dataset

OverMaps360-50 is the 360° companion of OverMaps-1K, which we released in 2025: 1,000 real-world scenes captured with smartphones through the same map2earn protocol, over 580,000 images, with COLMAP reconstructions, metric poses, LiDAR depth where the device had it, multi-view depth maps and 3DGS models. OverMaps-1K is itself a sample of a 155,000-scene collection, and it has passed 4,000 downloads on Hugging Face.

Where OverMaps-1K is built from photos taken around an anchor, every OverMaps360-50 scene is a continuous walk, and every scene ends in a physics-certified simulation bundle. The two share the same acquisition network, the same privacy pipeline and the same annotation fields (caption, scene type, weather, lighting, crowd density), so they can be used together as two views of how real places look and behave.

A collection that keeps growing

The 50 scenes are a snapshot. The full OverMaps-360Cam collection stands at 2,712 scenes across 29 countries and 125 cities, more than 500 hours of capture and 2,200 km of walked paths, about 60 TB in total. It grows continuously: every walk that passes validation is added to it, and its contributor is paid through map2earn. The dataset is growing with thousand of new scenes monthly.

That is the DePIN part of the story. The network of mappers is the capture team, the incentive is tied to validated output rather than to raw uploads, and the collection expands wherever people are willing to walk with a camera on a pole.

The full collection is available under a commercial license. Write to data@ovr.ai or request access from the project page.

Get started

One scene is about 22 GB. Download it with the Hugging Face CLI, or skip the image archives to get only the light assets (mesh, splat, IMU, poses, simulation bundle) for every scene:

pip install -U "huggingface_hub[cli]"

# One scene (about 22 GB)
hf download OverTheReality/OverMaps360_50 \
    --include "0313afa5-8b3c-4fb1-8245-c6b5abf738d2/*" \
    --repo-type dataset --local-dir ./OverMaps360-50

# Light assets of every scene, no image archives
hf download OverTheReality/OverMaps360_50 \
    --exclude "*.tar" \
    --repo-type dataset --local-dir ./OverMaps360-50

Then step a real street in MuJoCo:

import mujoco

scene = "OverMaps360-50/0313afa5-8b3c-4fb1-8245-c6b5abf738d2"
model = mujoco.MjModel.from_xml_path(f"{scene}/simulation/0313afa5.sre/scene.xml")
data = mujoco.MjData(model)
mujoco.mj_step(model, data)

Links:

If you use OverMaps360-50 in your research, please cite it:

@misc{OverMaps360_50,
  author = {OverTheReality},
  title = {{OverMaps360-50 Dataset}},
  howpublished = {Hugging Face Datasets},
  url = {https://huggingface.co/datasets/OverTheReality/OverMaps360_50},
  year = {2026},
}