OVER logo
The Last Scarce Asset

The Last Scarce Asset

2026-08-03

Robot foundation models train on orders of magnitude less data and compute than frontier LLMs. If they scale the way language models did, demand for real-world data grows by a similar factor — and that data has to be captured rather than scraped. This essay argues that decentralized capture networks are the supply model best matched to that demand, describes the 3D dataset OVER has built with one, and explains why falling costs for compute and research capability let a data owner do more than license.

~3,600×
Training-compute gap: frontier LLM (Grok 4) vs the best-disclosed robot foundation model (GR00T N1.7, 2026)
3–6 OOM
Gap between internet-scale data and all public robot data (2025 survey)
2% → ⅓
Share of the world’s roads mapped by Hivemapper’s DePIN network, 2023 → today
≥ 16 h
Autonomous-task horizon of the strongest 2026 model, beyond METR’s measurable range

Section 01

Demand: the scaling run hasn’t happened yet

The state of the art in physical AI trains on a small fraction of frontier-LLM compute. OpenVLA, still a reference open vision-language-action model, used 64 GPUs for two weeks; Grok 3 pre-trained on more than 100,000. The gap between a frontier LLM run and a state-of-the-art robot foundation model is three to four orders of magnitude.

Training compute — frontier LLMs vs physical-AI models
H100-eq GPU-hours, log scale

Log-scale bar chart of training compute in H100-equivalent GPU-hours: physical-AI models from VGGT (~7K) to GR00T N1.7 (~69K) versus frontier LLMs from GPT-5 (~20M) to Grok 4 (246M) — a roughly 3,600× gap between GR00T N1.7 and Grok 4.

All bars normalized to H100-equivalent GPU-hours (A100 ×0.5, H200 ×1.15, GB200 ×2.25 — BF16 training throughput; conversions ±30–50%). VGGT: 64 A100 × 9 days (paper). OpenVLA: 64 A100 × 14 days (paper). MapAnything: 64 H200 × 10 days, two-stage (paper). GR00T N1.7: 30,720 GB200-hours (NVIDIA model-card energy disclosure, Apr 2026). Llama 3.1 405B: ~31M GPU-hours (Meta model card). GPT-5: derived from Epoch AI’s ~3×10²⁵-FLOP pretraining estimate. Grok 4: ~246M H100-hours (Epoch AI estimate). Omitted for non-disclosure: π0.7, Gemini Robotics ER 1.6, Figure Helix 02, GR-3, RDT-2 — most frontier robotics labs publish no training compute. Cited, not charted: NVIDIA’s Cosmos program (Jan 2025) reported 10,000 H100s × 3 months ≈ 21.6M H100-h — a platform-level figure spanning the full 8-model WFM family plus tokenizers (text conditioning via a frozen pre-trained T5 encoder), not a single model; even that entire program ≈ GPT-4-scale FLOP, ~11× below Grok 4’s single run. Sources: model papers, NVIDIA model card, Epoch AI.[10–13, 36]

Two constraints keep these models small. The architectures are still maturing, and — the more binding limit — there isn’t enough data to train anything bigger. OpenVLA cycles through its entire dataset 27 times during training, where frontier LLMs see most of their data once. With small datasets, models stay small, and small models need little compute. By LLM standards, the field is still early.

The conditions for a scaling run are forming. DeepMind ships Gemini Robotics; NVIDIA has named robotics its next major market and presented a “Physical AI Data Factory” blueprint at GTC 2026; OpenAI is rebuilding its robotics program; robotics startups have raised $18.8B so far in 2026, more than in all of 2025. If architectures mature enough to absorb LLM-scale compute, balanced scaling implies data demand rises by roughly the same orders of magnitude: closing even half the current gap would require real-world data at 100–1,000× today’s entire public supply. Labs are already buying — XDOF launched with $70M and around 20 customers including frontier labs, General Intuition raised at a $2.3B valuation, and Tesla keeps more than 1,000 Optimus units in-house primarily to collect data.

Section 02

Supply: why physical data is scarce

Language models benefited from an unusual inheritance: thirty years of human activity already transcribed onto scrapeable servers. The indexed web holds roughly 500 trillion tokens, and it cost model builders nothing to accumulate. Nothing comparable exists for the physical world. The geometry of a street, or the way a door handle responds to force, was never uploaded anywhere; it has to be captured, sensor by sensor and place by place, at real marginal cost.

The data gap — internet vs the physical world
hours of experience, log scale

Log-scale bar chart of dataset sizes in hours of experience: public robot and 3D datasets from 350 hours (DROID) to about 100,000 hours (all public robot data), versus web-scale corpora up to roughly 25 billion hours of indexed web text — about six orders of magnitude apart.

Hours-native datasets at reported values: DROID 350 h; nuPlan 1,500 h (largest public driving dataset); AgiBot World 2,976 h; Ego4D 3,670 h; π0 corpus 10,000+ h; Panda-70M 167,000 h (video-generation corpus; larger curated corpora exist — e.g. InternVid ~760K h — all still ~10³× below one year of YouTube uploads). “3D-vision FM training mix” = the video/RGB-D portion of VGGT/MapAnything-class training data, derived from published frame counts (ScanNet 23 h + ScanNet++ ~25 h + DL3DV ~240–470 h + CO3D ~100–150 h + Aria Digital Twin 6.6 h + TartanAir ~9 h ≈ 400–700 h; plotted at ~600, upper bound ≲10³); excludes image-only and synthetic components (MegaDepth, Mapillary MPSD ~750K images, Kubric, Habitat), which have no natural hours measure. “All public robot data” ≈ 15M pooled episodes (2025 survey) ≈ 10⁵ h (derived at ~20–25 s/episode). YouTube: 500 h uploaded/minute ≈ 263M h/yr. Web text: ~500T tokens at 250 wpm ≈ 2.5×10¹⁰ reading-hours (derived, illustrative). The datasets that built 2D vision are measured in images, not hours — ImageNet 14M images, LAION-5B 5.85B pairs, SA-1B 1.1B masks — and are therefore not charted. Sources: dataset papers; Epoch AI; Statista.[15–22, 37]

The largest humanoid dataset assembled so far — AgiBot World, collected by 100 robots in a purpose-built Shanghai facility — totals 2,976 hours. YouTube receives that much video every six minutes. Web video is also not a substitute: Meta’s V-JEPA 2 pre-trained on a million hours of it and still needed real robot data for control, because video lacks metric scale, 3D geometry, and action labels.

The chart also leaves out the datasets that built 2D vision — ImageNet’s 14 million images, LAION’s 5.85 billion image-text pairs — because they’re measured in images rather than hours of experience. Meanwhile, the corpora behind today’s 3D-vision foundation models (VGGT, MapAnything) amount to under roughly 1,000 hours-equivalent of real capture.

One empirical result matters for what kind of data to collect. The most careful scaling study in robot learning found that policy generalization follows a power law in the diversity of environments and objects, not in the number of demonstrations collected in any one place. The binding constraint is coverage — many places, many objects, many conditions, refreshed over time — more than raw hours. The sections above cover demand; the rest of the essay is about supply.

Section 03

Four ways to capture reality

If demand develops as argued above, the practical question is how to capture physical data at scale. Four supply models exist today, with very different cost curves.

Four supply models
Supply model Who bears capex Marginal cost of a new place Environmental diversity Freshness Examples
Corporate fleet The company — cars, sensors, drivers High — a vehicle trip per location Follows corporate priorities; rich-world bias ~1 pass every 1–2 years Google Street View: 2007–2019 to reach 16.1M unique km
Robot fleet The company — billions in hardware Very high — a deployed robot Single embodiment; facility- and task-bound Continuous, but narrow Tesla keeps 1,000+ Optimus units in-house primarily for data
Teleop / data farm The company — facilities & operators Linear — $136–340 per operator-hour Limited to staged facilities On demand, at cost XDOF ($70M launch, ~20 lab customers), AgiBot’s 4,000 m² plant
DePIN network Contributors — capex moves to the edge Near zero — token emission + validation By construction — supply appears wherever people live Steerable — rewards boosted where data is stale Helium, Hivemapper, OVER

The first three models share the same limitation: cost scales with coverage. Every new neighborhood is another vehicle trip, another robot, another operator-hour. But coverage — environmental diversity — is exactly what the scaling results reward. Only the fourth model avoids this tension.

Section 04

Why DePIN’s cost curve is different

Decentralized Physical Infrastructure Networks approach capture with an incentive instrument instead of a fleet: a token that converts the network’s future demand into present supply. Contributors buy their own sensors, capture where they live, and are paid in upside tied to the network; the company’s marginal cost for a new location drops to validation and storage. Capex doesn’t disappear — it moves to the edge, spread across thousands of participants who each hold a stake in what they’re building.

There is now a track record to evaluate. Messari counts over 650 DePIN projects and 8.8 million active devices, with a growing share of the sector generating real revenue. Two precedents are most relevant. Helium bootstrapped a global wireless network to roughly one million IoT hotspots — infrastructure that would have cost a telecom billions in capex — by paying contributors in tokens to run radios at home. Hivemapper is the closest precedent for mapping: its dashcam network went from 2% of the world’s roads in March 2023 to 20% within eighteen months, and to roughly a third today. The company reports coverage growth about five times faster than Google Street View achieved with corporate fleets, with over 100 repeat passes where Street View averages one every year or two. These figures are company-reported, but even discounted they suggest decentralized capture can outpace fleets on coverage and freshness.

Token incentives reach places no fleet is ever dispatched to.

Three properties make DePIN a good fit for physical-AI data. First, diversity by construction: supply appears wherever contributors live, which matches the environment-diversity pattern the scaling results reward. Second, steerable supply: token rewards act as a dial, so a demand signal from model training — more indoor retail, more rain, more night scans — becomes a reward boost, and the network re-targets within days. Third, freshness: contributors re-scan because earning is continuous, so the corpus becomes a time series of places rather than a one-pass archive.

The main caveat: DePIN’s known failure mode is rewarding volume over quality. The mitigations are structural — validation before rewards vest, demand-side revenue that consumes supply, reward curves tuned to what training actually uses. The networks that survived the 2022–24 shakeout are the ones that built this discipline, and the sector’s compression from 1,000× to 10–25× revenue multiples suggests the market now prices fundamentals.

Section 05

What OVER has built

OVER has been applying this model to 3D mapping since well before the current physical-AI cycle, focused on metric-scale 3D reconstructions of real places rather than street video. Through Map2Earn, anyone with a smartphone can scan a location; the capture is validated, reconstructed into photogrammetry-grade 3D, and registered to the map, and the mapper earns OVR tokens. OVRLand — ownership of 300 m² spatial domains with publishing rights and revenue share — gives contributors a lasting stake in the network rather than piecework wages. The result is, to our knowledge, the largest crowdsourced 3D map:

275K+
locations mapped, across every continent
1,236 TB
of real-world 3D data — geometry, imagery, poses
105M+
spatially registered images
273K+
locations live on OVER’s Visual Positioning System

What matters about this corpus is less its size than its labels. Every capture is a 3D reconstruction with metric scale, camera poses, and spatial registration — the supervision signals that world models, VLA navigation stacks, and sim-to-real pipelines lack in web video (the V-JEPA 2 result above). The same asset serves three kinds of demand: training data for world and geospatial models; digital twins that seed simulation pipelines such as NVIDIA’s data factories; and a live localization service — VPS with centimeter-grade, TEE-encrypted pose estimation — already consumed by robots as an API.

01 · INCENTIVE

Token rewards

OVR emissions and OVRLand stakes recruit mappers wherever they live.

02 · CAPTURE

The world, scanned

Smartphone photogrammetry, validated into metric-scale 3D, without fleets or facilities.

03 · ASSET

Living 3D map

275K+ locations, re-scanned over time — a time series of real places.

04 · MODELS

LGMs & VPS

Large Geospatial Models and localization trained on the proprietary corpus.

05 · REVENUE

APIs & data sales

Robotics, XR and enterprise demand pays the network, funding richer rewards.

The cost structure follows from the loop. Adding a location doesn’t require dispatching hardware; it requires adjusting a reward curve, because contributors already own the sensors. This is also why coverage is geographically scattered — Toronto, Bangkok, Charleroi, Nghĩa Trụ, Atlanta and Shenzhen all appear in a single month of scans — rather than following a deployment plan.

What licensing alone is worth

Suppose OVER only ever licenses this data. There are recent comparables for what markets pay when data is the bottleneck: Meta paid $14.3B for 49% of Scale AI (implying ~$29B); Surge, bootstrapped to $1.2B in revenue, has reportedly negotiated at $25–30B; Mercor went from $2B to $10B in eight months and is reportedly in talks at $20B. Those valuations are for companies that refine text scraped for free; physical-AI data is proprietary from the moment it’s captured. Licensing at anywhere near those comparables would be a solid business by itself. We think there’s additional upside, because of what happened to the other two inputs over the last year.

Section 06

Research capability is becoming rentable

Until about a year ago, this essay would have ended at licensing. Moving up the value chain required two things a data company couldn’t easily get: a frontier research team and a nine-figure compute budget. Compute went first: an H100-hour that cost over $8 at the 2023 peak rents for $1.50–2.50 today, and Stargate, Colossus, and Meta’s gigawatt campuses are pushing supply further. The research constraint is now easing too, and the change is measurable.

METR tracks the length of software and ML-research tasks (measured in expert-human time) that an AI can complete autonomously at 50% reliability — close to the actual work of training a foundation model. That horizon doubled roughly every seven months from 2019 to 2025; since January 2024, it has doubled about every 105 days.

METR 50% time horizon
expert-human task length an AI completes at 50% reliability · log scale

Scatter plot of METR 50% time horizons by model release date on a log scale, from GPT-4 at 3.5 minutes in 2023 to Claude Mythos Preview at 16 or more hours in 2026, with a roughly 110-day doubling trend projected to cross a 40-hour work-week around October 2026.

Solid points: METR Time Horizon 1.1 measurements (228 software/ML/cyber tasks). Hollow: ranges reported on METR’s live dashboard. Mythos Preview: ≥16 h (95% CI 8.5–55 h); METR notes measurements above 16 h are unreliable with the current suite. Trend fitted on 2024+ measured points (~110-day doubling); dashed segment is a naive extrapolation, not a METR forecast — it crosses a 40-hour work-week ≈ Oct 2026. Sources: METR (May 2026); Kwa et al. 2025.[1–3]
Same data, linear scale
hours

The same METR time-horizon measurements on a linear scale, showing the sharp rise from minutes to 16 or more hours between 2023 and 2026.

The same measurements on a linear axis. The Claude Mythos Preview confidence interval extends to 55 h, beyond the top of the chart.

In May 2026, METR added Claude Mythos Preview — the model class behind Claude Fable 5 — and reported a horizon at or beyond 16 hours, past the range the current benchmark can measure reliably. Concrete examples of what this looks like in practice: Anthropic reports Mythos 5 running a week of largely autonomous research and producing an ML model that outperformed a recent Science-published model at a hundredth of the size; DeepMind’s AlphaEvolve improved on a 56-year-old matrix-multiplication result and recovers around 0.7% of Google’s global compute; on METR’s RE-Bench, AI agents score four times human experts at two-hour research budgets. On the open-weights side, Moonshot’s Kimi K3 — 2.8 trillion parameters, third on GDPval behind Claude Fable 5 and GPT-5.6 — rents at $15 per million output tokens, with the open-closed gap around 3–4 months and capability prices falling roughly 10× per year.

The specialized capability needed to train a state-of-the-art model — the scarcest input of the last decade — increasingly has a market price. To be clear, the claim is about market structure, not about human researchers becoming obsolete: the engineering capability required to train a competitive physical-AI model is ceasing to be a differentiator, because every serious team works with frontier research agents and training recipes diffuse through open weights within months. When an input is available to everyone at a price, it stops being a moat. The practical barrier between data vendor and AI lab — people and capex — is much lower than it was.

Section 07 · Conclusion

Three tiers of value

Taken together, the argument prices OVER’s position in three tiers. The demand and supply case justifies the first two on its own; the fall in research and compute costs opens the third — and the third is where foundation-model economics live.

Tier 1 · TodayLICENSING

License the data

Raw and curated 3D data, sold to labs facing the demand growth described above. Text-era comparables — Scale at ~$29B, Surge in talks at $25–30B, Mercor in talks at $20B — were built on refining freely scraped data; OVER’s data is proprietary from capture.

Tier 2 · LiveSERVICES

APIs on the same asset

VPS localization (centimeter-grade, TEE-encrypted, live on 273K+ locations), robotics endpoints, digital-twin feeds for simulation pipelines. Recurring, per-call revenue on the same asset.

Tier 3 · Opening nowMODELS

Train models on the corpus

Large Geospatial Models and world models trained on the proprietary corpus, with compute rented by the hour and research capability rented by the token. Labs without proprietary data compete with open weights on one side and data owners on the other; data owners who also train models keep more of the margin. OVER’s LGM program is already running.

ENABLED BY THE FALL IN RESEARCH AND COMPUTE COSTS — SECTION 06

Of a physical-AI model’s three inputs — compute, research capability, and data — the first two can now be rented. Diverse, metric-scale data of real places cannot: it has to be accumulated over years, and that is the asset OVER holds.

OVER Research · info@ovr.ai

Section 08

What could break this

Objection 01Synthetic data substitutes for capture

Simulation engines are themselves trained on real data — Cosmos consumed 20 million hours of real video — and NVIDIA’s own data-factory blueprint pairs synthetic generation with real seeds, because simulation needs anchoring to limit the sim-to-real gap. On current evidence, synthetic data multiplies ground truth rather than replacing it, and digital twins of real places are its input.

Objection 02Web video is enough

V-JEPA 2 tested this at million-hour scale: representations were strong, but real robot data was still needed for control, and performance degraded off-distribution. Video without geometry, scale, and action labels captures how the world looks, but not how it responds to action. The scarce — and therefore priced — layer is the grounded one.

Objection 03Labs vertically integrate their own fleets

They are, which itself suggests data is the scarce input. But fleet data is single-embodiment and facility-bound, while generalization depends on environmental diversity — which is why the same labs also buy from external providers. Fleets and diverse-capture networks complement each other, and external networks are the part labs can buy.

Objection 04Timing, and DePIN incentive quality

Timing is the main residual risk. Two mitigants: data demand is already repricing, even at today’s small compute scales, and capture networks take years to build — a head start like Hivemapper’s three years is hard for a later entrant to close. On incentive quality, the failure mode is rewarding volume over quality; the mitigations are validation before rewards vest, demand-side burn, and reward curves tuned to what training actually consumes.


Appendix

Sources

  1. METR — Task-Completion Time Horizons of Frontier AI Models (updated May 8, 2026). metr.org/time-horizons
  2. METR — Time Horizon 1.1 (Jan 29, 2026); Kwa et al., arXiv:2503.14499. metr.org
  3. OfficeChai — Claude Mythos shows 50% time horizon of 16+ hours on METR (May 9, 2026). officechai.com
  4. Anthropic — Introducing Claude Fable 5 & Mythos 5 (Jun 2026). anthropic.com
  5. Google DeepMind — AlphaEvolve (May 2025). deepmind.google
  6. METR — RE-Bench (2024). arXiv:2411.15114
  7. MarkTechPost — Moonshot AI releases Kimi K3 (Jul 16, 2026). marktechpost.com
  8. Artificial Analysis — GDPval-AA v2 snapshots (Jul 16–17, 2026). artificialanalysis.ai
  9. Epoch AI — Open vs closed frontier gap; LLM inference price trends; a16z “LLMflation”. epoch.ai · a16z.com
  10. Epoch AI — Grok 4 training resources (~246M H100-hours, Sep 2025); models over 1e25 FLOP. epoch.ai
  11. Meta — Llama 3.1: 16,384 H100s; ~31M GPU-hours (Jul 2024). ai.meta.com
  12. NVIDIA — GR00T N1 (~50,000 H100-hours, 2025). arXiv:2503.14734
  13. Kim et al. — OpenVLA: 21,500 A100-hours, 64 GPUs, 27 epochs (2024). arXiv:2406.09246
  14. NVIDIA — Cosmos: 20M hours of video, 10,000 H100s × 3 months (Jan 2025). arXiv:2501.03575
  15. Epoch AI — Can AI scaling continue through 2030? (~500T indexed tokens). epoch.ai
  16. HuggingFace — FineWeb (15T tokens); Villalobos et al., arXiv:2211.04325. huggingface.co
  17. Meta — V-JEPA 2 (Jun 2025). arXiv:2506.09985
  18. Open X-Embodiment (2023). arXiv:2310.08864
  19. AgiBot World — 1M trajectories / 2,976 h / 100 robots (2025). arXiv:2503.06669
  20. DROID (2024), arXiv:2403.12945; RT-1 (2022), arXiv:2212.06817; Ego4D, ego4d-data.org
  21. Statista — 500+ hours uploaded to YouTube per minute. statista.com
  22. Survey — Training Data for Generalist Robot Foundation Models (2025): ~15M open episodes, 3–6 OOM below language/vision. researchgate.net
  23. Lin et al. — Data Scaling Laws in Imitation Learning (CoRL 2024). arXiv:2410.18647
  24. Latent.Space — “$2 H100s” (2024); IntuitionLabs H100 rental comparison (2026). latent.space · intuitionlabs.ai
  25. OpenAI — Stargate ($500B / 10 GW, Jan 2025); xAI Colossus; Meta Hyperion. openai.com
  26. NVIDIA — GTC 2026: Physical AI Data Factory blueprint (Mar 2026); Crunchbase — robotics VC $18.8B in 2026. blogs.nvidia.com · crunchbase.com
  27. Messari — State of DePIN 2025: 650+ projects, 8.8M active devices, $72M FY25 on-chain revenue, 10–25× revenue multiples; Helium ≈1M IoT hotspots (GridEcon/Solana case studies). messari.io · gridecon.com
  28. TechCrunch — Hivemapper at 2% of global roads / 1M unique km (Mar 2023); Google Street View 16.1M unique km, 2007–2019. techcrunch.com
  29. Hivemapper / Bee Maps — 20% of global roads in 18 months; ~⅓ today; 5× Street View pace; 100+ repeat passes (company-reported). beemaps.com · hivemapper.com
  30. TechCrunch — XDOF launches with $70M, ~20 customers incl. frontier labs (Jun 2026); General Intuition at $2.3B (Jun 2026); SVRC — teleop $136–340/hr (directional). techcrunch.com
  31. Tesla Q4-2025 earnings — 1,000+ Optimus units used primarily for data collection (Jan 2026). finance.yahoo.com
  32. OVER — live platform statistics; Map2Earn; OVRLand; VPS & TEE; LGM & robotics APIs (accessed Jul 17, 2026). overthereality.ai · HF dataset
  33. Forbes — Meta invests $14.3B in Scale AI at ~$29B (Jun 2025). forbes.com
  34. Bloomberg — Surge AI in talks at $25B+; $1.2B 2024 revenue (Jul 2025). bloomberg.com
  35. TechCrunch — Mercor: $10B (Oct 2025); talks at $20B (Jul 2026). techcrunch.com
  36. 2026 compute disclosures — VGGT: 64 A100 × 9 days (arXiv:2503.11651); MapAnything: 64 H200 × 10 days (arXiv:2509.13414); NVIDIA GR00T N1.7 model card, 30,720 GB200-hours (huggingface.co); Epoch AI — GPT-5 training-compute notes (epochai.substack.com).
  37. Dataset scales — nuPlan: 1,500 h (arXiv:2106.11810); Panda-70M: 70.7M clips / 167K h (Snap Research); 3D-FM mix components: ScanNet (arXiv:1702.04405), DL3DV-10K (arXiv:2312.16256), CO3D (arXiv:2109.00512), Aria Digital Twin (arXiv:2306.06362), per the VGGT/MapAnything training-data lists; ImageNet (image-net.org); LAION-5B (laion.ai); InternVid ~760K h (2023).

OVER RESEARCH · JULY 2026 · v3 · Informational only; not investment advice or an offer of securities. Figures as of July 17, 2026; derived figures are marked in chart notes; deals described as “talks” were unclosed at publication; Hivemapper coverage figures beyond March 2023 are company-reported. Time-horizon projections are naive extrapolations of METR-measured trends, not METR forecasts.

© 2026 OVER · overthereality.ai · info@ovr.ai