OVER logo
Physical AI · Data Networks · DePIN

The Last Scarce Asset

Real-world data is Physical AI's last scarce asset — and its owner no longer has to sell it raw. A thesis in three acts: a demand shock, a supply monopoly, and the twist that turns a data business into an AI lab.

~3,600×
Training-compute gap: frontier LLM (Grok 4) vs the best-disclosed robot foundation model (GR00T N1.7, 2026)10,36
3–6 OOM
Gap between internet-scale data and all public robot data (2025 survey)22
2% → ⅓
Share of the world's roads mapped by Hivemapper's DePIN network, 2023 → today28,29
≥ 16 h
Autonomous-task horizon of the strongest 2026 model — past METR's ceiling. The twist.1,3
TL;DR — the whole argument

Physical AI is approaching its scaling moment and will arrive data-starved: today's robot foundation models train with 3–4 orders of magnitude less compute than frontier LLMs because there is 3–6 orders of magnitude less data to feed them. Closing even half that gap means demand for real-world data at 100–1,000× today's entire public supply — and reality cannot be scraped, only captured. OVER owns the world's largest crowdsourced 3D map, built on the one capture architecture — DePIN — whose economics improve with the environmental diversity scaling laws reward. Licensing that asset at the multiples markets already pay data companies is the floor. The prize is bigger: elite AI research intelligence just became a commodity — autonomy on research tasks doubles every ~105 days, the strongest models have outgrown METR's 16-hour ruler, near-frontier open weights rent for $15 per million tokens — and compute rents by the hour. The wall between data vendor and AI lab has fallen. OVER doesn't have to sell the crude. It can own the refinery.

Act I · Section I

The coming demand shock

Physical AI's state of the art trains on a rounding error of frontier-LLM compute. OpenVLA — still a reference open vision-language-action model — used 64 GPUs for two weeks. Grok 3 pre-trained on 100,000+. The gap between a frontier LLM run and a SOTA robot foundation model is three to four orders of magnitude.10–14

Training compute — frontier LLMs vs physical-AI models
GPU-hours, log scale
10³10⁴10⁵10⁶10⁷10⁸H100-eq hoursVGGT: ~7K H100-eq hoursVGGT3D geometry FM · 64 A100 × 9 d~7KOpenVLA 7B: ~11K H100-eq hoursOpenVLA 7Bopen VLA · 64 A100 × 14 d~11KMapAnything: ~18K H100-eq hoursMapAnythingmetric 3D FM · 64 H200 × 10 d~18KNVIDIA GR00T N1.7: ~69K H100-eq hoursNVIDIA GR00T N1.7humanoid FM · 30.7K GB200-h~69KGPT-5: ~20M H100-eq hoursGPT-5frontier LLM · 2025 · Epoch est.~20MLlama 3.1 405B: 31M H100-eq hoursLlama 3.1 405Bfrontier LLM · 2024 · 16K H10031MGrok 4: 246M H100-eq hoursGrok 4frontier LLM · ~200K GPUs (Epoch)246M≈ 3,600× — GR00T N1.7 vs Grok 4
All bars normalized to H100-equivalent GPU-hours (A100 ×0.5, H200 ×1.15, GB200 ×2.25 — BF16 training throughput; conversions ±30–50%). VGGT: 64 A100 × 9 days (paper). OpenVLA: 64 A100 × 14 days (paper). MapAnything: 64 H200 × 10 days, two-stage (paper). GR00T N1.7: 30,720 GB200-hours (NVIDIA model-card energy disclosure, Apr 2026). Llama 3.1 405B: ~31M GPU-hours (Meta model card). GPT-5: derived from Epoch AI's ~3×10²⁵-FLOP pretraining estimate. Grok 4: ~246M H100-hours (Epoch AI estimate). Omitted for non-disclosure: π0.7, Gemini Robotics ER 1.6, Figure Helix 02, GR-3, RDT-2 — most frontier robotics labs publish no training compute. Cited, not charted: NVIDIA's Cosmos program (Jan 2025) reported 10,000 H100s × 3 months ≈ 21.6M H100-h — a platform-level figure spanning the full 8-model WFM family plus tokenizers (text conditioning via a frozen pre-trained T5 encoder), not a single model; even that entire program ≈ GPT-4-scale FLOP, ~11× below Grok 4's single run. Sources: model papers, NVIDIA model card, Epoch AI.[10–13, 36]

This is not frugality. Two ceilings force it: the architectures haven't had their Transformer moment — and, more binding, there isn't enough data to feed a bigger model. OpenVLA recycles its entire dataset 27 times during training; frontier LLMs see most of their data once.13 Small data forces small models; small models need small compute. The entire industry is, by LLM standards, warming up.

Now run the tape forward. Every ingredient of the scaling moment is being staged: DeepMind ships Gemini Robotics, NVIDIA declared robotics its next major market and launched a "Physical AI Data Factory" blueprint at GTC 2026, OpenAI is rebuilding its robotics program, and robotics startups have raised $18.8B in 2026 so far — more than in all of 2025.26 When architectures mature enough to absorb LLM-scale compute, balanced scaling drags data demand up by the same orders of magnitude. Closing even half the gap implies demand for real-world data at 100–1,000× today's entire public supply. The labs are already buying: XDOF launched with $70M and ~20 customers including frontier labs; General Intuition raised at $2.3B; Tesla keeps 1,000+ Optimus units on its own floor doing nothing but collecting data.30,31


Act I · Section II

There is no Common Crawl for reality

Language models had a miracle underneath them: humanity spent thirty years transcribing itself onto scrapeable servers. The indexed web holds ~500 trillion tokens — a frontier corpus is the equivalent of ~86,000 years of continuous reading, and it was free.15,16 The physical world granted no such favor. Nobody pre-uploaded the geometry of your street or the friction of a door handle. Physical data is not scraped; it is captured — sensor by sensor, place by place, at real marginal cost.

The data gap — internet vs the physical world
hours of experience, log scale
10²10³10⁴10⁵10⁶10⁷10⁸10⁹10¹⁰hoursDROID: 350 h hoursDROIDrobot manipulation · 13 institutions350 h3D-vision FM training mix: ≲10³ h hours3D-vision FM training mixVGGT/MapAnything mixes (derived)≲10³ hnuPlan: 1,500 h hoursnuPlanlargest public driving dataset1,500 hAgiBot World: 2,976 h hoursAgiBot Worldlargest humanoid set · 100 robots2,976 hEgo4D: 3,670 h hoursEgo4Degocentric video · 9 countries3,670 hPhysical Intelligence π0: 10K h hoursPhysical Intelligence π0entire training corpus10K hAll public robot data: ~10⁵ h hoursAll public robot data~15M episodes pooled (est.)~10⁵ hPanda-70M: 167K h hoursPanda-70Mvideo-generation training corpus167K hYouTube — one year: 263M h hoursYouTube — one yearof uploads (500 h/min)263M hIndexed web text: ~25B h hoursIndexed web text≈500T tokens (derived hours)~25B h≈ 6 orders of magnitude
Hours-native datasets at reported values: DROID 350 h; nuPlan 1,500 h (largest public driving dataset); AgiBot World 2,976 h; Ego4D 3,670 h; π0 corpus 10,000+ h; Panda-70M 167,000 h (video-generation corpus; larger curated corpora exist — e.g. InternVid ~760K h — all still ~10³× below one year of YouTube uploads). “3D-vision FM training mix” = the video/RGB-D portion of VGGT/MapAnything-class training data, derived from published frame counts (ScanNet 23 h + ScanNet++ ~25 h + DL3DV ~240–470 h + CO3D ~100–150 h + Aria Digital Twin 6.6 h + TartanAir ~9 h ≈ 400–700 h; plotted at ~600, upper bound ≲10³); excludes image-only and synthetic components (MegaDepth, Mapillary MPSD ~750K images, Kubric, Habitat), which have no natural hours measure. “All public robot data” ≈ 15M pooled episodes (2025 survey) ≈ 10⁵ h (derived at ~20–25 s/episode). YouTube: 500 h uploaded/minute ≈ 263M h/yr. Web text: ~500T tokens at 250 wpm ≈ 2.5×10¹⁰ reading-hours (derived, illustrative). The datasets that built 2D vision are measured in images, not hours — ImageNet 14M images, LAION-5B 5.85B pairs, SA-1B 1.1B masks — and are therefore not charted. Sources: dataset papers; Epoch AI; Statista.[15–22, 37]

The largest humanoid dataset ever assembled — AgiBot World, 100 robots in a purpose-built Shanghai facility — totals 2,976 hours. YouTube ingests that every six minutes.19,21 And the shortcut doesn't work: Meta's V-JEPA 2 pre-trained on a million hours of web video and still required real robot data for control, because video lacks metric scale, 3D geometry, and action labels — the exact dimensions that make physical data physical.17

And note what is missing from this chart: the celebrated datasets that built 2D vision — ImageNet's 14 million images, LAION's 5.85 billion image-text pairs — are measured in images, not experience, and have no place on an hours axis. Meanwhile the corpora behind today's 3D-vision foundation models (VGGT, MapAnything) total ≲1,000 hours-equivalent of real capture — the entire spatial-AI field trains on less footage than a single television network broadcasts in six weeks.37

One more empirical result sharpens the requirement. The best scaling study in robot learning found that policy generalization follows a power law in the diversity of environments and objects — not the number of demonstrations collected in any one place.23 The binding constraint is not hours. It is coverage of reality: many places, many objects, many conditions, refreshed over time. Act I established the demand. The rest of this essay is about who can supply it — and what the supplier should do with the position.


Act II · Section III

The capture problem

Grant Act I and the question stops being philosophical and becomes industrial: how, exactly, do you build the physical world's Common Crawl? There are only four known supply models, and they have very different cost curves.

Four ways to capture reality
Supply modelWho bears capexMarginal cost of a new placeEnvironmental diversityFreshnessExamples
Corporate fleetThe company — cars, sensors, driversHigh — a vehicle-trip per locationFollows corporate priorities; rich-world bias~1 pass every 1–2 yearsGoogle Street View: 2007–2019 to reach 16.1M unique km28
Robot fleetThe company — billions in hardwareVery high — a deployed robotSingle embodiment; facility- and task-boundContinuous, but narrowTesla keeps 1,000+ Optimus units in-house purely for data31
Teleop / data farmThe company — facilities & operatorsLinear — $136–340 per operator-hourLimited to staged facilitiesOn demand, at costXDOF ($70M launch, ~20 lab customers), AgiBot's 4,000 m² plant19,30
DePIN networkContributors — capex pushed to the edgeNear zero — token emission + validationBy construction — supply appears wherever people liveSteerable — boost rewards where data is staleHelium, Hivemapper, OVER27–29

The first three models share a property that should alarm anyone underwriting them: their cost scales with coverage. Every new neighborhood is another vehicle-trip, another robot, another operator-hour. But coverage — environmental diversity — is precisely the variable the scaling laws reward. The economics and the science point in opposite directions. Except in one model.


Act II · Section IV

DePIN: the only supply curve that bends

Decentralized Physical Infrastructure Networks solve capture with a financial instrument instead of a fleet: a token that converts the network's future demand into present supply. Contributors buy their own sensors, capture where they live, and are paid in equity-like upside; the company's marginal cost of a new location collapses to validation and storage. Capex doesn't disappear — it is distributed to the edge, borne by thousands of participants who each hold a stake in the network they are building.

This is no longer a theory. Messari counts 650+ DePIN projects and 8.8 million active devices, a sector that has crossed from token speculation into revenue fundamentals.27 Two precedents matter most here:

Helium bootstrapped a global wireless network to roughly one million IoT hotspots — infrastructure a telecom would have needed billions in capex to deploy — by paying contributors in tokens to plug radios into their own windowsills.27

Hivemapper is the direct proof for mapping: its dashcam network went from 2% of the world's roads in March 2023 to 20% in eighteen months to roughly one-third today — coverage growth the company benchmarks at five times the pace Google Street View managed with corporate fleets, and with 100+ repeat passes where Google averages one every two years.28,29 A decentralized network out-mapped the best-capitalized mapping company in history on speed, coverage, and freshness, simultaneously.

Fleets buy coverage with capex. DePIN buys it with alignment — and alignment scales to places no fleet will ever be sent.

Three properties make DePIN uniquely fitted to the physical-AI data problem, mapping one-to-one onto Act I's requirements: diversity by construction — supply emerges wherever contributors live, and since generalization scales with environment diversity, a DePIN network's growth pattern is the scaling law's demand pattern;23 elastic, steerable supply — token rewards are a dial, so a demand signal from model training ("more indoor retail, more rain, more night scans") becomes a rewards boost and the network re-aims itself within days; and freshness as a native property — contributors re-scan because earning is continuous, turning the corpus into a living time-series of the world rather than a one-pass archive.

The honest caveat: DePIN's known failure mode is incentive misalignment — rewarding volume over quality. The mitigation is architectural: validation before rewards vest, demand-side revenue burning supply, reward curves tuned to what training actually consumes. The networks that survived the 2022–24 shakeout are the ones that built this discipline, and the sector's multiple compression (1,000× → 10–25× revenue) shows the market now prices fundamentals.27


Act II · Section V

OVER: the asset — and the floor

OVER has been running this playbook since before "physical AI" had a name — aimed at the hardest, highest-value layer: not road video, but metric-scale 3D reconstructions of real places. Through Map2Earn, anyone with a smartphone scans a location; the capture is validated, reconstructed into photogrammetry-grade 3D, and registered to the map; the mapper earns OVR tokens. OVRLand — ownership of 300 m² spatial domains with publishing rights and revenue share — gives contributors a durable stake in the network's success, not just piecework wages. The result is the world's largest crowdsourced 3D map:32

275K+
locations mapped, across every continent
1,236 TB
of real-world 3D data — geometry, imagery, poses
105M+
spatially registered images
273K+
locations live on OVER's Visual Positioning System

What makes this corpus AI-grade rather than merely large is that it carries the labels web video lacks. Every capture is a 3D reconstruction with metric scale, camera poses, and spatial registration — the exact supervision signals world models, VLA navigation stacks, and sim-to-real pipelines starve for (recall V-JEPA 2's million web-video hours failing to substitute for grounded data). The same asset serves three demand curves at once: training data for world and geospatial models, digital twins that seed simulation pipelines like NVIDIA's data factories, and a live localization service — VPS with centimeter-grade, TEE-encrypted pose estimation — that robots consume as an API, today.32

01 · INCENTIVE
Token rewards

OVR emissions + OVRLand stakes recruit mappers wherever they live.

02 · CAPTURE
The world, scanned

Smartphone photogrammetry → validated, metric-scale 3D — no fleets, no facilities.

03 · ASSET
Living 3D map

275K+ locations, re-scanned over time — a time-series of reality.

04 · MODELS
LGMs & VPS

Large Geospatial Models and localization trained on the proprietary corpus.

05 · REVENUE
APIs & data sales

Robotics, XR and enterprise demand pays the network — funding richer rewards.

Note the cost structure the flywheel implies. When a fleet company wants location #275,001, it dispatches hardware. When OVER wants it, it adjusts a reward curve — the capex was already bought, by someone who owns a piece of the outcome. That is why OVER's coverage map looks like the world (Toronto, Bangkok, Charleroi, Nghĩa Trụ, Atlanta, Shenzhen — this month's scans alone) rather than like a fleet-deployment plan.32

Pricing the floor

Suppose OVER only ever licenses this asset. We already know what markets pay when a platform's bottleneck is data — even when the raw material is free: Meta paid $14.3B for 49% of Scale AI (~$29B); Surge, bootstrapped to $1.2B of revenue, has negotiated at $25–30B; Mercor repriced from $2B to $10B in eight months and is reportedly in talks at $20B.33–35 Those are the multiples for refining scraped text. Physical AI's raw material is proprietary from the first byte. Licensing at these comparables is a real business — and it is the floor, not the thesis. What raises the ceiling is what happened to the other two inputs over the last twelve months.


Act III · Section VI

The twist: intelligence just became infrastructure

Until about a year ago, this essay would have ended at licensing. A data vendor's ceiling was a data vendor's multiple, because moving up the value chain required the two things no data company could get: a frontier research team and a nine-figure compute budget. The compute wall fell first — an H100-hour that cost $8+ at the 2023 peak rents for $1.50–2.50 today, while Stargate, Colossus and Meta's gigawatt campuses race to make FLOPs abundant.24,25 The research wall is falling now, and it is measurable.

METR tracks how long a software or ML-research task (in expert-human time) an AI completes autonomously at 50% reliability — precisely the substance of training a foundation model. That horizon doubled every seven months from 2019 to 2025. Since January 2024 it has doubled every ~105 days.2,3

The autonomy explosion — METR 50% time horizon
expert-human task length an AI completes at 50% reliability
08 h16 h24 h32 h40 h2023202420252026202740-h work-weekOct 2026GPT-4 — 50% horizon ≈ 3.5 minGPT-4GPT-4o — 50% horizon ≈ 7.0 minGPT-4oClaude 3.7 Sonnet — 50% horizon ≈ 60 min (1.0 h)Claude 3.7 Sonneto3 — 50% horizon ≈ 121 min (2.0 h)o3Claude Opus 4 — 50% horizon ≈ 101 min (1.7 h)Claude Opus 4GPT-5 — 50% horizon ≈ 214 min (3.6 h)GPT-5Claude Opus 4.5 — 50% horizon ≈ 320 min (5.3 h)Claude Opus 4.5GPT-5.2 — 50% horizon ≈ 394 min (6.6 h) (reported range)GPT-5.2Claude Opus 4.6 — 50% horizon ≈ 330 min (5.5 h) (reported range)Claude Opus 4.6Claude Mythos Preview — 50% horizon ≈ 960 min (16.0 h)Claude Mythos Preview≥ 16 h — past METR’s measurable ceiling95% CI 8.5 h – 55 hMETR TH1.1 measurementMETR-reported rangeprojection at current doubling (~110 days)
1 min10 min1 h8 h40 h160 h2023202420252026202740-h work-week160-h work-monthOct 2026May 2027GPT-4 — 50% horizon ≈ 3.5 minGPT-4GPT-4o — 50% horizon ≈ 7.0 minGPT-4oClaude 3.7 Sonnet — 50% horizon ≈ 60 min (1.0 h)Claude 3.7 Sonneto3 — 50% horizon ≈ 121 min (2.0 h)o3Claude Opus 4 — 50% horizon ≈ 101 min (1.7 h)Claude Opus 4GPT-5 — 50% horizon ≈ 214 min (3.6 h)GPT-5Claude Opus 4.5 — 50% horizon ≈ 320 min (5.3 h)Claude Opus 4.5GPT-5.2 — 50% horizon ≈ 394 min (6.6 h) (reported range)GPT-5.2Claude Opus 4.6 — 50% horizon ≈ 330 min (5.5 h) (reported range)Claude Opus 4.6Claude Mythos Preview — 50% horizon ≈ 960 min (16.0 h)Claude Mythos Preview≥ 16 h — past METR’s measurable ceiling95% CI 8.5 h – 55 hMETR TH1.1 measurementMETR-reported rangeprojection at current doubling (~110 days)
Solid points: METR Time Horizon 1.1 measurements (228 software/ML/cyber tasks). Hollow: ranges reported on METR's live dashboard. Mythos Preview: ≥16 h (95% CI 8.5–55 h); METR notes measurements above 16 h are unreliable with the current suite. Trend fitted on 2024+ measured points (~110-day doubling); dashed segment is a naive extrapolation, not a METR forecast — it crosses a 40-hour work-week ≈ Oct 2026. Sources: METR (May 2026); Kwa et al. 2025.[1–3]

Read the right edge. In May 2026, METR added Claude Mythos Preview — the model class behind Claude Fable 5 — and reported it at or beyond 16 hours, past the point where the benchmark can measure at all. The model outgrew the ruler.3 The receipts are concrete: Anthropic reports Mythos 5 running a week of largely autonomous research, designing and training an ML model that outperformed a recent Science-published model at 1/100th the size;4 DeepMind's AlphaEvolve broke a 56-year-old matrix-multiplication record and recovers ~0.7% of Google's global compute;5 on METR's RE-Bench, AI agents score 4× human experts at two-hour research budgets.6 And the day before this essay was finished, Moonshot released Kimi K3 — 2.8 trillion parameters, open weights, third on GDPval behind only Claude Fable 5 and GPT-5.6 — at $15 per million output tokens, with the open-closed gap at ~3–4 months and capability prices falling ~10× per year.7–9

The specialized intelligence needed to train a state-of-the-art model — the scarcest resource of the last decade — now has a list price. You can no longer build a moat out of PhDs. But you can build one out of reality.

To be precise: this is not a claim that human researchers are obsolete. It is a narrower claim about market structure — the engineering intelligence required to train a SOTA physical-AI model is ceasing to be a differentiator, because every serious team works with frontier research agents and the recipes diffuse through open weights within months. When everyone has the same intelligence on tap, it prices like electricity: essential, and margin-free. Which means the wall between "data vendor" and "AI lab" — a wall made of people and capex — is gone.


Conclusion · Section VII

The value ladder

Put the three acts together and OVER's position prices in tiers. Acts I and II justify the first two rungs on their own. Act III unlocks the third — and the third is where foundation-model economics live.

Rung 1 · TodayTHE FLOOR

License the scarce asset

Raw and curated 3D data, sold to labs racing a 100–1,000× demand shock. Comparables already printed in the text era — Scale ~$29B, Surge $25–30B talks, Mercor $20B talks — for refining data that was free. OVER's raw material is proprietary from the first byte.33–35

Rung 2 · LiveTHE SERVICES

Sell reality as an API

VPS localization (centimeter-grade, TEE-encrypted, live on 273K+ locations), robotics endpoints, digital-twin feeds for simulation pipelines. Recurring, per-call revenue on the same asset.32

Rung 3 · UnlockedTHE PRIZE

Train the models. Own the economics.

Large Geospatial Models and world models trained on the proprietary corpus — with compute rented by the hour and research intelligence rented by the token. Labs without proprietary data get squeezed between open weights and data owners; data owners with models capture the margin of the era. OVER's LGM program is already running.32

▲ UNLOCKED BY ACT III — INTELLIGENCE AS A COMMODITY

Compute you can rent. Intelligence you can now rent. Reality you cannot. So don't sell the crude — own the refinery. It just became rentable.

The thesis, closed · info@ovr.ai

Section VIII

What could break this

Objection 01Synthetic data substitutes for capture

The sim engines themselves are trained on reality — Cosmos consumed 20 million hours of real video — and NVIDIA's own data-factory blueprint pairs synthetic generation with real seeds, because simulation must be anchored to escape the sim-to-real gap. Synthetic data is a multiplier on ground truth, not a replacement; digital twins of real places are its feedstock.14,26

Objection 02Web video is enough

V-JEPA 2 ran this experiment at a million-hour scale: strong representations, but real robot data was still required for control, and performance degraded off-distribution. Video without geometry, scale, and action labels teaches what the world looks like — not how it responds. The priced layer is the grounded one.17

Objection 03Labs vertically integrate their own fleets

They are — which is the strongest evidence data is the scarce input. But fleet data is single-embodiment and facility-bound, while generalization is bought with environmental diversity; that's why the same labs simultaneously buy from external providers. Fleets and diverse-capture networks are complements, and only one of them is for sale.23,31

Objection 04Timing — and DePIN incentive quality

Timing is the honest residual risk; two mitigants: data demand is repricing now, on rookie-numbers compute, and capture networks compound slowly — Hivemapper's three-year head start is the kind incumbents never closed. On quality: DePIN's failure mode is rewarding volume over quality; the mitigation is validation before rewards vest, demand-side burn, and reward curves tuned to what training consumes — the discipline that separates surviving networks from the 2022 cohort.27

Appendix

Sources

  1. METR — Task-Completion Time Horizons of Frontier AI Models (updated May 8, 2026). metr.org/time-horizons
  2. METR — Time Horizon 1.1 (Jan 29, 2026); Kwa et al., arXiv:2503.14499. metr.org
  3. OfficeChai — Claude Mythos shows 50% time horizon of 16+ hours on METR (May 9, 2026). officechai.com
  4. Anthropic — Introducing Claude Fable 5 & Mythos 5 (Jun 2026). anthropic.com
  5. Google DeepMind — AlphaEvolve (May 2025). deepmind.google
  6. METR — RE-Bench (2024). arXiv:2411.15114
  7. MarkTechPost — Moonshot AI releases Kimi K3 (Jul 16, 2026). marktechpost.com
  8. Artificial Analysis — GDPval-AA v2 snapshots (Jul 16–17, 2026). artificialanalysis.ai
  9. Epoch AI — Open vs closed frontier gap; LLM inference price trends; a16z "LLMflation". epoch.ai · a16z.com
  10. Epoch AI — Grok 4 training resources (~246M H100-hours, Sep 2025); models over 1e25 FLOP. epoch.ai
  11. Meta — Llama 3.1: 16,384 H100s; ~31M GPU-hours (Jul 2024). ai.meta.com
  12. NVIDIA — GR00T N1 (~50,000 H100-hours, 2025). arXiv:2503.14734
  13. Kim et al. — OpenVLA: 21,500 A100-hours, 64 GPUs, 27 epochs (2024). arXiv:2406.09246
  14. NVIDIA — Cosmos: 20M hours of video, 10,000 H100s × 3 months (Jan 2025). arXiv:2501.03575
  15. Epoch AI — Can AI scaling continue through 2030? (~500T indexed tokens). epoch.ai
  16. HuggingFace — FineWeb (15T tokens); Villalobos et al., arXiv:2211.04325. huggingface.co
  17. Meta — V-JEPA 2 (Jun 2025). arXiv:2506.09985
  18. Open X-Embodiment (2023). arXiv:2310.08864
  19. AgiBot World — 1M trajectories / 2,976 h / 100 robots (2025). arXiv:2503.06669
  20. DROID (2024), arXiv:2403.12945; RT-1 (2022), arXiv:2212.06817; Ego4D, ego4d-data.org
  21. Statista — 500+ hours uploaded to YouTube per minute. statista.com
  22. Survey — Training Data for Generalist Robot Foundation Models (2025): ~15M open episodes, 3–6 OOM below language/vision. researchgate.net
  23. Lin et al. — Data Scaling Laws in Imitation Learning (CoRL 2024). arXiv:2410.18647
  24. Latent.Space — "$2 H100s" (2024); IntuitionLabs H100 rental comparison (2026). latent.space · intuitionlabs.ai
  25. OpenAI — Stargate ($500B / 10 GW, Jan 2025); xAI Colossus; Meta Hyperion. openai.com
  26. NVIDIA — GTC 2026: Physical AI Data Factory blueprint (Mar 2026); Crunchbase — robotics VC $18.8B in 2026. blogs.nvidia.com · crunchbase.com
  27. Messari — State of DePIN 2025: 650+ projects, 8.8M active devices, $72M FY25 on-chain revenue, 10–25× revenue multiples; Helium ≈1M IoT hotspots (GridEcon/Solana case studies). messari.io · gridecon.com
  28. TechCrunch — Hivemapper at 2% of global roads / 1M unique km (Mar 2023); Google Street View 16.1M unique km, 2007–2019. techcrunch.com
  29. Hivemapper / Bee Maps — 20% of global roads in 18 months; ~⅓ today; 5× Street View pace; 100+ repeat passes (company-reported). beemaps.com · hivemapper.com
  30. TechCrunch — XDOF launches with $70M, ~20 customers incl. frontier labs (Jun 2026); General Intuition at $2.3B (Jun 2026); SVRC — teleop $136–340/hr (directional). techcrunch.com
  31. Tesla Q4-2025 earnings — 1,000+ Optimus units used primarily for data collection (Jan 2026). finance.yahoo.com
  32. OVER — live platform statistics; Map2Earn; OVRLand; VPS & TEE; LGM & robotics APIs (accessed Jul 17, 2026). overthereality.ai · HF dataset
  33. Forbes — Meta invests $14.3B in Scale AI at ~$29B (Jun 2025). forbes.com
  34. Bloomberg — Surge AI in talks at $25B+; $1.2B 2024 revenue (Jul 2025). bloomberg.com
  35. TechCrunch — Mercor: $10B (Oct 2025); talks at $20B (Jul 2026). techcrunch.com
  36. 2026 compute disclosures — VGGT: 64 A100 × 9 days (arXiv:2503.11651); MapAnything: 64 H200 × 10 days (arXiv:2509.13414); NVIDIA GR00T N1.7 model card, 30,720 GB200-hours (huggingface.co); Epoch AI — GPT-5 training-compute notes (epochai.substack.com).
  37. Dataset scales — nuPlan: 1,500 h (arXiv:2106.11810); Panda-70M: 70.7M clips / 167K h (Snap Research); 3D-FM mix components: ScanNet (arXiv:1702.04405), DL3DV-10K (arXiv:2312.16256), CO3D (arXiv:2109.00512), Aria Digital Twin (arXiv:2306.06362), per the VGGT/MapAnything training-data lists; ImageNet (image-net.org); LAION-5B (laion.ai); InternVid ~760K h (2023).