WorldData / Strategy Publication 01 Research frozen 28 July 2026
Competitive atlas + plan of record

Own the physical data plane.

A source-grounded map of the real-data, teleoperation, robot-infrastructure, simulation, and evaluation markets—and the retrofit strategy that turns a fragmented stack into a defensible company.

01
Managed data programsScale · Encord · Roborax · specialists
02
Capture & teleoperationTrossen · UMI · ALOHA · Isaac Teleop
03
Robot-data infrastructureFoxglove · Roboto · Rerun · Viam
04
Simulation & validationNVIDIA · Applied · Parallel Domain · open stacks
05
Internal buildROS 2 · MCAP · scripts · cloud · simulator
BuildA hardware-neutral evidence and calibration layer
SellMeasured capability, not recording hours
AvoidUniversal control liability and another raw format
01 / Executive decision

Build the retrofit. Constrain its job.

The winning company is not “another dataset vendor.” It is the neutral layer that makes heterogeneous physical runs trustworthy, comparable, trainable, simulatable, and evaluable.

Recommendation

Standardize the data plane first. Keep the robot’s certified controller and safety system in charge. Monetize the verified capability above the appliance.

The retrofit should connect existing robots, teleoperation devices, and external sensors; establish a common clock and coordinate system; preserve raw evidence; emit model-ready views; create calibrated simulator counterparts; and close the loop from failure to targeted recollection. Commanded teleoperation belongs in a separate, hardware-specific safety envelope—not in the universal core. [E‑007]

1 loopReal capture → canonical evidence → synthetic expansion → hidden evaluation → targeted recollection
6Initial integration families recommended before broad industrial expansion
90 daysProposed commercial pilot measured by capability lift, not hours recorded
The core strategic insight: sensors, edge computers, operators, and simulation compute can all be rented. The compatibility graph, calibration ledger, transformation lineage, real↔sim residuals, and private evaluation network are what compound. [E‑009]
02 / Market anatomy

“Robot data” is seven different goods.

A buyer does not purchase bytes in the abstract. Different data planes teach different things, carry different scarcity, and support different business models.

AWorld priors

Video, images, audio, maps, CAD and spatial context. Abundant generically; scarce in rights-cleared specialist environments.

BObservation

Calibrated RGB-D, LiDAR, radar, thermal, IMU and geometry. Valuable when sensor-matched and synchronized.

CAction & contact

Robot state, commands, force, tactile, slip, controller mode and latency. Premium learning signal.

DFailure & recovery

Interventions, near misses, partial progress and correction. Scarce and disproportionately useful.

ESimulation & eval

Assets, physics, tasks, synthetic rollouts, counterfactuals and hidden tests. Reusable infrastructure.

The full taxonomy also separates simulation ingredients from generated rollouts and evaluation evidence. This matters because simulation throughput is increasingly abundant while task authoring, contact fidelity, system identification, scenario validity, and real-world correlation remain difficult. [E‑001]

Real and synthetic data are complements

NVIDIA reported generating 780,000 synthetic trajectories—described as 6,500 human-demonstration hours—in 11 hours, and reported a 40% performance improvement from mixing synthetic and real data relative to real-only training. The correct conclusion is not that real data disappears: scarce demonstrations, target distributions, system identification, and real holdouts determine whether synthetic scale is useful. [S‑01]

The minimum valuable episode is causal

A useful episode contains task intent and success predicates; embodiment, firmware and controller state; calibrated observations; native robot state and requested/executed actions; outcome and intervention; clock and transform uncertainty; provenance, rights and lineage. LeRobot and RLDS provide important interchange structures, but the commercial evidence contract must also preserve calibration, contact, action semantics, privacy, and evaluation exclusion. [S‑02] [S‑03]

03 / Competitive atlas

Five budgets. No single winner.

The buyer’s shortlist changes with the job: outsource collection, buy a rig, manage logs, simulate worlds, or keep building internally. The company must position against each substitution set.

Tier 1A — managed physical-data programs

ProductPublic surfaceStrongest advantageOpening for the retrofit companyThreat
Scale Physical AIGlobal factories and collectors, Scale Harness, bimanual demonstrations, annotation, platform and internal validationOperating scale, enterprise trust and frontier-lab relationships; claims 1,000+ demonstration hours dailyCustomer-owned appliance, installed mixed-fleet compatibility, simulator neutrality and independent evaluationVery high
Encord Physical AILabs, in-field operators, leader/follower, UMI, multimodal ingestion, curation, annotation and deployment feedbackStrong end-to-end data platform and enterprise deployment optionsDeeper controller normalization, signed calibration, cross-simulator delivery and pervasive hardware applianceVery high
RoboraxTeleop, demonstrations, sensor capture, annotation, simulation, evaluation and several operating modelsBPO-backed labor and compliance breadth; claims 2.4M+ trajectoriesPublic product proof and customer outcomes require diligence; win on technical evidence and installed-base depthHigh
Scaled MotionContact-rich, deformable and mobile-bimanual fine-tuning packs with custom capture kitsCorrect high-value task wedgeExecution maturity and platform breadth need validationMedium
Defined.aiMarketplace plus custom multimodal physical-AI collectionExisting data operations, audio assets and enterprise sourcingLess visible controller-level action and real↔sim specializationMedium
Specialists
Tactum · Paddy · Operant
Focused site, vertical, task or regional programsSpeed, specialization and priceEarly public proof; potential partners, acquisition targets or wedge competitorsWatch

Tier 1B — capture and teleoperation

Hardware is getting cheap

  • Trossen Mobile AI packages a bimanual leader/follower station, synchronized RGB-D, MCAP and LeRobot/HDF5 outputs.
  • TRumi commercializes a portable UMI-style workflow, publicly listing dual kits below $2,200 with cameras.
  • ALOHA 2, UMI and GELLO provide open collection designs and strong research baselines.
  • Isaac Teleop provides a unified device and retargeting framework across ROS 2, Isaac Sim and Isaac Lab.
  • LeRobot connects hardware, datasets, policies, simulation and community distribution.
Implication

A recorder is not a moat.

The defensible chain is sensor identity → clock → calibration → controller state → action semantics → outcome → training transform → twin → hidden result.

Tier 2A — data and robot operations

ProductOwns todayRecommended posture
Foxglove + MCAPOpen raw container, visualization, ingest, search, governance and evaluationEmbrace MCAP; integrate Foxglove rather than inventing a container/viewer
RobotoCross-format log ingestion, automated analysis, anomalies, root cause, curation and fleet reliabilityPotential analytics partner; differentiate in acquisition, calibration, simulation and capability delivery
RerunOpen visualization plus query, transformation, catalog and lineageUse as an optional inspection surface
Formant / InOrbitFleet operations, telemetry, teleoperation, missions, incidents and orchestrationCoexist: they optimize uptime; the foundry certifies learning evidence
ViamHardware abstraction, modules, connectivity, capture, sync, query, ML and deploymentClosest horizontal edge threat; stay narrow and deeper on action/contact/evaluation integrity

Tier 2B — simulation and validation

PlatformCore advantageStrategic response
NVIDIA Isaac / Cosmos / GR00T / Data FactoryBroadest integrated real/sim teleop, OpenUSD worlds, physics, synthetic generation, training and evaluationFirst-class target, never the only target; own field execution, calibration and neutral evaluation
Applied IntuitionEnterprise autonomy data, simulation, reconstruction and V&VAvoid vehicle/autonomy head-on; focus contact-rich manipulation and installed robot cells
Parallel DomainPhotoreal sensor simulation and real-scene replicasComplement for perception; differentiate in manipulation dynamics and action data
ForetellixCoverage-driven scenario verification for safety-critical autonomyPotential evaluation partner and methodological model
MuJoCo Playground / ManiSkill3 / RoboCasaOpen, fast physics, training, tasks, demonstrations and benchmarksUse them; sell customer-specific task authoring, system identification and real correlation
The most important platform threat is NVIDIA. Its 2026 Physical AI Data Factory explicitly unifies processing, curation, synthetic generation, reinforcement learning and evaluation. A startup whose entire story is “we connect real data to Isaac” has no durable boundary. [S‑04]
04 / Strategic opening

The open seam is evidence continuity.

Managed providers operate programs. Hardware vendors provide rigs. Data platforms organize logs. Simulators create worlds. The company should own the integrity contract across all four.

What compounds

Build these assets

  1. Robot/controller/firmware compatibility graph
  2. Longitudinal clock and calibration ledger
  3. Native-preserving action and task ontology
  4. Validated training-view transformation recipes
  5. Real↔sim dynamics, sensor and latency residuals
  6. Coverage and data-influence models
  7. Leakage-controlled private evaluation network
  8. Task-specific collection operating curve
What commoditizes

Do not confuse these for moat

  • Generic first-person video hours
  • Camera and edge-compute resale
  • A ROS bag uploader or new raw format
  • Labor arbitrage without technical QA
  • Public assets without real calibration
  • Normalized actions that erase native dynamics
  • A marketplace before rights and transfer value are proven

The data-rights posture should be conservative: the customer owns its raw and task-derived data; the company retains adapter code, generic schemas, de-identified calibration statistics, aggregate quality patterns, generic residual models, and benchmark methods. Cross-customer trajectory reuse should be explicit opt-in. This creates trust without giving away the compounding technical system. [E‑009]

05 / Proposed product

The Physical Data Plane

A rugged edge appliance is the distribution mechanism. The actual product is a certified chain from a physical run to a model result.

Closed-loop architecture for the Physical Data Plane Existing robots, sensors and teleoperation devices flow into an edge capture and calibration layer, then canonical evidence, training views, digital twins, synthetic expansion, customer models, hidden evaluation and failure-directed recollection. PHYSICAL INPUTS Existing robot state · controller · safety Retrofit sensors RGB-D · IMU · F/T · tactile Teleoperation leader · XR · UMI · exo TRUST BOUNDARY Vendor adapter read-only default · guarded control Clock + calibration frames · latency · uncertainty Episode QA task · outcome · failure · rights EVIDENCE + LEARNING Canonical evidence store raw MCAP · immutable metadata · lineage Training views LeRobot · RLDS · native Twin sim + ID Customer model train · adapt · deploy Synthetic expand Hidden real + simulation evaluation capability · safety · OOD · regression FAILURE → TARGETED RECOLLECTION

FIG. 02 — The product boundary. The universal core owns evidence integrity; hardware-specific control remains behind an entitled safety gateway.

01
Capture NodeOffline-first, encrypted edge recording; ROS/DDS, vendor API and direct sensor ingestion; hardware timestamps where possible.
02
Robot AdapterNative state, requested/accepted/executed action, controller mode, safety, I/O, firmware and latency.
03
Calibration AuthorityIntrinsics, extrinsics, tool frame, clock offset, F/T zero, payload, compliance and session uncertainty.
04
Episode & Quality EngineTask graph, success, intervention, failure, accepted yield, drift, dropped frames and feasibility checks.
05
Dataset FabricImmutable raw evidence plus versioned LeRobot, RLDS, HDF5/Zarr/Parquet and customer-native views.
06
Twin & Synthetic ServiceIsaac and MuJoCo/SAPIEN adapters, controller and sensor models, task logic and quantified real↔sim residuals.
07
Capability EvaluationFrozen real holdouts, scenario suites, OOD variants, model registry and reproducible release gates.
06 / Retrofit strategy

Capture broadly. Control selectively.

The safe product expansion is a ladder. Each rung adds value—and a different class of engineering, safety and liability.

L0External observationFastest · lowest risk
L1Robot state + controller contextLaunch default
L2Guarded vendor-supported teleoperationPlatform certification
L3Customer policy execution gatewayLater product
L4Low-level torque/current controlResearch exception only
P0 / Months 0–6UR · Franka · xArm

Accessible interfaces, rich manipulation data and strong overlap between learning labs and industry.

P0 / PartnerKinova · Trossen / ALOHA

Research, medical/assistive relevance, bimanual collection and packaged leader/follower workflows.

P1 / Months 7–9Unitree · Clearpath · PX4

Humanoid, legged, mobile and aerial expansion through ROS 2/DDS-compatible surfaces.

P2 / ChannelFANUC · ABB · KUKA

Industrial installed base, but controller generations, options, integrators and safety cells demand certification.

FamilyOfficial interface evidenceCaptureTeleopRecommended move
Universal RobotsRTDE + official ROS 2HighHighP0 production reference
Franka FR3 / PandaFCI/libfranka, 1 kHz state and control, ROS integrationsHighHigh with real-time constraintsP0 calibration reference
UFactory xArmPython/C++ SDK and ROS/ROS 2 packagesHighHighP0 affordable dual-arm/vertical rig
Kinova Gen3official Kortex ROS 2HighHighP0/P1 medical and research
UnitreeSDK2 + ROS 2/DDS across several familiesHighMedium–highP1 humanoid reference
Clearpathunified official ROS 2 APIHighHigh for guarded base velocityP1 mobile reference
FANUCofficial ROS 2 driver, high-rate control and I/OHigh on supported stackMediumP2 with integrator/OEM
ABB / KUKA / peersABB RWS, ROS-Industrial/vendor-specific optionsMedium–highController-specificPaid-demand certification
Closed humanoidsPartnership-oriented accessLow without OEMPartnership onlyDo not promise aftermarket control

Interface availability is not permission to bypass safety. Real deployments must preserve vendor controllers, safety PLCs, interlocks, e-stops, watchdogs, workspace and force limits, and site risk assessment. [E‑008]

07 / Business design

Sell a TaskPack, not a terabyte.

The commercial unit should describe a measurable capability. Recording hours and trajectories remain cost metrics inside the factory.

90-day Capability Retrofit Pilot

  1. Retrofit two customer robots and one teleoperation modality.
  2. Instrument one commercially important contact-rich task.
  3. Collect accepted successes, failures and recoveries.
  4. Deliver a calibrated simulator task and synthetic expansion.
  5. Train or support a baseline and run frozen evaluations.
  6. Report lift, accepted yield, coverage, cost and remaining failure slices.
$150k–$350kHypothesis · validate with six design partners

Positioning

For physical-AI teams that already own robots but cannot turn heterogeneous runs into reliable capability fast enough, the company is the hardware-neutral physical data plane and capability foundry. It retrofits the installed fleet, certifies every observation-action episode, expands scarce interactions in calibrated simulation, and proves model improvement on hidden tests—without replacing the customer’s robots, models, cloud or simulator.

When the company wins

  • The customer has several robots or embodiments and repeats integration work.
  • Temporal integrity, calibration, contact and outcome semantics affect policy quality.
  • Simulation exists but is not demonstrably correlated with the real task.
  • Failures and interventions are abundant but not converted into training and evaluation slices.
  • Neutrality, customer data ownership and private evaluation matter.

When it loses

  • A small lab has one stable robot and can operate an open Trossen/LeRobot workflow itself.
  • The buyer primarily needs commodity annotation or thousands of operators immediately.
  • The only pain is log visualization, fleet uptime or orchestration.
  • The customer is fully standardized on a single vertical stack and already has excellent physical-data operations.

Battlecard shorthand

Against Scale or Encord

Do not claim superior raw scale. Position the appliance as the customer-owned evidence standard across internal and outsourced sources. Win on hardware depth, signed calibration, simulator neutrality and private capability evaluation. Lose when procurement wants one established managed-service vendor.

Against NVIDIA

Isaac is a first-class target, not the product boundary. Win through field operations, non-NVIDIA support, rights-cleared site access, real hardware calibration and independent evaluation. Lose when the buyer is fully standardized on NVIDIA and staffs the last mile internally.

Against Trossen, ALOHA, UMI or LeRobot

Keep them. The company connects those sources to every other robot, applies one evidence contract, and handles enterprise operations, transformations, twins and tests. Lose when one open rig is all the customer needs.

Against Foxglove, Roboto, Rerun, Formant or Viam

They help teams see, search, operate or abstract robots. The foundry makes episodes trainable and links them to simulator and model outcomes. Integrate rather than replace unless a platform expands directly into certified acquisition and evaluation.

Against internal build

Never say “you cannot build it.” Say: keep your models and control stack; stop spending senior roboticist time repeating clock, calibration, schema, conversion and evaluation infrastructure for every robot and site.

08 / Plan of record

Twelve months to prove compounding.

The objective is not broad compatibility on paper. It is repeated deployment with falling marginal integration cost and demonstrated model lift.

Months 0–3 / Evidence contract

Make one physical run trustworthy

Canonical episode and calibration schema; ROS 2/DDS + MCAP edge recorder; LeRobot and RLDS views; UR, Franka and xArm L1 adapters; automated skew, frame, kinematic and outcome QA; one Isaac and one MuJoCo task from the same setup.

Months 4–6 / Multi-embodiment

Prove the abstraction survives new hardware

Add Kinova and Trossen/ALOHA, leader/follower, UMI/TRumi and XR ingestion; make calibration technician-operable; establish task/failure ontology; close three paid Capability Retrofit Pilots.

Months 7–9 / Installed-base distribution

Operate a small heterogeneous fleet

Add Unitree and Clearpath; production capture-node health; signed session certificates; customer VPC/on-prem deployment; Foxglove/Rerun/Roboto connectors; real↔sim residual dashboard.

Months 10–12 / Industrial channels

Certify through partners

FANUC and ABB L1 reference adapters with integrators; guarded L2 teleoperation on only two validated arm families; third-party adapter test suite; private evaluation renewal; open a narrow Physical Episode Conformance specification.

Operating metrics

MetricWhat it proves
Time to first accepted episode on a new robotAdapter and deployment leverage
Accepted episodes per staffed hourHardware, reset, operator and QA productivity
Certified sessions / total sessionsProtection from silent clock and calibration defects
Cost per accepted novel sliceCoverage value rather than repetition
Failure-to-targeted-data cycle timeClosed-loop responsiveness
Real↔sim scenario rank correlationWhether simulation guides selection and evaluation
Hidden real-test liftWhether the TaskPack bought capability
Adapter reuse ratioWhether the services business is becoming a platform
09 / Falsification

Design the company to be killable.

The retrofit thesis is attractive, but it can collapse into bespoke services, unsafe control work, or an appliance customers refuse to install. Six design partners should produce decisive evidence.

Main risks and mitigations

RiskMitigation
NVIDIA absorbs the workflowOwn the real-world last mile, multi-simulator neutrality, calibration and independent tests
Every adapter is bespokeConformance levels, adapter SDK, certification harness, explicit support matrix and paid exceptions
Control liability dominatesL1 read-only default; isolate L2; preserve vendor safety; use certified integrators
Normalization destroys signalImmutable native streams plus derived views with transformation loss and uncertainty
Customers will not share dataCustomer ownership; build moat from adapters, aggregate quality metadata and evaluation methods
Synthetic data does not transferReal seed, paired calibration trials, real holdouts and correlation thresholds for every program
Collection becomes low-margin BPOPrice accepted capability evidence; automate QA/resets; partner for labor

Six-part kill test

  1. Median new-arm L1 integration still takes more than four engineer-weeks after three reference adapters.
  2. Fewer than three of six partners will install the node on more than five robots.
  3. Customers value annotation and search but will not pay for calibration and outcome instrumentation.
  4. The system cannot find quality defects that materially affect policy performance.
  5. Real↔sim calibration fails to improve scenario ranking or transfer on two task families.
  6. Security and procurement consistently reject the appliance while accepting a pure software agent.

If the appliance fails but the evidence layer wins, pivot to an adapter SDK and managed calibration/evaluation service running on customer-owned edge hardware.

10 / Evidence ledger

Sources, confidence and limits

This publication synthesizes two local research reports totaling 17,958 words and 114 unique public links. The frozen source snapshot records checksums, evidence IDs, privacy and the publication contract. Company scale, performance and customer statements are attributed to vendors, not treated as independently audited. Strategic ratings and all prices are inferences or hypotheses.

  1. E‑001 / E‑002 — State of Physical-AI Data. Local source: PHYSICAL_AI_DATA_LANDSCAPE.md, SHA-256 5dc791449135edf641455f4022e2d66c9aeb097894fdd0571003e25a5f4fcb0f.
  2. E‑003–E‑010 — Competitive Products and Retrofit Strategy. Local source: PHYSICAL_AI_COMPETITIVE_PRODUCTS_AND_RETROFIT_STRATEGY.md, SHA-256 e059b33b8ef72f91268d2410755c001cddccc3f7aaf52f80759380b1f4e8b3f2.
  3. S‑01 — Synthetic motion generation. NVIDIA, Building a Synthetic Motion Generation Pipeline for Humanoid Robot Learning.
  4. S‑02 — Training and data ecosystem. Hugging Face, LeRobot documentation.
  5. S‑03 — Episodic sequential-data structure. Google Research, RLDS ecosystem.
  6. S‑04 — Integrated data-factory threat. NVIDIA, Physical AI Data Factory Blueprint.
  7. Direct data providers. Scale, Encord, Roborax, Scaled Motion, Defined.ai, Tactum Labs, Paddy.
  8. Capture and teleoperation. Trossen Mobile AI, TRumi, ALOHA 2, UMI, GELLO, Isaac Teleop.
  9. Robot-data infrastructure. Foxglove, Roboto, Rerun, Formant, InOrbit, Viam, Sift.
  10. Simulation and validation. Isaac Sim, Applied Intuition, Parallel Domain, Foretellix, Bifrost, MuJoCo Playground, ManiSkill3, RoboCasa.
Boundary: this is a company-building decision document, not a robot safety certification, legal opinion, or independently audited vendor benchmark. Controller access, safety requirements, data rights and performance must be verified per contract, firmware, robot and site.