Lab-ready isn't deployment-ready.
We work the sim-to-real loop with you: refine your data, run your policy on real hardware against pass criteria agreed up front, and review the failures together.
The sim-to-real gap is a loop problem, not a model problem.
Teams rarely stall on model capability. They stall between a policy that works in the lab and one that survives a real workcell.
- Physics doesn't transferReal actuators can't hold the torque profiles sim asked for.
- Teleop data is off-policyDemonstrations come from a person, not your model.
- Nobody defines "working"Lab success rates meet the customer's criteria in the long tail.
01 · Actuator physics doesn't transfer
A policy trained in sim requests torque profiles real actuators can't sustain. Thermal throttling, joint backlash and unit-to-unit variance mean two of the same arm off the line don't share dynamics. "Embodiment-agnostic" turns out embodiment-specific.
02 · Teleop data is structurally off-policy
Bootstrap demonstrations come from a human operator, not your model. That covariate shift doesn't shrink by collecting more of the same. It shrinks through on-policy rollouts on real hardware, with failures fed back into targeted collection.
03 · Nobody defines "working" before deploy
Teams ship on lab success rates, then meet the customer's acceptance criteria in the long tail: the wrinkled shirt, the 4pm glare, the pallet 3cm off spec. Without an eval harness that encodes those criteria, readiness is a guess.
Supplying hardware and data to robotics teams, we keep seeing projects stall in these three places. This is our read of the field, not a measured result.
RL²: reinforcement learning in real life.
Five stages on real hardware, run as a cycle. Each maps to a component we've built; today we run them as a supervised engagement, not an automated loop.
01 · Deploy on real hardware
Run the policy on a real arm, hand and camera, not a sim twin, as a supervised session. Real actuator dynamics, real latency, and the real failure surface.
- On-policy rollouts on our hardware, run as a supervised session
- Shadow or operator-supervised (copilot) mode
- Every predicted action and operator decision logged
02 · Capture the failure distribution
Log the whole rollout: the misses, not just the wins. MCAP ingestion with time-synced multi-sensor alignment keeps every stream on one clock.
- Rollout and failure logging
- MCAP ingestion
- Time-synced multi-sensor streams
03 · Evaluate against agreed pass criteria
Score against pass thresholds agreed with you: success-rate and latency checks, plus operator review of each action for tasks a unit test can't grade.
- Pass/fail checks on success rate and latency
- Operator review of each action
- Eval runs recorded per model version
04 · Scope the next collection
Our team reviews failed rollouts with you and scopes the next collection: the scene, object or condition worth capturing. Today this is a hands-on service, not an automated step.
- Failure review with your team
- Collection by our in-house operators
- Not more of the same
05 · Retrain, check, redeploy
Retrain with the targeted data, re-run the eval, and go again. A regression can open a fine-tune job from the last good checkpoint for review.
- Fine-tune from the last good checkpoint, reviewed
- Regression checks before promoting
- Loop repeats
One stack, from raw capture to a readiness report.
Built on the data infrastructure we already run. Each row says where that capability stands today.
Data refinery
Egocentric and teleop capture, refined on request: quality control, 21-point hand landmarks, action segmentation and LeRobot-v2 packaging. Run it from the workbench or the centeros refinery CLI.
Managed teleop
Our in-house operators and a WebSocket teleop layer with MCAP ingestion and time-synced multi-sensor alignment. This is the source of your on-policy and targeted data.
Model registry
Register policy versions and run supervised inference sessions, in shadow or copilot mode, with every predicted action and operator decision logged for evaluation.
Evaluation runs
Pass criteria on success rate and latency, per model version; recovery and per-task criteria are next. A regression can open a fine-tune job from the last good checkpoint for review.
CLI · API · MCP
Scriptable through the centeros CLI, with key operations also available as MCP tools, so the loop can run from your own pipeline or an agent, not just a dashboard.
Dataset access
Index, inspect and share datasets with signed, time-limited access.
We're not a foundation-model lab. We help get a policy from demo to production with evaluation runs and the targeted data that make it better, run today as a supervised engagement. A readiness engagement is a paid improvement service, not neutral third-party certification. Disclosure: RC also sells the hardware and data referenced on this page.
Acceptance criteria, not vanity metrics.
We agree numeric pass thresholds with you and re-run them on every model version. Structured task, environment and edge-case criteria are in development.
Regression checks
Metric thresholds recorded across versions, so you can run a regression check on a policy before promoting it.
Operator-in-the-loop evaluation
A shirt folded "correctly" isn't a unit test. Our operators supervise real rollouts and approve or reject each action, so the score reflects the workcell, not a proxy.
Failure review → next collection
Our team reviews failed rollouts with you and scopes the next collection. This is a hands-on service today: the eval reports failure types, not a data order.
Measuring hands? Our hand-measurement protocol is public (draft v0.1, open for comment) →
Each eval run produces a readiness report.
It answers one question: how often, and how fast, does this policy complete this task on real hardware?
Readiness report- Model
- sample policy · v0.3
- Task
- hypothetical
- Sessions
- 3 seeded
- Decisive actions
- 52
- Success rate88.5%46 ✓ / 6 ✗
- Mean action latency161 ms
- Top failure modeoperator_rejected×4
- Secondarytimeout×2
- VerdictPASSvs sample thresholds
Fully automated policy scoring across arbitrary tasks is where we're investing next. Today, readiness combines automated checks with operator review, so the number means something in the real world.
Need the hardware to run evals on?
Buy or lease the exact rig you'll deploy on, or start from an existing dataset.
- Arms & handsBuy the exact rig you'll deploy on.Store
- Lease a rigRun the loop without the capex.Leasing
- DatasetsBootstrap collection from real data.Marketplace
- Dexterous handsCompare and measure hands before you choose one.Dexterity
Works in the lab but not the workcell? Let's run the loop.
Book a walkthrough of the refinery, managed teleop and evaluation runs. We host builders at 90 Welsh St in San Francisco.