
Most Robotic is a San Francisco robotics lab building Instance, a verification layer for robot learning, backed by Y Combinator.
Instance, the product built by the San Francisco robotics research lab Most Robotic, acts as a verification layer for robot learning. Teams submit a task description together with camera footage of a rollout, and the system returns a success verdict backed by grounded subtask evidence for that episode. The offering began as a success detector for robot policies and was benchmarked on more than ten thousand human-labeled episodes spanning seven robot platforms, where it reached higher accuracy than Claude Opus 4.8 at a fraction of the latency.
Access is deliberately simple: a public demo at demo.instancelabs.ai lets anyone check a rollout, while a plain HTTP POST to the /v1/verify endpoint serves teams that want verification wired into their training loop. Rollout footage can be captured live from a phone, so a hardware team does not need a dedicated instrumentation pipeline before it can start verifying policies. The provider, Most Robotic, is a San Francisco robotics research lab founded by MIT computer scientists Claire Mao and Lucy Cai and backed by Y Combinator's Summer 2026 batch.
Source: instancelabs.ai
Robotics teams today spend human hours watching rollouts, marking successes, and resetting scenes between runs, and that burden grows as robot foundation models scale to whole fleets and dozens of candidate policies.
Most Robotic positions Instance's verification layer as the substitute for that manual loop. Adjacent providers are converging on the same robot-learning evaluation budget from different directions: Fern Robotics Evaluation Platform through simulation, Scale AI's Physical AI data engine through annotation, and shotwell.ai's observability layer through monitoring.
Source: ycombinator.com
Instance verifies whether a robot completed a task by reading only the task description and camera footage of the rollout, and it returns a success verdict backed by grounded subtask evidence on any robot platform. The approach removes the human watch-and-label loop, and the verifier was benchmarked against more than ten thousand held-out, human-labeled episodes spanning seven robot platforms.
On the public benchmark table, the model's macro success-class F1 leads a zero-shot Claude Opus 4.8 baseline across all eight held-out test sets. It also runs on a single local GPU with no per-call cost and roughly two seconds of latency, where the paid API baseline sits above five seconds.
Source: demo.instancelabs.ai