All work

CASE STUDY 01 / Building now

A common ground for robot models.

Standardized robotics benchmarks on real hardware.

  • Robotics
  • Evaluation
  • Real hardware
PROJECTCoop
ROLE / SCOPECo-founder
CONTEXTa16z speedrun · SR007
A shared testbed for robotics evaluation Conceptual line drawing of a robot arm moving a block on a gridded test surface. Different models use a shared task and protocol for comparable results. No real performance data is shown. FIG. 01 / SHARED TESTBEDCONCEPTUAL VIEW MODEL → ROBOTSHARED PROTOCOL ZX
Conceptual illustration of the work. No measured results or actual hardware setup shown.

Coop builds standardized, third-party robotics benchmarks so humanoid and manipulation model developers can run their models on the same tasks and get comparable scores.

Results need a shared context.

Robotics labs use different robots, setups, and task definitions. That makes it difficult to understand how model results compare across teams.

Same tasks. Comparable scores.

Coop gives model developers a shared set of tasks on real hardware. Results are organized by robot, model, and task, keeping the protocol, scores, and failures useful as the benchmark grows.

Building the evaluation infrastructure.

My focus is building Coop as a co-founder. Coop is backed by a16z speedrun, SR007. This overview covers our public positioning; customer information, benchmark results, and internal implementation details are not included.

Talk about robotics evaluation
NEXT CASE STUDYBCI & assistive robotics