// REAL CAPTURE · NOT A RENDER
One clip.
UMI right gripper-cam · fisheye · bottle → box
// the delivered view · 6 task families · full fisheye · 6-DoF tracked · Kav-Lab
The action layer.
Language models had the whole internet to learn from. Robots don’t — the web is full of video that shows what happened, but almost none of it carries the actions that produced it. Kav-Lab captures that missing layer: real-world activity and manipulation, tracked to clean action data — the fuel robot policies actually train on.
- 2.7K / 60fps multi-view — synchronized ego + gripper fisheye
- 6-DoF pose + gripper-width → clean, cm-level action labels
- LeRobot-format, GR00T-compatible — QA’d on device
Volume to pre-train. Quality to post-train.
Robot foundation models need two very different diets. Pre-training wants scale and diversity — huge amounts of real human activity across every setting. Post-training wants precision — clean, task-specific demonstrations with exact action labels. Kav-Lab delivers both.
An egocentric collection app.
Puts capture in anyone’s hands — turning operators into a distributed collection network that gathers diverse, first-person activity data at volume, across every setting a robot might one day work in.
- ▹First-person (egocentric) — the view frontier VLAs pre-train on
- ▹Every sector, one pipeline
- ▹The diverse scale foundation models are hungry for
QA’d manipulation capture.
UMI-style rigs record manipulation with 6-DoF pose tracking and on-device quality gates — clean, action-labeled state–action trajectories a policy can fine-tune on, task by task.
- ▹6-DoF end-effector pose + gripper width
- ▹On-device QA & anti-fraud gates before it ships
- ▹Task-specific demos — the diet for fine-tuning
Built to scale to a million hours.
The egocentric app puts capture in anyone’s pocket — a distributed collection network designed to gather first-person activity at million-hour scale, across every sector a robot might one day work in. Every clip flows straight into an automated pipeline and comes out refined, formatted, labelled, and annotated — delivered on time, hassle-free.
// one hands-free path: the operator records — the pipeline does the rest, and the lab gets a dataset.
Every way in.
From a head-cam in the wild to a millimetre-accurate bimanual rig — seven capture modes, each tuned to a different trade between reach, precision, and diversity. They all resolve to the same clean 6-DoF action format.
One clean record. Every stream, synced.
Whatever the capture mode, each episode ships the same way — every stream timestamped to a common clock, undistorted, and calibrated.
The label stack keeps growing.
Every release adds another signal to the ego/UMI episode, beyond pose and gripper width. Here’s what’s shipping, what’s in validation, and what’s next — a channel is only marked built once it’s validated on real data.
See every episode move.
Ops watch the pipeline the way a lab would — throughput at every stage, server QA decisions, the Failure Atlas, and deliveries, all in one place.
// illustrative dashboard — synthetic sample figures, not live numbers. The surfaces shown (RBAC ops dashboards, taxonomy + benchmark-split gate, Failure Atlas) are built and tested.
Load it in one command. Verify it in one more.
Audit-ready, every time.
Every set exports in the exact LeRobot v2.1 layout NVIDIA Isaac-GR00T and LeRobot consume — modality.json, per-data stats (never hardcoded), parquet + video frame-aligned. Then it’s wrapped in a signed bundle: a lab runs the SHA256SUMS check and the load_test.ipynb, and the data loads and audits on the first try.
- GR00T LeRobot v2.1 exporter with modality.json — stats computed from data
- Full lab-bundle packager — consent, units, calibration + MANIFEST + SHA256SUMS
- Datacard + load_test.ipynb — auditable and loadable out of the box
From demonstration to policy.
A robot learns in two moves. Imitation learning clones demonstrations into a base policy; reinforcement learning then refines it against a reward — and can even learn offline, straight from logged trajectories. In both, the demonstrations are what make real-world learning tractable — they pretrain the model, seed exploration, and bootstrap RL that’s otherwise far too sample-hungry on real hardware.
For a given task, generalization scales with the range of environments and objects seen — a power law — not raw demo count. Which is why a field-collection engine across every sector wins. — Data Scaling Laws in Imitation Learning, ICLR 2025
Treated loops (grade · vignette · grain · HUD · © watermark) from raw 2.7K footage. All footage © 2026 Kavron.ai.