Two sensors, one clock
Video and inertial data are aligned in hardware, not stitched afterwards. Every clip ships with its measured sync offset so you can trust the timestamps you train on.
Human demonstration data | India
Zenyth Labs collects the data physical AI runs out of first: real people doing real tasks, filmed from the head and time-synced to inertial motion. Consent collected up front, faces removed, metadata your team can read before you spend a rupee.
We film the world so robots don’t have to guess.
Position
Most of this market is selling a roadmap: signed intentions, planned capacity, hours that do not exist yet. Zenyth Labs is built the other way round. Collection, consent and QA are designed first, so the moment a camera is on a head, every hour that ships already meets the bar your diligence team will check.
We are solving the boring half of the problem first, the half that stalls internal data programs: recruiting and training operators, getting commercial consent that a legal team will actually accept, blurring what has to be blurred, and holding a schema steady across thousands of hours so your loader does not need a special case per batch.
Video and inertial data are aligned in hardware, not stitched afterwards. Every clip ships with its measured sync offset so you can trust the timestamps you train on.
Bimanual work, deformable material, liquids, tool use, cluttered surfaces. The long tail that synthetic pipelines and staged studio capture keep missing.
Opt-in commercial training consent per operator, per session, referenced by ID in the metadata. Nothing reaches you that we cannot show provenance for.
Capability
Collection runs to a brief. You define the task list, object set, environment and schema; we recruit, capture, review and deliver. These are the environments already running.
ENV / INDUSTRIAL
Assembly, inspection, packing, sorting, hand tools, small-batch manufacturing lines.
ENV / DOMESTIC
Cooking, pouring, cutting, wiping, laundry folding, washing, wet and cluttered surfaces.
ENV / COMMERCE
Stocking, billing, bagging, product handling, counter workflows, dense shelf scenes.
ENV / TO BRIEF
Your task list, your objects, your camera geometry, your annotation schema. Quoted in days.
Specs
Every delivery is a directory, not a demo. Here is one clip from a household batch, opened up: what was captured, at what fidelity, with what rights, and how it passed QA.
| Clip ID | znl_kit_0193_a |
|---|---|
| Task | pour_liquid_from_kettle |
| Environment | Home kitchen, natural and mixed lighting |
| Duration | 42.6 s |
| Operator | Trained collector, pseudonymous ID retained |
| Hands in frame | Left and right, bimanual |
| Consent reference | cns_2f81c4 commercial training |
| Anonymization | Faces and screens blurred applied |
| Batch | household_2026_08 | 1,240 clips |
| Channel | Specification | Notes |
|---|---|---|
| Video | 3840 x 2160 | 30 fps | H.264, 8-bit, constant frame rate |
| Field of view | 108 deg | Head-mounted, wearer eye line |
| IMU | 200 Hz | 6-axis | 3-axis accelerometer, 3-axis gyroscope |
| Sync offset | 3.1 ms measured | Reported per clip in metadata |
| Audio | 48 kHz mono | Optional, muted by default on delivery |
| Second view | Exocentric, on request | Third-person camera, shared clock |
| Session length | 45 to 90 min | Split into task-scoped clips at QA |
One JSON sidecar per clip. Keys are stable across batches; extra keys are additive, never renamed.
{
"clip_id": "znl_kit_0193_a",
"task": "pour_liquid_from_kettle",
"environment": "home_kitchen",
"duration_s": 42.6,
"capture": {
"video": { "file": "znl_kit_0193_a.mp4", "fps": 30, "resolution": "3840x2160" },
"imu": { "file": "znl_kit_0193_a_imu.csv", "rate_hz": 200, "axes": 6 },
"sync_offset_ms": 3.1
},
"annotations": {
"hands": ["left", "right"],
"contact_frames": [412, 508, 1170],
"phases": [
{ "label": "reach", "t0": 1.2, "t1": 3.9 },
{ "label": "grasp", "t0": 3.9, "t1": 5.1 },
{ "label": "pour", "t0": 5.1, "t1": 19.4 },
{ "label": "place", "t0": 19.4, "t1": 22.0 }
]
},
"rights": {
"consent_id": "cns_2f81c4",
"commercial_training": true,
"anonymized": ["faces", "screens"]
},
"qa": { "sharpness": 0.94, "exposure_flags": 0, "sync_drift_ms": 2.4 }
}
Every clip is reviewed by a human against fixed gates before it enters a delivery. Failed clips are logged with a reason, not silently dropped.
Laplacian variance, normalised. Gate at 0.80.
Video to IMU, end of session. Bar shows headroom to the 10 ms gate.
Required keys present on every sidecar.
Share of captured clips that clear all gates.
| Formats | MP4 plus JSON by default. CSV, HDF5, WebDataset or your internal schema on request. |
|---|---|
| Transfer | Object storage bucket you own, or ours with scoped credentials. Checksummed manifest per batch. |
| Evaluation sample | 10 to 15 h, free |
| Pilot | 100 to 500 h | priced per hour, task dependent |
| Ongoing program | Monthly volume against a standing brief, with a named delivery owner. |
| Turnaround on a brief | 24 h to quote |
| Licence | Non-exclusive commercial training licence. Exclusivity negotiable per program. |
Values marked this way are representative and pending final numbers.
Pipeline
You send a brief. Everything between the brief and a loadable dataset is ours.
01
Tasks, objects, environments, camera geometry, volume, schema. We come back with a plan and a price.
02
Opt-in commercial AI training consent from every operator, recorded before a single frame is captured.
03
Trained collectors run sessions in real environments. Video and IMU on one clock, checked at the end of each session.
04
Faces, screens, plates, name tags and anything else your counsel flags, blurred before the data leaves us.
05
Task, phase, hand contact and timestamps. Human review against fixed gates, with rejects logged by reason.
06
Checksummed batch into your bucket, with a manifest, a QA report and the consent references.
How we prove it
Zenyth Labs is a new company, building this pipeline from first principles: no inherited footage, no borrowed numbers. What that buys you: nothing on this page is a projection dressed up as a fact. What ships is verifiable at the source, operator consent logged per session, QA gates documented per batch, full sensor provenance in every metadata sidecar. Inspect the actual schema below.
Anyone can promise data. Provenance is what ships.
Contact
Tell us what you are training and we will put together an evaluation package from existing footage, or quote a collection brief. Replies come from a founder, usually same day.