Briefly
Many teams trying to improve a model reach for architecture or compute first, but what has the most influence over the model is almost always the data.
A dataset can be entirely correctly labeled and still be a poor dataset to train on. Many datasets are highly redundant, or imbalanced in ways that only show up in one class. It can be annotated inconsistently at the edges, where the hard cases live, far larger than it needs to be, which is expensive in both ROI and GPU.
3LC tells you which of your training samples are hurting the model and how to act on those insights. Attached to the training run you were already going to do, it records what happens to every individual sample, ranks what is worth fixing, and then shows you that fixing it worked. This entire process stays inside your own infrastructure.
What 3LC does
It attaches to the training loop you already have
It is seamless to integrate, with no rewrite of your training loop, your model, or your dataloaders. It works with PyTorch, TensorFlow, Ultralytics YOLO, HuggingFace, SAM 3, VLAs and VLMs and any custom model you built yourself, because the usual cost of a data platform is the migration.
Per-sample metrics are the foundation
3LC captures metrics per sample, per epoch, across the whole run rather than just the aggregate.
Aggregate accuracy of 91% tells you nothing about the class sitting at 40%, or about the handful of samples the model spent every epoch failing to learn. Per-sample tracking finds both. In the Dashboard, you can lasso a cluster in embedding space where the model is struggling, filter live, fix labels in batch, reweight what matters, and commit the result as a sparse revision.
This runs across every computer vision task type: classification, object detection & oriented bounding boxes, pose estimation, semantic & instance segmentation, and LiDAR and 3D point clouds. Multiple modalities are supported as well; a robotics team, for example, collects data from many of these sources concurrently.
Insights ranks what to fix
Insights scores your data against your model, in order to surface the issues degrading it. These are ranked by severity: missing annotations, label errors, edge cases, class imbalance. Each one comes with a one-click path to fix it.
This process reveals that the diagnosis is tied to your model’s actual problems rather than to generic dataset statistics. It allows anyone on the team to act on it, rather than only the person who runs the training code.
Fixes are revisions, and revisions are comparable
Every edit creates a new dataset revision instead of overwriting the old one. Revisions are immutable and linked; knowing what was trained on is easily accessible.
Comparing two runs shows what changed in the data, what annotations were added, moved or deleted, alongside the per-class deltas those changes produced. The health score advances after each round of fixing, retraining, and comparing.
The Hub opens this to the rest of the team
3LC includes a no-code layer over this loop. Projects, experiments, datasets, training runs, and lineage exist within a browser tab. It connects to local drives, S3, Azure Blob and Google Cloud with versioning of data, and plugins arrive as hot-loaded and can be installed in one click with environments created automatically.
DriftCatcher, the production layer
DriftCatcher runs on models already deployed, by scoring the trust of every individual inference rather than random sampling or aggregating, catching distribution shift as it happens in the field, and auto-curating the samples worth retraining on. DriftCatcher determines, out of everything the model observes while in production, which of it is worth labeling.
This allows for production, curation, retraining, and deployment on one artifact. It is patent pending and not generally available yet, with close partners testing it out.
Where it runs
Most platforms in this category start by asking you to upload your data, but 3LC never does. Every deployment runs on your infrastructure, and no data reaches us.
- In your cloud account. The full stack inside your own tenancy on AWS, Azure or GCP. Single-tenant, so nothing runs on our infrastructure.
- On premise. Self-hosted on your own servers behind your firewall. Your VPC, your IAM, your network policies.
- Air-gapped. Fully operational without outbound network connectivity, in production today with classified workflows.
- Either centrally or per person. Deployed for the whole team, or on individual machines.
3LC is designed to adhere to GDPR, SOC 2 Type II, ISO 27001 and HIPAA, with SSO through SAML, role-based access, and encryption in transit and at rest.
Who 3LC is for
Teams working with robot or egocentric video. A sensor-rich robot produces tens of terabytes a day. A fleet produces petabytes. A classic computer vision dataset is a few gigabytes. You cannot store, label or train on all of it, and you should not want to, because only a small fraction carries anything new. The problem stops being collection and becomes selection.
Teams labeling continuous video rather than stills. Video is not just more data. Every clip is many frames with temporal context, which makes it far slower to annotate than an image. The cost has to be attacked from both directions at once: label fewer clips, and label each one faster.
Teams whose data cannot leave the building. Plant floors, defense programs, medical imaging, anything that’s under a data residency requirement.
Teams training small models for edge deployment. A smaller model needs a more relevant dataset, not a bigger one. Smart curation is what makes a small model viable.
Teams who have already cleaned their data and still want more accuracy. Verified-correct labels and optimal training data are different things.
Teams who need to explain why a model improved. Regulated industries, reporting, especially Enterprises that don’t accept the “we changed some things and it got better” narrative. Every model traces to a dataset revision, every revision to the specific human decisions inside it.
Teams training on synthetic data. More synthetic data is not reliably better data. The question is which generated samples improve the model, and that is answerable rather than a matter of faith.
Does it work
2026-03-18-20-27-14-181747 · 768f @ 30fpsOn public robotics data
In 2026 we ran the full curation loop on EgoVerse, a public egocentric human dataset built for robot learning by researchers at Georgia Tech, Stanford, UC San Diego, ETH Zürich, MIT, Meta Reality Labs, Mecka AI, and Scale AI. It holds roughly 1,360 hours of demonstrations across 1,965 tasks, 240 scenes and 2,087 demonstrators.
The method: the most informative clips are selected from the unlabeled pool, a vision-language model captions each segment, a human corrects only what is wrong, the model is fine-tuned on the corrected captions, and the next selection is sharper than the last. We ran five iterations across four tasks.
Across those five iterations:
- Human labeling time fell from 5 minutes 9 seconds to 52 seconds per minute of video, a reduction of 83%.
- Caption quality, scored with SODA against ground truth on both content and temporal order, rose from 0.241 to 0.744.
Most of the gain came early: after just the first iteration, five labeled examples per task, labeling time had already dropped to 3 minutes 12 seconds and SODA to 0.49.
SODA (Fujita et al., ECCV 2020) is an evaluation metric for dense video captioning that considers semantic content, temporal split, and temporal ordering. It scores captions on whether they tell the whole story with the right level of detail and in the right order, rather than grading each caption independently.
The result worth dwelling on is that a fine-tuned 8B model outperformed a 32B base model on the same clips. A smaller model, finetuned on curated data, beat a model four times its size trained on everything.
SODA scores the whole caption sequence vs ground truth (content + temporal order); per-caption similarity via SPICE.
Across industries, the data quality problem takes a different shape each time:
- Energy. Robotic and drone inspection at fleet scale, where data quality is paramount.
- Aerospace and defense. Less data, smaller models, sovereign by design, no cloud roundtrips.
- Robotics and physical AI. A robot fleet streams more data in a day than classic computer vision saw in a year. The work is finding the few thousand samples that improve the model.
- Agriculture. Field robots, sprayers and crop classification, where managing natural variation is what separates a trial from a product.
- Automotive. Multi-camera vision in every lighting condition, pinpointing the samples driving false positives.
- Infrastructure. Computer vision across thousands of edge points, where drift and domain shift are constant rather than exceptional.
Where we are going
With DriftCatcher, we’re working toward a fully automated data flywheel: selectively determining which data at the edge will teach the model something in the next iteration, feeding it into automatic labeling, verifying and automatically retraining the model, and using a new DriftCatcher deployment to guide the next iteration.
Trust has to be built into that loop. We want a clear record of where data came from, what was changed, how it was checked, and where uncertainty remains, with reporting that supports human oversight, and regulatory documentation as more of the process becomes automated.


