Production AI starts with the dataset, not the model
3LC (Three Lines of Code) builds turnkey tooling for real-time debugging, diagnosis, and correction of data-related issues in computer-vision and ML pipelines. Paul’s mission line on the call was blunt: help companies understand the data they already have, how to use it, and what not to use.
That framing matters for industrial programmes where labeling and retraining budgets are real. Hosts mentioned order-of-magnitude training-time reductions and better true/false positive rates from data work in the intro - treat those as session claims to confirm with 3LC before using them as marketing facts.
When more labels make the model worse
Paul walked from a toy visual example into industrial-relevant cases. One competition / labeling story stuck: a model peaked at about 0.5% of the labeled set; adding more data made results worse. The takeaway is selection, not volume - most labels may never need to be created if you know which examples carry the signal.
As datasets explode for physical AI and robotics, he argued, intelligent data creation and selection matter more than the classic computer-vision scale of tens of thousands of images. The shortcut is fix and select before throwing a larger model at a noisy set.
Why this fits an industrial AI hour
June’s call paired data-centric ML with agent-built software for a reason: plants do not get a working system from either step alone. 3LC’s segment is the data side of that pipeline - cheaper labeling and training when you know what to keep.
Practitioners should leave with a sharper question than “which foundation model?”: which subset of your data is actually worth training on, and how fast can you see and correct the rest?


