Every supervised learning model needs labeled data, and labeling is where most AI projects stall. The model is the easy part — collecting, cleaning, and annotating training data at sufficient quality and volume is where the real work lives. I've seen projects with sophisticated architectures fail because their training data was garbage.
Self-hosted, highly customizable, supports every annotation type. I use this for most projects because you own the data and can customize the annotation interface. Setup: Docker container, connect your storage (S3, GCS, local), configure the labeling template.
When you need thousands of annotations fast and have budget. Their workforce handles the labeling — you define the task, review samples, and iterate on quality. Good for object detection and image segmentation at scale. Expensive for text tasks.
Purpose-built for NLP annotation. Its active learning loop is the killer feature: the model suggests labels, you confirm or correct, the model improves, better suggestions follow. One annotator with Prodigy can match the output of 3-4 using traditional tools.
Bad labels produce bad models. Quality control is non-negotiable:
Active learning asks: which unlabeled examples would be most valuable to label next? Instead of randomly selecting items to annotate, the model identifies samples where it's most uncertain. This reduces the total labels needed by 30-60% compared to random sampling.
Implementation: train an initial model on a small labeled set → predict on unlabeled data → sort by uncertainty → send the most uncertain examples to annotators → retrain → repeat.
Synthetic data (generated by another model or by augmentation) works in specific scenarios:
Synthetic data is a supplement, not a replacement. Models trained exclusively on synthetic data perform worse than those trained on real data. The sweet spot is real data augmented with synthetics.