Deployment · Infrastructure · Production · 14 min read · October 8, 2026
AI Model Deployment: Serverless, Containers, Edge, and Choosing the Right Strategy
The Deployment Decision Tree
The best deployment strategy depends on three factors: latency requirements, request volume, and model size. Get these wrong and you're either overpaying for idle GPUs or losing users to slow responses.
Serverless Inference
Pay per request, zero infrastructure management. Best for:
- Low to medium traffic (under 100 requests/second)
- Spiky workloads (busy during business hours, dead at night)
- Small to medium models (under 5GB)
Platforms: AWS Lambda + SageMaker Serverless, Google Cloud Run, Modal, Replicate, Banana.dev. The trade-off is cold start latency — the first request after idle can take 5-30 seconds while the model loads into memory.
I deployed a text classification model on Modal for a client: $0.002 per request, zero maintenance, scales automatically. Perfect for their use case (500 requests/day with spikes to 2,000 during marketing campaigns).
Container-Based Deployment
Docker containers running on Kubernetes or ECS. Best for:
- Consistent traffic that justifies always-on instances
- Latency-sensitive applications (no cold starts)
- Complex serving pipelines (preprocessing → model → postprocessing)
Use a model serving framework inside the container:
- vLLM: For LLM inference — continuous batching, PagedAttention, best throughput per GPU dollar
- TorchServe: General-purpose PyTorch model serving
- Triton Inference Server: Multi-framework support, dynamic batching, GPU sharing between models
Edge Deployment
Running models directly on user devices or edge servers. Best for:
- Privacy-critical applications (data never leaves the device)
- Offline functionality requirements
- Ultra-low latency (no network round trip)
Challenges: model must be small enough for the target hardware, quantization is usually necessary, different devices need different optimizations (ARM vs x86, iOS vs Android).
Tools: ONNX Runtime (cross-platform), TensorFlow Lite (mobile/embedded), Core ML (Apple devices), TensorRT (NVIDIA edge devices).
GPU Provisioning and Cost
GPU costs dominate AI deployment budgets. Strategies to optimize:
- Right-size your GPU: An A10G handles most inference workloads. You don't need an A100 unless you're serving a 70B+ parameter model
- Use spot/preemptible instances: 60-70% savings for fault-tolerant workloads
- Batch requests: Dynamic batching groups multiple requests into one GPU forward pass. 4x throughput improvement is common
- Quantize: INT8 quantization halves model size and doubles throughput with minimal quality loss
- Share GPUs: Triton Inference Server can host multiple models on one GPU, each getting a time-slice
Monitoring in Production
Deployed models need monitoring beyond standard application metrics:
- Prediction latency: P50 and P99 latency per request
- Input data drift: Distribution of incoming features changing over time
- Output distribution shift: Model predictions skewing toward certain classes
- GPU utilization: Under 30% means you're overpaying. Over 90% means you need to scale
- Error rates: Invalid inputs, timeouts, OOM errors