Local AI vs Cloud API: When to Self-Host and When to Pay
Every week someone asks me whether they should run their own LLM or just pay OpenAI. The answer is never one or the other. It is always both — but the ratio depends on your workload, your volume, and how much you value keeping data on your own network. I have been running a hybrid stack for over a year now, and I have the numbers to prove exactly where each approach wins.
This is not theory. I run production workloads on a self-hosted Qwen 27B model that costs me $0/month in API fees, and I route complex reasoning tasks to Claude and GPT-4 via cloud API. Here is the complete breakdown.
The Real Cost Comparison
Cloud AI pricing is simple: you pay per token. Self-hosted pricing is complicated: you pay for hardware, electricity, your time, and the opportunity cost of that GPU doing nothing when your workload is idle. Most "self-hosting is free" articles ignore half of these costs. Most "just use the API" articles ignore what happens when you scale past 10,000 calls a day.
Cloud API Costs (October 2026)
- Claude Sonnet 4: $3 / 1M input tokens, $15 / 1M output tokens
- Claude Opus 4: $15 / 1M input, $75 / 1M output
- GPT-4o: $2.50 / 1M input, $10 / 1M output
- GPT-4o-mini: $0.15 / 1M input, $0.60 / 1M output
- OpenRouter (Qwen 27B hosted): ~$0.15 / 1M input, ~$0.15 / 1M output
A typical business chatbot interaction is about 800 input tokens and 400 output tokens. At Sonnet 4 rates, that is $0.0084 per conversation. Run 1,000 conversations a day and you are looking at $252/month. Run 5,000 and it is $1,260/month.
Self-Hosted Costs
- RTX 3060 12GB (used): ~$200 one-time
- Tesla M40 24GB (used): ~$80 one-time
- Electricity: ~$15-20/month (system draws ~200W at load, Ohio rates ~$0.12/kWh)
- Maintenance time: ~2 hours/month (model updates, monitoring, occasional restart)
- Total first year: ~$500 hardware + ~$200 electricity = ~$700
- Each subsequent year: ~$200 electricity only
At cloud API rates for a comparable open-source model ($0.15/1M tokens through OpenRouter), you break even on self-hosting at roughly 4.7 million tokens per month — about 150,000 tokens per day. That is approximately 300 chatbot conversations or 50 long-form content generation runs. If you are doing less than that, the API is cheaper. If you are doing more, self-hosting pays for itself within months.
The Break-Even Analysis
I built a calculator for this. The math depends on which cloud model you are replacing:
# Break-even: self-hosted vs cloud API
# Hardware: $280 GPU + $200/yr electricity = $480 first year
# vs GPT-4o-mini ($0.60/M output tokens)
# 480 / 0.0006 = 800M output tokens/yr to break even
# = 2.19M tokens/day = NOT worth self-hosting for mini-tier
# vs Claude Sonnet ($15/M output tokens)
# 480 / 0.015 = 32M output tokens/yr to break even
# = 87,671 tokens/day = ~175 conversations/day
# = VERY achievable for any production system
# vs Claude Opus ($75/M output tokens)
# 480 / 0.075 = 6.4M output tokens/yr
# = 17,534 tokens/day = ~35 conversations/day
# = Self-hosting wins almost immediately
The key insight: you are not replacing Opus or Sonnet with your self-hosted model. You are replacing the tasks that do not need Opus or Sonnet. The question is not "which is cheaper per token" — it is "which tasks can a 27B open-source model handle just as well?"
Hardware Options: What Actually Works
Consumer GPUs
- RTX 3060 12GB (~$200 used): My primary inference card. Runs Qwen 27B at Q2_K_P quantization at 20 tok/s. 12GB is tight but workable for models up to ~13B at Q4 or ~27B at Q2. Best price-to-performance for inference.
- RTX 4090 24GB (~$1,400 used): The gold standard for home inference. 24GB fits 27B at Q4 comfortably. 2-3x faster generation than a 3060. If you can afford it, this is the card to buy.
- RTX 3090 24GB (~$700 used): Same 24GB as the 4090 at half the price, but slower. Good middle ground.
Datacenter Cards (Used)
- Tesla M40 24GB (~$80): Incredible VRAM-per-dollar, but Maxwell architecture (2015) means no FP16, slow compute, and GDDR5 bandwidth. I use one for KV cache overflow and secondary workloads. Do not expect fast inference.
- Tesla P40 24GB (~$120): Pascal architecture, marginal improvement over M40. Still no FP16 tensor cores.
- A100 40GB (~$5,000-8,000 used): The real deal. If you are running a business on local inference and need throughput, this is worth the investment. 40-80GB HBM2e memory, 2TB/s bandwidth.
Cloud GPU Rental
- vast.ai: Spot instances from $0.15/hr for an RTX 3090. Great for burst workloads or experimentation. Not reliable enough for 24/7 production.
- RunPod: More reliable than vast.ai, slightly higher prices. Good for production-grade serverless inference with auto-scaling.
- Lambda Labs: A100 instances at $1.10/hr. Best for training and heavy batch inference, not cost-effective for always-on serving.
My recommendation for most people starting out: buy a used RTX 3060 12GB for $200 and run llama.cpp. If your workload grows, upgrade to a 3090 or 4090. Do not start with datacenter cards unless you already have the infrastructure to power and cool them.
The Quality Gap: When It Matters
This is the part most self-hosting enthusiasts gloss over. There is a real quality difference between frontier models and open-source alternatives, and pretending otherwise will cost you customers.
Where Open-Source Wins (or Ties)
- Text classification: Sentiment analysis, intent detection, category tagging. A fine-tuned 7B model beats GPT-4 at these tasks because you can train it on your exact categories.
- Structured extraction: Pulling names, dates, prices, addresses out of unstructured text. Qwen 27B handles this reliably with good prompting.
- Embeddings: Sentence-transformers models run locally at thousands of documents per second. No reason to pay for cloud embeddings.
- Simple Q&A: FAQ bots, knowledge base search, straightforward information retrieval over your own documents.
- Code generation: For boilerplate, test scaffolding, and standard patterns, Qwen 27B and Llama 3 70B are very competent.
Where Cloud APIs Still Win
- Complex multi-step reasoning: Tax calculations, legal analysis, intricate business logic. Claude Opus and GPT-4 are measurably better at chaining reasoning steps without losing context.
- Persuasive writing: Sales copy, objection handling, nuanced customer communication. The AI sales agent I built uses Claude because the quality difference in natural conversation is immediately noticeable.
- Long-context synthesis: Summarizing a 50-page contract, analyzing a full codebase, comparing multiple documents simultaneously. Frontier models handle 100K+ context more reliably.
- Safety and alignment: Cloud APIs have extensive RLHF training and safety filters. For customer-facing applications, this matters.
Latency: The Hidden Advantage of Local
Cloud APIs have faster raw generation speed — an A100 cluster generates tokens faster than a single 3060. But latency is not just generation speed. It includes:
- Network round-trip: 50-200ms per request, depending on your location and the provider
- Cold start: Serverless endpoints (like OpenAI) can add 500ms-2s on the first request after idle
- Queue time: During peak hours, cloud providers throttle or queue requests. I have seen 5-10 second delays on Claude during high-traffic periods
- Rate limits: Hit your rate limit and you are waiting, or paying for a higher tier
Local inference has none of these problems. My llama.cpp server responds in under 50ms to first token, every time, with zero network overhead. For latency-sensitive applications like real-time chat or voice AI pipelines, this matters more than raw throughput.
# Latency comparison (measured, not theoretical)
# Task: 200-token prompt, 100-token response
# Local (Qwen 27B on 3060, llama.cpp)
# Time to first token: 40ms
# Total generation: 5.0s (20 tok/s)
# Total latency: 5.04s
# Cloud (Claude Sonnet via API, from Ohio)
# Network + queue: 180ms
# Time to first token: 320ms
# Total generation: 1.8s (~55 tok/s)
# Total latency: 2.3s
# Cloud wins on short responses. Local wins on reliability.
Privacy and Compliance
This is the easiest decision in the whole analysis. If your data cannot leave your network, self-host. Period.
- HIPAA: Patient data processed by a cloud API means you need a BAA with the provider. OpenAI offers this on Enterprise plans ($$$). Or you run inference locally and the data never leaves your server.
- SOC 2: Your compliance auditor will have a much easier time if AI processing stays within your existing security boundary.
- Client confidentiality: Law firms, financial advisors, HR departments — any industry where sending client data to a third-party API creates liability.
- Proprietary data: Trade secrets, internal strategy documents, competitive intelligence. Do you really want your M&A analysis going through OpenAI's servers?
I have had three clients choose self-hosted inference purely for compliance reasons, even though the cloud API would have been cheaper and higher quality. When legal says the data stays on-prem, the data stays on-prem.
The Hybrid Approach I Actually Use
Here is exactly how I split workloads across local and cloud:
Self-Hosted (Qwen 27B, $0/month)
- Content generation for blogs, social media, marketing copy
- Script generation for video production pipelines
- Data extraction and classification from documents
- Embeddings for RAG systems (sentence-transformers, also local)
- Internal development assistance and code review
- Batch processing: nightly report generation, inventory analysis
Cloud API (Claude/GPT-4, ~$50-150/month)
- Customer-facing sales chat (Rick, the AI salesman)
- Complex reasoning: pricing analysis, competitive research synthesis
- Voice AI agent responses where nuance and tone matter
- Escalation handling when the local model is uncertain
The Router Pattern
I use a simple routing layer that decides which backend handles each request:
def route_request(task_type, complexity, customer_facing):
# High-stakes, customer-facing = cloud API
if customer_facing and complexity == "high":
return "claude-sonnet"
# Complex reasoning regardless of audience
if complexity == "high":
return "claude-sonnet"
# Customer-facing but simple (FAQ, status checks)
if customer_facing and complexity == "low":
return "local-qwen-27b"
# Internal or batch = always local
return "local-qwen-27b"
This routing cuts my cloud API bill by roughly 70% compared to sending everything to Claude. The local model handles about 80% of all requests by volume, and the quality is indistinguishable from cloud for those task types.
Real Numbers From My Setup
Here is a typical month of inference across my production systems:
# September 2026 inference stats
# ================================
# Local (Qwen 27B on 3060 + M40)
# Requests: 18,400
# Tokens generated: 7.2M
# Uptime: 99.7% (2h downtime for model update)
# Cost: $18 electricity
# Equivalent cloud: $1,080 (at Sonnet rates)
#
# Cloud (Claude Sonnet + Opus via API)
# Requests: 4,200
# Tokens generated: 1.8M
# Cost: $89
#
# Total monthly AI cost: $107
# Without local: $1,169
# Savings: $1,062/month (91%)
Operational Overhead: The Real Cost
Self-hosting is not free in terms of time. Here is what maintenance actually looks like:
- Model updates: New quantizations and model releases every few weeks. Testing a new model takes 1-2 hours. I do this monthly, not for every release.
- Monitoring: A simple health check endpoint that pings the llama.cpp server every 60 seconds. If it is down, I get a notification via ntfy. Setup took 30 minutes, runs unattended.
- Failover: If the local server goes down, requests automatically route to cloud API. This has happened twice in six months. Both times the systemd service restarted automatically within 10 seconds.
- CUDA/driver updates: I update GPU drivers maybe twice a year. Each update takes 20 minutes including a reboot.
- Hardware failures: Zero so far. Consumer GPUs are remarkably reliable when not overclocked. The M40, a 10-year-old datacenter card, has run continuously without issue.
Total time investment: about 2-3 hours per month. At any reasonable hourly rate, that is still far cheaper than the $1,000+ in API savings.
Decision Framework: Four Questions
Before you build anything, answer these four questions:
- What is your monthly token volume? Under 5M tokens/month — just use cloud APIs. Over 5M — self-hosting starts making financial sense. Over 20M — self-hosting is almost certainly cheaper.
- Does your data need to stay on-prem? If yes, self-host. There is no cloud workaround for a compliance requirement.
- Do your tasks require frontier-model quality? If every request needs Claude Opus-level reasoning, self-hosting a 27B model will not cut it. If most requests are classification, extraction, or templated generation, a local model is fine.
- Do you have (or want) the infrastructure? If you already run a homelab or have a server closet, adding a GPU is trivial. If you are starting from zero, factor in the learning curve and the $500-1,500 in hardware before you see any savings.
The right answer for most businesses is: start with cloud APIs to validate your use case, measure your actual volume and quality requirements, then migrate the high-volume commodity tasks to self-hosted once you have real data. Do not over-engineer on day one.
Want Help Building a Hybrid AI Stack?
I design and deploy hybrid inference architectures — self-hosted models for the bulk of your workload, cloud APIs for the tasks that need frontier quality, and a routing layer that makes the decision automatically. The result is typically a 60-90% reduction in AI costs with no loss in output quality for end users.
Let's talk about your AI infrastructure
Related Articles
How I Run a 27B Parameter LLM on Consumer GPUs for $0/Month
Running Qwen 27B on a Tesla M40 and RTX 3060 in a home lab. Real numbers, real gotchas, and why self-hosted inference makes sense for production AI.
Building a Multi-GPU Inference Stack on Consumer Hardware
How to split models across multiple GPUs, when tensor parallelism helps vs hurts, and the PCIe topology pitfalls nobody warns you about.
AI Cost Optimization: Cutting Your API Bill by 80%
Practical strategies for reducing AI inference costs without sacrificing quality — caching, routing, batching, and knowing when to downgrade models.