Why Vector Databases Matter

Every AI application that needs to find relevant information — search, recommendation, RAG, deduplication — eventually needs vector search. Traditional databases search by exact matches or keywords. Vector databases search by meaning, finding content that's semantically similar even when the words are different.

The concept: convert text (or images, audio, code) into high-dimensional vectors using an embedding model. Store those vectors. When a query comes in, embed the query and find the nearest vectors. Those nearest vectors are your most relevant results.

The Options in 2026

Managed (Hosted)

  • Pinecone — the most mature managed option. Serverless pricing, automatic scaling, filtering, namespaces. Best for teams that don't want to manage infrastructure. Downsides: vendor lock-in, costs scale linearly with data
  • Weaviate Cloud — open-source core with a managed option. Built-in vectorization, GraphQL API, multi-modal support. More complex but more flexible than Pinecone
  • Qdrant Cloud — fast, Rust-based, good filtering. Open-source with a managed option. Growing quickly in the ecosystem

Self-Hosted

  • Chroma — simplest to start with. Python-native, runs embedded or client-server. Perfect for prototypes and small datasets (<1M vectors). Gets slow at scale
  • pgvector — PostgreSQL extension. Use your existing Postgres database. Best for teams already on Postgres who don't want another database. Performance is good up to ~5M vectors with proper indexing
  • Milvus — enterprise-grade, built for billions of vectors. Complex to operate but handles scale that others can't

Decision Framework

<100K vectors, prototype   → Chroma (embedded mode)
<1M vectors, Postgres shop → pgvector
<10M vectors, production   → Pinecone or Qdrant
10M+ vectors, custom needs → Milvus or Weaviate
Budget-constrained         → pgvector or self-hosted Qdrant

Embedding Models Matter More Than You Think

The vector database is just storage. The embedding model determines search quality. The right model can make a cheap database outperform an expensive one with a bad embedder.

  • OpenAI text-embedding-3-small — good default, cheap ($0.02/1M tokens), 1536 dimensions
  • OpenAI text-embedding-3-large — better quality, $0.13/1M tokens, configurable dimensions (256-3072)
  • Cohere embed-v3 — excellent for multilingual, comparable to OpenAI
  • BGE-M3 (self-hosted) — free, multilingual, runs on a single GPU. Quality rivals commercial options
  • Nomic Embed (self-hosted) — open-source, 8K context, very good for its size
Rule of thumb: always benchmark your embedding model against your actual data before committing. A model that scores well on benchmarks may perform poorly on your specific domain.

Production Gotchas

1. Chunking Strategy

How you split documents into chunks before embedding matters enormously. Too big and you lose precision. Too small and you lose context.

  • Start with 512 tokens per chunk with 50-token overlap
  • Respect document structure — split on paragraphs and headings, not character counts
  • Store metadata with each chunk (source document, section, page number) for filtering

2. Re-embedding Is Expensive

When you switch embedding models (and you will), you have to re-embed your entire dataset. Plan for this: store raw text alongside vectors so re-embedding is a batch job, not a crisis.

3. Hybrid Search

Pure vector search misses exact matches (product IDs, names, codes). The best production systems combine vector search with keyword search (BM25). Pinecone and Weaviate support this natively. For pgvector, combine with Postgres full-text search.

4. Filtering Before vs After

Pre-filtering (narrow the candidate set before vector search) is much faster than post-filtering (search everything, then filter results). Design your metadata schema around the filters you'll need.

Cost at Scale

A real example from a production RAG system I built:

  • 500K documents, average 3 chunks each = 1.5M vectors
  • Embedding cost: $18 one-time (OpenAI text-embedding-3-small)
  • Pinecone: ~$70/month (serverless, us-east-1)
  • Self-hosted Qdrant: ~$25/month (4GB RAM VPS)
  • pgvector: $0/month additional (existing Postgres instance)

The embedding cost is trivial. The ongoing database cost is where budgets matter. Self-hosting saves 60-70% but adds operational burden.

Related Articles

RAGEmbeddings

Building a RAG System That Actually Works

The complete architecture for retrieval-augmented generation that produces accurate, sourced answers.

Data PipelineAutomation

Building AI Data Pipelines That Run Themselves

Ingest, transform, enrich with LLMs, embed, and serve.

Cost OptimizationLLM

Your AI Bill Is 10x What It Should Be

Model routing, caching, and optimization that cut AI costs by 70-90%.