Skip to content
Proudly based in Nova Scotia, Canada · clients welcome from every countryContact usClient login
CodeLumaDevelopment Inc.

Home / Blog / Article

CodeLuma insights · April 11, 2026 · 6 min read

Vector Databases Explained: Powering Semantic Search

A plain-language guide to embeddings, vector databases and similarity search: how they work, when you need one, indexing, filtering and costs.

Almost every recent AI feature that "knows" about your documents, products or customer history relies on the same underlying technique: turning content into vectors and finding the nearest ones. The storage layer for this is called a vector database, and the term has attracted a great deal of hype. This article explains, without mathematics, what vectors and embeddings are, what a vector database does, when you really need one, and how to build retrieval that works reliably.

The same embedding model must be used for indexing content and for queries.
The same embedding model must be used for indexing content and for queries.

Embeddings in plain language

An embedding model reads a piece of text, an image or audio, and outputs a list of numbers, typically several hundred to a few thousand, which represent its meaning. Content with similar meaning produces vectors that point in similar directions. "How do I reset my password?" and "I forgot my login credentials" end up close together, even though they share almost no words. You can measure closeness with a mathematical distance, such as cosine similarity, which is fast to compute. That is the entire trick: convert everything to vectors once, then compare.

A laptop showing source code on a desk
Original CodeLuma 3D render: a laptop showing source code on a desk.

What a vector database does

Comparing a query vector with millions of stored vectors one by one is too slow, so vector databases build specialised indexes for approximate nearest neighbour search, trading a tiny amount of accuracy for large speed-ups. Popular index types include graph-based structures such as HNSW, and clustering approaches such as IVF, sometimes combined with compression called quantisation to reduce memory. Beyond raw search, a vector database stores metadata alongside each vector (source, owner, date, category, permissions), supports filtering by that metadata, handles inserts, updates and deletes, and scales and replicates like other data stores.

What you use it for

  • Semantic search: find documents by meaning, not exact words; see our AI search guide.
  • Retrieval-augmented generation (RAG): before asking a language model a question, retrieve the most relevant passages from your knowledge base and include them in the prompt so answers are grounded in your data instead of guesses; see our guide to generative AI integration.
  • Recommendations and similarity: "more like this" for products, articles or images.
  • Deduplication and clustering: find near-duplicate records or group similar support tickets.
  • AI memory: store summaries of past interactions for assistants, retrieving relevant memories for each new conversation.
  • Classification by example: assign categories based on similarity to labelled examples.

Do you actually need a dedicated vector database?

Perhaps not. If your data is up to a few million vectors, a vector extension for the relational database you already run may be entirely adequate, and it lets you join vectors with normal data and keep transactions, backups and permissions in one place. Search engines increasingly offer vector fields and hybrid queries, which is valuable when you also need keyword matching. Dedicated vector databases earn their place at very large scale, when you need advanced filtering performance, multi-tenant isolation features or very low latency under heavy load. Choose based on scale, filtering needs, operations capacity and cost, using the table above, and avoid adding infrastructure your team must learn and secure without a clear benefit.

Many products never need more than a vector extension in the database they already run.
Many products never need more than a vector extension in the database they already run.

Chunking: the underrated craft

Documents must be split into chunks before embedding, because models have input limits and because a vector for a whole long document blurs its meaning. Good chunking respects structure: split by headings, paragraphs and sections, keep chunks in the range of a few hundred tokens, use modest overlap so ideas are not cut in half, and attach context such as the document title and section heading to each chunk. Keep metadata to trace every chunk back to its source, so you can cite it. For tables, code and PDFs, use extraction techniques that preserve structure. Poor chunking is the most common reason retrieval-based assistants disappoint.

Choosing an embedding model

Models differ in quality, language coverage, vector size, speed, cost and whether they run through an API or on your own servers. Test candidates on your own data with real queries rather than trusting leaderboards. Larger vectors improve nuance but increase storage and query cost. Multilingual content needs a multilingual model. Whatever you choose, record the model name and version with each vector, because changing models means re-embedding everything, and vectors from different models are not comparable. Plan for that migration.

Filtering and permissions

Real applications need "only search this customer's documents" or "only articles the user may read." Store access metadata with each chunk and apply filters within the search, not after it, so results are both correct and fast. Never rely on the language model to keep secrets: enforce permissions in retrieval so restricted text never enters the prompt. Multi-tenant designs should isolate tenants strictly; the patterns in our white-label article apply. Audit access as in our audit logging article.

A two-monitor developer workstation
Original CodeLuma 3D render: a two-monitor developer workstation.

Improving retrieval quality

  • Hybrid search that merges keyword and vector results, for identifiers and rare terms.
  • Re-ranking the top candidates with a stronger model.
  • Query rewriting to turn conversational questions into good search queries.
  • Metadata boosts for freshness and authority.
  • Evaluation sets of real questions with known correct sources, measuring whether the right chunk appears in the top results. Improve retrieval first; a better prompt cannot fix missing context.

Keeping the index fresh

Content changes. Re-embed only changed items, delete vectors for removed content, and schedule reconciliation jobs that compare the source of truth to the index. Use queues for bulk embedding, since it is rate limited and costs money; see our background jobs article. Version the index so you can rebuild in the background and swap without downtime.

Costs and operations

Costs come from embedding generation (per token, mostly one-off plus updates), storage (vectors of a thousand or more numbers each add up), memory for fast indexes, and query compute. Estimate before committing, monitor latency and recall, and back up the index or, better, be able to rebuild it from the source data. See our backup guide and our monitoring guide.

Pitfalls

  • Treating a vector database as a magic knowledge store: garbage in, confident garbage out.
  • Skipping evaluation and tuning by intuition.
  • Embedding sensitive data with a provider without checking data-handling terms.
  • Ignoring updates and deletions, so stale or removed content keeps appearing.
  • Over-engineering scale for a dataset that would fit in a single table.

Start small: embed a well-chosen collection, store it in the simplest suitable place, evaluate with real questions, and grow from there. Our software team builds retrieval systems and AI assistants over your own content, and our hosting and maintenance services operate the infrastructure securely.


Put this into practice with CodeLuma

CodeLuma designs retrieval layers, from a vector extension on your existing database to a dedicated store, with chunking, permissions and evaluation done properly so AI features answer from your data reliably.

Start a conversation. Tell us about your project and we will reply with practical next steps, or browse all CodeLuma services. CodeLuma Development Inc. is based in Nova Scotia and works with teams across Canada and remotely.

Keep reading

Share this article: Facebook · LinkedIn · X · Email

← All articles

Ready to put this into practice?

Talk to a Nova Scotia full-stack team that builds complex, connected systems for clients across Canada and worldwide.

Start a project