Introduction

Hello there, welcome to the second lesson of our "Scaling Up RAG with Vector Databases" course! In the previous unit, you explored how to break large documents into smaller chunks and attach useful metadata (like doc_id, chunk_id, and labels such as category). These chunks are essential for structuring data in a way that makes retrieval easier. In this lesson, we'll build on that groundwork by showing you how to store them in a vector database. One popular choice is ChromaDB — a specialized, open-source database designed for high-speed, semantic querying of vectors. By switching from keyword-based searches to semantic searches, your RAG system will retrieve relevant information more efficiently. Let's dive in!

Understanding Vector Databases

A vector database stores data in the form of numerical vectors that capture the semantic essence of texts (or other data). The database then uses similarity metrics — rather than literal word matches — so that conceptually similar items are stored close together. This means searches on vector databases can retrieve contextually relevant results even when keywords are absent. By leveraging approximate or exact nearest-neighbor strategies for similarity, vector databases can scale to handle millions or billions of vectors while still providing quick query responses. This makes them especially suitable for RAG systems, which rely on fast semantic lookups across large collections of text.

Why We Need Vector Databases for RAG

Before we explore how to set up a vector database, let's look at why it's a crucial component of a RAG pipeline:

  1. Semantic Retrieval: By embedding text into vectors, queries can match documents based on meaning rather than strict keyword matches. This yields more accurate and context-sensitive search results.
  2. Scalability: Specialized vector databases handle large datasets efficiently, allowing you to store and query vast libraries of text chunks without sacrificing performance.
  3. Richer Context: Embeddings capture nuanced relationships among chunks, ensuring that related information is surfaced even when it doesn't use the exact same terms.
  4. Easy Updates: Vector databases (like ChromaDB) often allow you to add and remove chunks on the fly, so your collection stays in sync with new or evolving information.
Setting Up ChromaDB and Basic Configuration

Now, let's jump into coding with ChromaDB, our chosen vector database. Here's how to set up a ChromaDB client using JavaScript:

import fs from 'fs';
import path from 'path';
import { ChromaClient, OpenAIEmbeddingFunction } from 'chromadb';

async function buildChroma // Use an OpenAI model for embeddings
    const embedder = new OpenAIEmbeddingFunction({
        model_name: "text-embedding-ada-002"
    });

    // Create a ChromaDB client with a specified path
    const chroma = new ChromaClient({ path: "http://localhost:8000" });

    // Check if the collection already exists or create a new one
    const collectionName = "rag_collection";
    const collections = await chroma.listCollections();
    let collection;
    if (collections.find((coll) => coll === collectionName)) {
        console.log(`Collection "${collectionName}" already exists. Retrieving it...`);
        collection = await chroma.getCollection({ name: collectionName });
        if (!collection.embeddingFunction) {
            collection.embeddingFunction = embedder;
        }
    } else {
        console.log(`Collection "${collectionName}" does not exist. Creating collection...`);
        collection = await chroma.createCollection({
            name: collectionName,
            embeddingFunction: embedder,
        });
    }

    // ... continues
}

How It Works:

  • Embedding Setup: We define an OpenAIEmbeddingFunction to generate vectors for the text chunks. The model we're using, text-embedding-ada-002, is a powerful model that maps sentences to a dense vector space. It's popular for RAG applications because it balances efficiency with strong semantic understanding capabilities.
  • Client Configuration: new ChromaClient({ path: "http://localhost:8000" }) connects to ChromaDB with a specified path. This setup allows for a persistent connection to a ChromaDB server running locally or remotely.
  • Collection Management: We check if a collection named "rag_collection" exists; if not, we create a new one. It's worth noting that listCollections() returns only the collection names, not their full metadata. So when checking for existence with find((coll) => coll === collectionName), ensure that you're comparing against strings — not objects — to avoid mismatches. This check avoids accidentally creating duplicates and keeps your namespace clean. A collection in ChromaDB is a logical container that groups related documents and their embeddings together, similar to a table in a traditional database but optimized for vector similarity operations. Collections allow you to organize your vector data into separate namespaces, making it possible to maintain multiple distinct sets of documents with different embedding models or for different use cases.
Preparing Data and Adding Chunks to ChromaDB
Updating and Managing Documents
Conclusion and Next Steps

By storing text chunks in a vector database, you've laid the foundation for faster, more semantically aware retrieval. You know how to create, update, and manage a ChromaDB collection — crucial skills for any large-scale RAG system.

In the next lesson, you'll learn how to query the vector database to fetch the most relevant chunks and feed them into a language model. That's where the real magic of producing context-rich, accurate responses shines! For now, feel free to explore different embedding models or try adding and deleting a variety of chunks. When you're ready, proceed to the practice exercises to cement these concepts and further refine your RAG workflow.

Sign up
Join the 1M+ learners on CodeSignal
Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignal