Hybrid Retrieval: Combining Metadata and Vector Search

Introduction to Hybrid Retrieval

Welcome back! In the previous lesson, we explored the concept of multi-query expansion, which enhances search results by broadening the scope of a user's query. Today, we will delve into hybrid retrieval, a powerful technique that combines metadata and vector search to improve the accuracy and relevance of search results. This lesson will build on your existing knowledge and introduce you to the practical implementation of hybrid retrieval using ChromaDB.

Hybrid retrieval leverages the strengths of both metadata and vector search. Metadata provides structured information that can refine search queries, while vector search uses embeddings to understand the semantic meaning of text. By combining these approaches, we can achieve more precise and relevant search outcomes. Let's explore how this works in practice.

Understanding Metadata and Vector Search

Before we dive into the implementation, let's briefly revisit the concepts of metadata and vector search. Metadata refers to structured information that describes the content of a document, such as categories, tags, or author names. It allows us to filter and refine search queries based on specific attributes.

Vector search, on the other hand, uses embeddings to capture the semantic meaning of text. By representing text as vectors, we can measure the similarity between different pieces of text, enabling us to perform semantic searches that go beyond simple keyword matching.

Combining metadata and vector search allows us to leverage the strengths of both approaches. Metadata helps us narrow down the search space, while vector search ensures that the results are semantically relevant. This synergy is what makes hybrid retrieval a powerful tool in semantic search systems.

Setting Up Data and Collection

Before we dive into the implementation of hybrid retrieval, let's set up some sample data using ChromaDB. This will help us understand how the data is structured and how it can be queried.

from chromadb import Client
from chromadb.utils import embedding_functions
from data import load_documents

# Load documents
documents = load_documents("./data/corpus.json")
print(f"Loaded {len(documents)} documents.")

# Set up Chroma
model_name = "sentence-transformers/all-MiniLM-L6-v2"
embed_func = embedding_functions.SentenceTransformerEmbeddingFunction(model_name=model_name)

client = Client()
collection = client.get_or_create_collection("document_collection", embedding_function=embed_func)

# Batch add function
def batch_add_to_chroma(collection, documents, batch_size=50):
    print("Starting batch insert into ChromaDB...")
    for i in range(0, len(documents), batch_size):
        batch = documents[i:i + batch_size]
        print(f"Inserting batch {i} to {i + len(batch)}...")
        collection.add(
            documents=[doc["content"] for doc in batch],
            ids=[str(doc["id"]) for doc in batch],
            metadatas=[{
                "title": doc["title"],
                "category": doc.get("category", "unknown"),
                "tags": ",".join(doc.get("tags", [])),
                "date": doc.get("date", "unknown")
            } for doc in batch]
        )
    print("Finished inserting all batches.")

# Add data to Chroma
batch_add_to_chroma(collection, documents)

In this setup, we load documents from a JSON file and initialize a ChromaDB client. We create a collection named "document_collection" with an embedding function. We then add documents to the collection in batches, each with associated metadata such as title, category, tags, and date. This data will be used in our hybrid retrieval examples.

Sign up

Join the 1M+ learners on CodeSignal

Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignal