ChromaDB + Authentication and Collections

Translated from the Spanish original. Read in Spanish

In this tutorial we’ll learn how to work with ChromaDB, a database for handling embedding vectors, with authentication and collection management.

iStock AI Generator

Setting up a ChromaDB server

Here, we’ll use Docker to deploy the server. To run a ChromaDB server with authentication, you must set the authentication variables correctly in the `docker-compose.yml` file. It’s essential not to store the token or secret key in plain text, for security reasons; this demonstration is only meant to illustrate how to go about the configuration.

version: "3.9"
services:
  chroma:
    image: ghcr.io/chroma-core/chroma:latest
    environment:
      CHROMA_SERVER_AUTHN_CREDENTIALS: "test-token"
      CHROMA_SERVER_AUTHN_PROVIDER: "chromadb.auth.token_authn.TokenAuthServerProvider"
    volumes:
      - index_data:/chroma/.chroma/index
    ports:
      - 8001:8000
    networks:
      - net
volumes:
  index_data:
    driver: local
  backups:
    driver: local

networks:
  net:
    driver: bridge

To start the ChromaDB server, run the following command:

docker-compose up

Parameters for the collection and the query

We define a few parameters we’ll use for the collection and the query:

EMBEDDINGS_MAX_RESULTS = 2
CHROMA_SERVER_AUTHN_CREDENTIALS = os.getenv('CHROMA_SERVER_AUTHN_CREDENTIALS')
CHROMA_CLIENT_AUTHN_PROVIDER = 'chromadb.auth.token_authn.TokenAuthClientProvider'
VECTOR_EMBEDDING_HOST = 'localhost'
directory = 'Zero/'

Loading and splitting documents

We create one function to load documents from a directory and another to split them into chunks:

Function to load documents from a directory

def load_docs(directory):
    loader = DirectoryLoader(directory, glob="**/*.txt", loader_cls=TextLoader)
    documents = loader.load()
    return documents

Function to split documents into chunks

def split_docs(documents, chunk_size=1000, chunk_overlap=20):
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=chunk_size, chunk_overlap=chunk_overlap)
    docs = text_splitter.split_documents(documents)
    return docs

Loading and splitting the documents

documents = load_docs(directory)
docs = split_docs(documents)

Generating Embeddings

We generate embeddings for the document chunks using a pre-trained model:

embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")

ChromaDB client and operations

To connect to the ChromaDB server and authenticate, we first need to create the client. In this example, we set the port and provide the required authentication credentials.

chroma_client = chromadb.HttpClient(
    port=8001,
    settings=Settings(
    chroma_client_auth_provider=CHROMA_CLIENT_AUTHN_PROVIDER,
    chroma_client_auth_credentials=CHROMA_SERVER_AUTHN_CREDENTIALS)
)

Creating or Retrieving a Collection

Here we’ll see how to check whether a collection exists and, if it doesn’t, how to create it.

collection = chroma_client.get_or_create_collection(name='myCollection')

Generating Embeddings for the Documents

We generate the embeddings for the documents, which will later let us search within the collection.

doc_embeddings = embeddings.embed_documents([doc.page_content for doc in docs])

Adding Documents and their Metadata to the Collection

Once the embeddings are generated, we add the documents and their metadata to the collection we just created or retrieved.

collection.add(
    ids = [str(uuid.uuid4()) for _ in docs],
    embeddings = doc_embeddings,
    documents = [doc.page_content for doc in docs],
    metadatas = [{'timestamp': timestamp, 'chapter': 'A', 'region': 'AMER', 'book': 'REGISTRY'} for _ in docs]
)

Running an Example Query

We run an example query: first we create the query embedding, then we use it to get results from the database.

embed_query = embeddings.embed_documents('The search query')

results = collection.query(
    query_embeddings = embed_query,
    n_results= EMBEDDINGS_MAX_RESULTS,
    where= {'$and': [{'chapter': 'A'}, {'region': 'AMER'}]},
)

Finally, we print the query results.

print(f"Results: {results}")

Complete Code

Below is the complete code used in this tutorial:

from langchain_community.document_loaders import DirectoryLoader, TextLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.embeddings import HuggingFaceEmbeddings
from dotenv import load_dotenv
import chromadb, uuid, os, datetime
from chromadb.config import Settings
load_dotenv()

EMBEDDINGS_MAX_RESULTS = 2
CHROMA_SERVER_AUTHN_CREDENTIALS = os.getenv('CHROMA_SERVER_AUTHN_CREDENTIALS')
CHROMA_CLIENT_AUTHN_PROVIDER = 'chromadb.auth.token_authn.TokenAuthClientProvider'
VECTOR_EMBEDDING_HOST = 'localhost'
directory = 'Zero/'

def load_docs(directory):
    loader = DirectoryLoader(directory, glob="**/*.txt", loader_cls=TextLoader)
    documents = loader.load()
    return documents

def split_docs(documents, chunk_size=1000, chunk_overlap=20):
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=chunk_size, chunk_overlap=chunk_overlap)
    docs = text_splitter.split_documents(documents)
    return docs

documents = load_docs(directory)
docs = split_docs(documents)
embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")

try:
    chroma_client = chromadb.HttpClient(
        port=8001,
        settings=Settings(
        chroma_client_auth_provider=CHROMA_CLIENT_AUTHN_PROVIDER,
        chroma_client_auth_credentials=CHROMA_SERVER_AUTHN_CREDENTIALS)
    )

    collection = chroma_client.get_or_create_collection(name='myCollection')

    timestamp = datetime.datetime.now().isoformat()
    doc_embeddings = embeddings.embed_documents([doc.page_content for doc in docs])

    collection.add(
        ids = [str(uuid.uuid4()) for _ in docs],
        embeddings = doc_embeddings,
        documents = [doc.page_content for doc in docs],
        metadatas = [{'timestamp': timestamp, 'chapter': 'A', 'region': 'AMER', 'book': 'REGISTRY'} for _ in docs]
    )

    embed_query = embeddings.embed_documents('The search query')
    results = collection.query(
        query_embeddings = embed_query,
        n_results= EMBEDDINGS_MAX_RESULTS,
        where= {'$and': [{'chapter': 'A'}, {'region': 'AMER'}]},
    )

    print(f"Results: {results}")
except Exception as error:
    print(f"Error: {error}")

Maximiliano Díaz Doglia

AI Platform Engineer & Full-Stack Developer
Building Enterprise Integrations & Automations

Published in: AI