AI bot with documents and memory
Translated from the Spanish original. Read in Spanish
This tutorial shows how to set up an AI bot with documents and memory, using Azure OpenAI and LangChain.

Setting up the credentials
Before you start, make sure you have a .env file in the same directory as your code with the following information, used to authenticate against the Azure services:
AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_SECRET = "xxxxx"
In the code, we load the environment variables and set the credentials using Azure’s ChainedTokenCredential and EnvironmentCredential:
import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from dotenv import load_dotenv
load_dotenv()
credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")
Setting up the model and the embedding
We define the model and the embedding we’ll use, then set the Azure environment variables needed to work with the OpenAI API:
# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"
# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment
Initialising the OpenAIEmbeddings class
The AzureOpenAIEmbeddings class is initialised with the information configured above:
from langchain.embeddings import AzureOpenAIEmbeddings
embeddings = AzureOpenAIEmbeddings(
azure_deployment=embedding_deployment,
openai_api_version="2023-07-01-preview",
chunk_size=1
)
Loading documents and setting up the retriever
We load the documents from a persistent directory and set up the retriever to run similarity searches:
from langchain.vectorstores import Chroma
persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()
retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})
Initialising the Azure chat model
We create an instance of AzureChatOpenAI with the corresponding settings:
from langchain.chat_models import AzureChatOpenAI
llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)
Formatting documents and creating chat templates
We define a function that formats the retrieved documents and create templates for the question-and-answer system using ChatPromptTemplate:
from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage
def format_docs(docs):
if docs:
return "\n\n".join(doc.page_content for doc in docs)
# Cadenas y plantillas para la condensación de preguntas y el sistema de Q&A
condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
[
("system", condense_q_system_prompt),
MessagesPlaceholder(variable_name="chat_history"),
("human", "{question}"),
]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()
qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
[
("system", qa_system_prompt),
MessagesPlaceholder(variable_name="chat_history"),
("human", "{question} Let’s think step by step"),
]
)
def condense_question(input: dict):
if input.get("chat_history"):
return condense_q_chain
else:
return input["question"]
Initialising the RAG chain
We initialise the RAG chain to generate answers based on the documents and the chat history:
rag_chain = (
RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
| qa_prompt
| llm
)
Interaction between the user and the AI bot
We run a loop that lets the user ask the bot questions and get answers, saving the chat history:
chat_history = []
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
question = input("\033[93m" + "You: " + "\033[0m")
print("\n")
if question == "quit":
break
ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})
print("\033[92m" + ai_msg.content + "\033[0m")
print("\n")
chat_history.extend([HumanMessage(content=question), ai_msg])
Now you have everything you need to build a RAG AI bot with memory. Here is the complete code from the tutorial:
import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from langchain.vectorstores import Chroma
from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage
from langchain.embeddings import AzureOpenAIEmbeddings
from langchain.chat_models import AzureChatOpenAI
from dotenv import load_dotenv
load_dotenv()
# Place a .env file within the same folder with the following information:
# AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_SECRET = "xxxxx"
credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")
# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"
# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment
# Initialize the OpenAIEmbeddings class
embeddings = AzureOpenAIEmbeddings(
azure_deployment=embedding_deployment,
openai_api_version="2023-07-01-preview",
chunk_size=1
)
# Load documents from the persisted directory
persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()
retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})
llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)
# Join all the documents from the retriever together with newlines
def format_docs(docs):
if docs:
return "\n\n".join(doc.page_content for doc in docs)
condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
[
("system", condense_q_system_prompt),
MessagesPlaceholder(variable_name="chat_history"),
("human", "{question}"),
]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()
qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
[
("system", qa_system_prompt),
MessagesPlaceholder(variable_name="chat_history"),
("human", "{question} Let’s think step by step"),
]
)
def condense_question(input: dict):
if input.get("chat_history"):
return condense_q_chain
else:
return input["question"]
# Initialize RAG chain
rag_chain = (
RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
| qa_prompt
| llm
)
chat_history = []
# print welcome message in green
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
#input in yellow
question = input("\033[93m" + "You: " + "\033[0m")
print("\n")
if question == "quit":
break
ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})
# do a simple search with the retriever and the questio
print("\033[92m" + ai_msg.content + "\033[0m")
print("\n")
chat_history.extend([HumanMessage(content=question), ai_msg])
