Bot IA avec documents et mémoire

Traduit de l'original en espagnol. Lire en espagnol

Ce tutoriel montre comment configurer un bot IA avec documents et mémoire, à l’aide d’Azure OpenAI et de LangChain.

iStock AI Generator

Configuration des identifiants

Avant de commencer, assurez-vous d’avoir un fichier .env dans le même répertoire que votre code, avec les informations suivantes, utilisées pour s’authentifier auprès des services Azure :

AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_SECRET = "xxxxx"

Dans le code, nous chargeons les variables d’environnement et définissons les identifiants avec ChainedTokenCredential et EnvironmentCredential d’Azure :

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from dotenv import load_dotenv
load_dotenv()

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

Configuration du modèle et de l’embedding

Nous définissons le modèle et l’embedding à utiliser, puis les variables d’environnement Azure nécessaires pour interagir avec l’API OpenAI :

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

Initialisation de la classe OpenAIEmbeddings

La classe AzureOpenAIEmbeddings est initialisée avec les informations configurées précédemment :

from langchain.embeddings import AzureOpenAIEmbeddings

embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

Chargement des documents et configuration du retriever

Nous chargeons les documents depuis un répertoire persistant et configurons le retriever pour effectuer des recherches par similarité :

from langchain.vectorstores import Chroma

persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

Initialisation du modèle de conversation Azure

Nous créons une instance d’AzureChatOpenAI avec la configuration correspondante :

from langchain.chat_models import AzureChatOpenAI

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

Formatage des documents et création des modèles de chat

Nous définissons une fonction qui formate les documents récupérés et créons des modèles pour le système de questions-réponses avec ChatPromptTemplate :

from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage

def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

# Cadenas y plantillas para la condensación de preguntas y el sistema de Q&A
condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

Initialisation de la chaîne RAG

Nous initialisons la chaîne RAG pour générer des réponses à partir des documents et de l’historique de la conversation :

rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

Interaction entre l’utilisateur et le bot IA

Nous exécutons une boucle qui permet à l’utilisateur de poser des questions au bot et d’obtenir des réponses, en enregistrant l’historique de la conversation :

chat_history = []
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Vous disposez maintenant de tout le nécessaire pour créer un bot IA RAG avec mémoire. Voici le code complet du tutoriel :

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from langchain.vectorstores import Chroma
from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage
from langchain.embeddings import AzureOpenAIEmbeddings
from langchain.chat_models import AzureChatOpenAI
from dotenv import load_dotenv
load_dotenv()

# Place a .env file within the same folder with the following information:
# AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_SECRET = "xxxxx"

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

# Initialize the OpenAIEmbeddings class
embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

# Load documents from the persisted directory
persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

# Join all the documents from the retriever together with newlines
def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

# Initialize RAG chain
rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

chat_history = []
# print welcome message in green
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    #input in yellow
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})

    # do a simple search with the retriever and the questio
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Maximiliano Díaz Doglia

AI Platform Engineer & Full-Stack Developer
Building Enterprise Integrations & Automations

Publié dans : IA