Bot de IA com documentos e memória

Traduzido do original em espanhol. Ler em espanhol

Este tutorial mostra como configurar um bot de IA com documentos e memória, usando Azure OpenAI e LangChain.

iStock AI Generator

Configuração das credenciais

Antes de começar, certifique-se de ter um arquivo .env no mesmo diretório do seu código com as seguintes informações, usadas para autenticar nos serviços do Azure:

AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_SECRET = "xxxxx"

No código, carregamos as variáveis de ambiente e definimos as credenciais usando ChainedTokenCredential e EnvironmentCredential do Azure:

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from dotenv import load_dotenv
load_dotenv()

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

Configuração do modelo e do embedding

Definimos o modelo e o embedding que vamos usar e depois definimos as variáveis de ambiente do Azure necessárias para interagir com a API da OpenAI:

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

Inicialização da classe OpenAIEmbeddings

A classe AzureOpenAIEmbeddings é inicializada com as informações configuradas anteriormente:

from langchain.embeddings import AzureOpenAIEmbeddings

embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

Carregamento dos documentos e configuração do retriever

Carregamos os documentos de um diretório persistente e configuramos o retriever para fazer buscas por similaridade:

from langchain.vectorstores import Chroma

persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

Inicialização do modelo de conversa do Azure

Criamos uma instância de AzureChatOpenAI com a configuração correspondente:

from langchain.chat_models import AzureChatOpenAI

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

Formatação dos documentos e criação de templates de chat

Definimos uma função que formata os documentos recuperados e criamos templates para o sistema de perguntas e respostas usando ChatPromptTemplate:

from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage

def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

# Cadenas y plantillas para la condensación de preguntas y el sistema de Q&A
condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

Inicialização da cadeia RAG

Inicializamos a cadeia RAG para gerar respostas com base nos documentos e no histórico do chat:

rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

Interação entre o usuário e o bot de IA

Executamos um loop que permite ao usuário fazer perguntas ao bot e receber respostas, salvando o histórico do chat:

chat_history = []
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Agora você tem tudo o que precisa para criar um bot de IA RAG com memória. A seguir, o código completo do tutorial:

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from langchain.vectorstores import Chroma
from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage
from langchain.embeddings import AzureOpenAIEmbeddings
from langchain.chat_models import AzureChatOpenAI
from dotenv import load_dotenv
load_dotenv()

# Place a .env file within the same folder with the following information:
# AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_SECRET = "xxxxx"

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

# Initialize the OpenAIEmbeddings class
embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

# Load documents from the persisted directory
persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

# Join all the documents from the retriever together with newlines
def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

# Initialize RAG chain
rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

chat_history = []
# print welcome message in green
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    #input in yellow
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})

    # do a simple search with the retriever and the questio
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Maximiliano Díaz Doglia

AI Platform Engineer & Full-Stack Developer
Building Enterprise Integrations & Automations

Publicado em: IA