Bot de IA con documentos y memoria

Este tutorial muestra cómo configurar un bot de IA con documentos y memoria , usando Azure OpenAI y Langchain.

iStock AI Generator

Configuración de las credenciales

Antes de empezar, asegúrate de tener un archivo .env en el mismo directorio que tu código con la siguiente información, que se utiliza para autenticar contra los servicios de Azure:

AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_SECRET = "xxxxx"

En el código, cargamos las variables de entorno y establecemos las credenciales utilizando ChainedTokenCredential y EnvironmentCredential de Azure:

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from dotenv import load_dotenv
load_dotenv()

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

Configuración del modelo y el embedding

Definimos el modelo y el embedding que vamos a utilizar, luego establecemos las variables de entorno de Azure necesarias para interactuar con la API de OpenAI:

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

Inicialización de la clase OpenAIEmbeddings

Se inicializa la clase AzureOpenAIEmbeddings con la información previamente configurada:

from langchain.embeddings import AzureOpenAIEmbeddings

embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

Carga de documentos y configuración del retriever

Cargamos los documentos desde un directorio persistente y configuramos el retriever para realizar búsquedas por similitud:

from langchain.vectorstores import Chroma

persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

Inicialización del modelo de conversación de Azure

Creamos una instancia de AzureChatOpenAI con la configuración correspondiente:

from langchain.chat_models import AzureChatOpenAI

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

Formateo de documentos y creación de plantillas de chat

Definimos una función que formatea los documentos recuperados y creamos plantillas para el sistema de preguntas y respuestas usando ChatPromptTemplate:

from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage

def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

# Cadenas y plantillas para la condensación de preguntas y el sistema de Q&A
condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

Inicialización de la cadena RAG

Inicializamos la cadena RAG para la generación de respuestas basadas en los documentos y el historial de chat:

rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

Interacción con el usuario y el bot de IA

Ejecutamos un bucle que permite al usuario hacer preguntas al bot y recibir respuestas, guardando el historial de chat:

chat_history = []
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Ahora dispones de todo lo necesario para crear un bot de IA RAG con memoria. A continuación, se muestra el código completo del tutorial:

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from langchain.vectorstores import Chroma
from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage
from langchain.embeddings import AzureOpenAIEmbeddings
from langchain.chat_models import AzureChatOpenAI
from dotenv import load_dotenv
load_dotenv()

# Place a .env file within the same folder with the following information:
# AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_SECRET = "xxxxx"

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

# Initialize the OpenAIEmbeddings class
embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

# Load documents from the persisted directory
persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

# Join all the documents from the retriever together with newlines
def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

# Initialize RAG chain
rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

chat_history = []
# print welcome message in green
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    #input in yellow
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})

    # do a simple search with the retriever and the questio
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Maximiliano Díaz Doglia

AI Platform Engineer & Full-Stack Developer
Building Enterprise Integrations & Automations

Publicado en: AI