Bot IA con documenti e memoria

Tradotto dall'originale in spagnolo. Leggi in spagnolo

Questo tutorial mostra come configurare un bot IA con documenti e memoria, usando Azure OpenAI e LangChain.

iStock AI Generator

Configurazione delle credenziali

Prima di iniziare, assicurati di avere un file .env nella stessa cartella del codice con le seguenti informazioni, usate per autenticarsi sui servizi di Azure:

AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
AZURE_CLIENT_SECRET = "xxxxx"

Nel codice carichiamo le variabili d’ambiente e impostiamo le credenziali usando ChainedTokenCredential ed EnvironmentCredential di Azure:

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from dotenv import load_dotenv
load_dotenv()

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

Configurazione del modello e dell’embedding

Definiamo il modello e l’embedding da usare, poi impostiamo le variabili d’ambiente di Azure necessarie per interagire con l’API di OpenAI:

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

Inizializzazione della classe OpenAIEmbeddings

La classe AzureOpenAIEmbeddings viene inizializzata con le informazioni configurate in precedenza:

from langchain.embeddings import AzureOpenAIEmbeddings

embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

Caricamento dei documenti e configurazione del retriever

Carichiamo i documenti da una cartella persistente e configuriamo il retriever per eseguire ricerche per similarità:

from langchain.vectorstores import Chroma

persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

Inizializzazione del modello di conversazione di Azure

Creiamo un’istanza di AzureChatOpenAI con la configurazione corrispondente:

from langchain.chat_models import AzureChatOpenAI

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

Formattazione dei documenti e creazione dei template di chat

Definiamo una funzione che formatta i documenti recuperati e creiamo i template per il sistema di domande e risposte usando ChatPromptTemplate:

from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage

def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

# Cadenas y plantillas para la condensación de preguntas y el sistema de Q&A
condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

Inizializzazione della catena RAG

Inizializziamo la catena RAG per generare risposte basate sui documenti e sulla cronologia della chat:

rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

Interazione tra l’utente e il bot IA

Eseguiamo un ciclo che permette all’utente di fare domande al bot e ricevere risposte, salvando la cronologia della chat:

chat_history = []
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Ora hai tutto il necessario per creare un bot IA RAG con memoria. Ecco il codice completo del tutorial:

import os
from azure.identity import ChainedTokenCredential, EnvironmentCredential
from langchain.vectorstores import Chroma
from langchain.schema import StrOutputParser
from langchain.schema.runnable import RunnablePassthrough
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.schema.messages import HumanMessage
from langchain.embeddings import AzureOpenAIEmbeddings
from langchain.chat_models import AzureChatOpenAI
from dotenv import load_dotenv
load_dotenv()

# Place a .env file within the same folder with the following information:
# AZURE_TENANT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_ID = "00000000-0000-0000-0000-000000000000"
# AZURE_CLIENT_SECRET = "xxxxx"

credential = ChainedTokenCredential(EnvironmentCredential())
access_token = credential.get_token("https://cognitiveservices.azure.com/.default")

# Model
deployment = "zerogap-openai-ue2-gpt4-turbo"
# Model text-embedding-ada-002
embedding_deployment = "zerogap-openai-ue2-ada"

# Set OS environment variables
os.environ["AZURE_OPENAI_ENDPOINT"] = "https://zerogap.openai.azure.com/"
os.environ["AZURE_OPENAI_API_KEY"] = access_token.token
os.environ["OPENAI_API_TYPE"] = "azure_ad"
os.environ["OPENAI_DEPLOYMENT"] = deployment

# Initialize the OpenAIEmbeddings class
embeddings = AzureOpenAIEmbeddings(
    azure_deployment=embedding_deployment,
    openai_api_version="2023-07-01-preview",
    chunk_size=1
)

# Load documents from the persisted directory
persist_directory = "chroma_db_generic"
vectordb = Chroma(persist_directory=persist_directory, embedding_function=embeddings)
vectordb.get()

retriever = vectordb.as_retriever(search_type="similarity", search_kwargs={"k": 5})

llm = AzureChatOpenAI(openai_api_version="2023-07-01-preview", azure_deployment=deployment, temperature=0.1)

# Join all the documents from the retriever together with newlines
def format_docs(docs):
    if docs:
        return "\n\n".join(doc.page_content for doc in docs)

condense_q_system_prompt = """Given a chat history and the latest user question \
which might reference the chat history, formulate a standalone question \
which can be understood without the chat history. Do NOT answer the question, \
just reformulate it if needed and otherwise return it as is."""
condense_q_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", condense_q_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question}"),
    ]
)
condense_q_chain = condense_q_prompt | llm | StrOutputParser()

qa_system_prompt = """You are an assistant for question-answering tasks. \
Answer the user question based on provided context and history only. \
If you don't know the answer, just say that you don't know. \
Use three sentences maximum and keep the answer concise.\
All the information in the context is information from documents from the user. \
The context is: {context}"""
qa_prompt = ChatPromptTemplate.from_messages(
    [
        ("system", qa_system_prompt),
        MessagesPlaceholder(variable_name="chat_history"),
        ("human", "{question} Let’s think step by step"),
    ]
)

def condense_question(input: dict):
    if input.get("chat_history"):
        return condense_q_chain
    else:
        return input["question"]

# Initialize RAG chain
rag_chain = (
    RunnablePassthrough.assign(context=condense_question | retriever | format_docs)
    | qa_prompt
    | llm
)

chat_history = []
# print welcome message in green
print("\033[92m" + "Welcome to the Zerogap AI Bot!" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\033[92m" + "###############################" + "\033[0m")
print("\n")
while True:
    #input in yellow
    question = input("\033[93m" + "You: " + "\033[0m")
    print("\n")
    if question == "quit":
        break
    ai_msg = rag_chain.invoke({"question": question, "chat_history": chat_history})

    # do a simple search with the retriever and the questio
    print("\033[92m" + ai_msg.content + "\033[0m")
    print("\n")

    chat_history.extend([HumanMessage(content=question), ai_msg])

Maximiliano Díaz Doglia

AI Platform Engineer & Full-Stack Developer
Building Enterprise Integrations & Automations

Pubblicato in: IA