Come creare una suite di sicurezza con l'IA partendo da strumenti open source

Tradotto dall'originale in spagnolo. Leggi in spagnolo

La nostra configurazione è un server Linux che esegue diversi strumenti di sicurezza open source: fail2ban, ClamAV, rkhunter, maldet e lynis. Ogni strumento scrive i propri log in una cartella comune, insieme a uno script bash che aggrega gli eventi di fail2ban, SSH e altre fonti. Per dare un senso a questi dati usiamo uno script Python che si appoggia ad Azure OpenAI tramite la libreria LangChain.

Disclaimer:
Questo è solo un esempio. Se vuoi usarlo in produzione, analizza prima i tuoi requisiti specifici, esegui test approfonditi e adatta il codice alle esigenze e alle policy della tua organizzazione. Non usare questo codice così com’è senza averlo rivisto e adattato al tuo ambiente. Proteggi con cura i file di configurazione, le credenziali e i report generati.


iStock AI Generator

Cosa sono questi strumenti e perché eseguirli insieme?

In questo esempio combiniamo diversi strumenti open source ampiamente riconosciuti per la sicurezza dei server Linux:

  • fail2ban: monitora i log di servizi come SSH per rilevare tentativi di accesso falliti o attacchi di forza bruta e blocca automaticamente gli IP responsabili tramite regole del firewall.
  • ClamAV: un antivirus open source che analizza i file alla ricerca di malware, virus e trojan conosciuti.
  • rkhunter (Rootkit Hunter): cerca rootkit, backdoor ed exploit locali esaminando file di sistema, binari e configurazioni sospette.
  • maldet (Linux Malware Detect): specializzato nel rilevare malware noto e minacce web in ambienti Linux, particolarmente utile per i server web condivisi.
  • lynis: uno strumento di audit di sicurezza e conformità che analizza la configurazione, cerca vulnerabilità e fornisce raccomandazioni per rafforzare il sistema.

Perché eseguirli insieme?
Ognuno di questi strumenti copre un aspetto diverso della sicurezza del sistema: dal rilevamento delle intrusioni e dal blocco degli accessi non autorizzati fino all’analisi di malware e rootkit e agli audit di configurazione. Eseguendoli tutti insieme e combinandone i risultati otteniamo una visione molto più completa e approfondita dello stato di sicurezza del server. Inoltre, centralizzare i log e analizzarli con l’intelligenza artificiale permette di correlare gli eventi, scoprire schemi che potrebbero passare inosservati e dare priorità alle azioni di risposta in modo più efficiente.

1. Caricamento dell’ambiente e della configurazione

Lo script inizia caricando le variabili d’ambiente da un file .env (se esiste) o dall’ambiente del sistema. Poi verifica che siano presenti tutte le variabili richieste, interrompendo l’esecuzione se ne manca qualcuna.

Pratica consigliata:
Assicurati che il file .env e qualsiasi file di configurazione siano ben protetti e accessibili solo all’utente necessario. Valuta di cifrare i file che contengono credenziali.

def load_environment():
    """Load environment variables from dotenv file and check required ones."""
    if Path(ENV_FILE).exists():
        load_dotenv(dotenv_path=ENV_FILE)
        logging.info(f"Loaded environment from {ENV_FILE}")
    else:
        logging.warning(f".env file not found at {ENV_FILE}; relying on current environment.")

    missing_vars = [var for var in REQUIRED_ENV if not os.getenv(var)]
    if missing_vars:
        logging.error(f"Missing required environment variables: {', '.join(missing_vars)}")
        sys.exit(1)

2. Verifica e creazione delle cartelle

Si assicura che le cartelle dei log, dell’output e dell’archivio esistano, creandole se necessario. Se manca la cartella dei log, lo script si interrompe.

Importante:
Verifica i permessi delle cartelle e segmenta gli accessi per evitare che utenti non autorizzati possano manipolare i file di log e i report.

def ensure_directories():
    """Ensure log, output, and archive directories exist."""
    if not LOG_DIR.is_dir():
        logging.error(f"Log directory not found: {LOG_DIR}")
        sys.exit(1)
    OUT_DIR.mkdir(parents=True, exist_ok=True)
    ARCHIVE_DIR.mkdir(parents=True, exist_ok=True)

3. Autenticazione con Azure

La funzione get_azure_ad_token ottiene un token di accesso da Azure Active Directory usando le client credentials.

Nota:
Assicurati che i permessi su Azure siano i minimi necessari per il tuo caso d’uso. Le credenziali usate devono essere conservate e trasmesse in modo sicuro.

get_azure_ad_token():
    """Acquire Azure AD Token using client credentials."""
    token_url = f"https://login.microsoftonline.com/{os.environ['AZURE_TENANT_ID']}/oauth2/v2.0/token"
    token_data = {
        "grant_type": "client_credentials",
        "client_id": os.environ['AZURE_CLIENT_ID'],
        "client_secret": os.environ['AZURE_CLIENT_SECRET'],
        "scope": os.environ.get('AZURE_OPENAI_SCOPE', 'https://cognitiveservices.azure.com/.default'),
    }
    response = requests.post(token_url, data=token_data)
    response.raise_for_status()
    return response.json()["access_token"]

4. Caricamento dei prompt per l’IA

Lo script carica due prompt da file di testo: uno per l’analisi individuale di ogni log e l’altro per il report finale aggregato.

Importante:
Il contenuto dei prompt determina la qualità e l’impostazione dei report. Rivedi e adatta i prompt al livello tecnico e alle esigenze della tua organizzazione.

load_prompt(prompt_file: Path) -> str:
    """Load a prompt from a file."""
    if not prompt_file.is_file():
        logging.error(f"{prompt_file} not found!")
        sys.exit(1)
    with prompt_file.open("r", encoding="utf-8") as f:
        return f.read()

5. Raccolta dei log

Raccoglie tutti i file .log e .txt della cartella dei log per analizzarli.

Nota:
Assicurati che nella cartella ci siano solo i log attesi, per evitare di analizzare file irrilevanti o potenzialmente malevoli.

collect_log_files() -> list:
    """Collect log files (*.log, *.txt) from LOG_DIR."""
    log_files = list(LOG_DIR.glob("*.log")) + list(LOG_DIR.glob("*.txt"))
    return log_files

6. Analisi IA di ogni log

Legge ogni log, lo invia all’IA insieme al prompt individuale e salva il risultato.

Attenzione:
L’IA può commettere errori di interpretazione o non rilevare minacce avanzate. Non sostituisce la supervisione e la verifica umana nei contesti critici.

log_file in log_files:
    try:
        with log_file.open("r", encoding="utf-8") as f:
            log_content = f.read()
        messages = [
            SystemMessage(content=per_log_system_prompt),
            HumanMessage(content=log_content)
        ]
        response = llm(messages)
        report_text = response.content if hasattr(response, 'content') else str(response)
        individual_reports.append({
            "log_file": log_file.name,
            "report": report_text
        })
    except Exception as e:
        logging.error(f"Error processing {log_file}: {e}")

7. Salvataggio dei report individuali

Salva in un file di testo i report individuali generati dall’IA.

Buona pratica:
Rivedi periodicamente i report e tienili sotto controllo di versione se la tua organizzazione lo richiede.

individual_reports_txt = ""
for report in individual_reports:
    individual_reports_txt += f"===== Report for {report['log_file']} =====\n{report['report']}\n\n"

try:
    with REPORT_TXT.open("w", encoding="utf-8") as f:
        f.write(individual_reports_txt)
except Exception as e:
    logging.error(f"Error writing individual reports: {e}")

8. Report finale aggregato dall’IA

L’IA riceve tutti i report individuali e genera un’analisi o un riepilogo complessivo.

Importante:
Convalida i risultati dell’IA e tieni conto dei suoi limiti. Considera l’output come un supporto e non come una decisione automatica definitiva.

reports_content = "\n".join(
    f"Report for {r['log_file']}:\n{r['report']}" for r in individual_reports
)
final_messages = [
    SystemMessage(content=final_system_prompt),
    HumanMessage(content=reports_content)
]
try:
    final_response = llm(final_messages)
    final_report = final_response.content if hasattr(final_response, 'content') else str(final_response)
except Exception as e:
    logging.error(f"Error during final aggregation OpenAI call: {e}")
    final_report = "Aggregation failed due to error."

9. Salvataggio e permessi dei report

Il report finale e quelli individuali vengono salvati in file .txt e .json, con i permessi appropriati.

Raccomandazione:
Valuta di cifrare i report se contengono informazioni sensibili e configura backup automatici in un ambiente sicuro.

try:
    with REPORT_TXT.open("a", encoding="utf-8") as f:
        f.write("\n===== FINAL AGGREGATED REPORT =====\n")
        f.write(final_report)
    with REPORT_JSON.open("w", encoding="utf-8") as f:
        json.dump({
            "individual_reports": individual_reports,
            "final_report": final_report
        }, f, indent=2)
except Exception as e:
    logging.error(f"Error saving final report: {e}")

def set_permissions():
    """Set secure permissions on report files."""
    try:
        REPORT_JSON.chmod(0o640)
        REPORT_TXT.chmod(0o640)
    except Exception as e:
        logging.error(f"Error setting permissions on report files: {e}")

10. Invio automatico via email (SMTP)

Invia il report di sicurezza via email usando SMTP.

Attenzione:
Assicurati di usare password per le app e connessioni TLS/SSL. Non usare account personali e non esporre credenziali in file non sicuri.

def send_email_report(report_path: Path):
    """Send the TXT report as the email body using smtplib."""
    SMTP_SERVER = os.getenv('SMTP_SERVER')
    SMTP_PORT = int(os.getenv('SMTP_PORT', '587'))
    SMTP_USERNAME = os.getenv('SMTP_USERNAME')  
    SMTP_PASSWORD = os.getenv('SMTP_PASSWORD')

    if not SMTP_SERVER or not MAIL_TO or not MAIL_FROM:
        logging.error("SMTP configuration or email addresses missing.")
        logging.info(f"Reports saved at:\n  JSON: {REPORT_JSON}\n  Text: {REPORT_TXT}")
        sys.exit(3)
    try:
        with report_path.open("r", encoding="utf-8") as f:
            report_content = f.read()

        msg = EmailMessage()
        msg['Subject'] = MAIL_SUBJECT
        msg['From'] = MAIL_FROM
        msg['To'] = MAIL_TO
        msg.set_content(report_content)

        with smtplib.SMTP(SMTP_SERVER, SMTP_PORT) as server:
            server.starttls()
            if SMTP_USERNAME and SMTP_PASSWORD:
                server.login(SMTP_USERNAME, SMTP_PASSWORD)
            server.send_message(msg)
        logging.info(f"Security report emailed to {MAIL_TO} via SMTP")
    except Exception as e:
        logging.error(f"Failed to email report via SMTP: {e}")

11. Archiviazione e pulizia

Tutti i log e i report vengono compressi in un file ZIP e i file vecchi vengono eliminati secondo la policy di conservazione.

Importante:
Controlla regolarmente lo spazio su disco e lo stato dei file archiviati, e monitora le prestazioni del sistema, soprattutto se il volume dei log è elevato.

def archive_and_cleanup():
    """Archive logs and reports, and clean up old files."""
    zip_name = f"security-logs-{datetime.datetime.now().strftime('%Y%m%d-%H%M%S')}.zip"
    logs_to_archive = list(LOG_DIR.glob("*.txt")) + list(LOG_DIR.glob("*.log")) + [REPORT_JSON]
    files_to_archive = [f for f in logs_to_archive if f.is_file()]
    archive_path = ARCHIVE_DIR / zip_name

    if files_to_archive:
        try:
            with zipfile.ZipFile(archive_path, 'w', zipfile.ZIP_DEFLATED) as zipf:
                for file in files_to_archive:
                    zipf.write(file, arcname=file.name)
            for file in files_to_archive:
                file.unlink()
            logging.info(f"Archived and removed: {' '.join(str(f) for f in files_to_archive)}")
        except Exception as e:
            logging.error(f"Error archiving files: {e}")
    else:
        logging.info(f"No .txt, .log, or .json files to archive in {LOG_DIR}.")

    # Remove old archives
    now = datetime.datetime.now().timestamp()
    for f in ARCHIVE_DIR.glob("*.zip"):
        try:
            if f.stat().st_mtime < now - ARCHIVE_RETENTION_DAYS * 86400:
                f.unlink()
        except Exception as e:
            logging.error(f"Error deleting archive {f}: {e}")

    # Remove other files/directories in LOG_DIR except .log_offsets
    for f in LOG_DIR.iterdir():
        if f.name == '.log_offsets':
            continue
        try:
            if f.is_dir():
                shutil.rmtree(f)
            else:
                f.unlink()
        except Exception as e:
            logging.error(f"Error removing {f}: {e}")

12. System prompt utilizzati

// MAIN LLM - This will handle the final report
/*
Act as a senior security analyst.
You are given output from an AI analyst.
Logs are from Fail2ban, ClamAV, RKHunter, Lynis, Maldet,
and SSH authentication logs from /var/log/security-scans.
Tasks:

Vulnerabilities: Identify misconfigurations, outdated practices,
open attack surfaces, or missing controls in the system.
Provide severity, verbatim evidence, and remediation steps for each finding.

Detections: Extract concrete security events, such as repeated failed SSH attempts,
or banned IP address activity.
Include first_seen, last_seen, counts, and impacted entities as evidence.

Patterns & Analysis: Summarize attacker behavior patterns,
such as targeted usernames and source IP ranges observed.
Include defenses triggered, like bans, lockouts, and possible evasion indicators.

Executive Summary: Provide a concise, non-technical summary of risk posture,
including urgent recommended actions if found.

Appendices:

a. List Indicators of Compromise (IPs, users, keys)
and note any data coverage limitations encountered in analysis.
b. Output a single JSON object containing:
{
“executive_summary”: “…”,
“detections”: [
{
“type”: “…”,
“details”: “…”,
“evidence”: [“…”],
“first_seen”: “…”,
“last_seen”: “…”,
“count”: 0,
“severity”: “low|medium|high|critical”
}
],
“vulnerabilities”: [
{
“issue”: “…”,
“severity”: “…”,
“evidence”: [“…”],
“remediation”: “…”
}
],
“patterns_analysis”: {
“attacker_behaviors”: “…”,
“defenses”: “…”,
“timeline”: “…”,
“notable_patterns”: “…”
},
“ioc”: {
“ips”: [“…”],
“users”: [“…”],
“other”: [“…”]
},
“limitations”: “…”
}
c. Use only concrete evidence from logs
and quote relevant lines in the evidence arrays.
*/

// JUNIOR LLM - This will handle individual log analysis
/*
You are a Linux security log analyst.
Review a single Linux security log entry for issues.
Identify security issues, anomalies, or suspicious activities present in the log.
Summarize findings concisely and clearly for further review.
Focus on:

All relevant security events or indicators in the log entry.

Potential impact or risk level identified from the analysis.

Actionable insights or next steps, if applicable, for remediating issues.

Do not include the full log entry in your response.
Avoid unnecessary verbosity to keep the summary focused and clear.
Your summary will be used by senior analysts for review and aggregation.

Format response as:

Findings:
Insights:
Recommendations (if any):
*/

Esempio: script Bash per creare log di attività personalizzati da analizzare

  • Cambia /var/log/fail2ban.log, ecc. in base ai percorsi dei log della tua distribuzione, se necessario.
  • Per motivi di privacy, evita di eseguirlo o condividerlo su sistemi con dati sensibili o regolamentati senza un’adeguata revisione.
#!/usr/bin/env bash
set -euo pipefail

# Set output directory in a generic temp/logs path
OUT_DIR="/tmp/security-scans"
TS="$(date +'%Y%m%d-%H%M%S')"
mkdir -p "$OUT_DIR"

# Helper to safely run commands that might not exist
have() { command -v "$1" >/dev/null 2>&1; }

# -----------------------------
# Fail2ban: jails and bans (last 48h from fail2ban.log)
# -----------------------------
if have fail2ban-client && [ -r /var/log/fail2ban.log ]; then
  {
    echo "== Fail2ban bans (last 48 hours) =="
    awk -v d="$(date --date='-48 hours' '+%Y-%m-%d %H:%M:%S')" \
      '$0 >= d && /Ban /' /var/log/fail2ban.log
  } > "$OUT_DIR/fail2ban-bans-activity-$TS.txt" 2>&1
fi

# -----------------------------
# SSH: successful and failed logins (last 48h)
# -----------------------------
AUTH_FILE=""
for cand in /var/log/auth.log /var/log/secure; do
  if [ -r "$cand" ]; then AUTH_FILE="$cand"; break; fi
done

if [ -n "$AUTH_FILE" ]; then
  {
    echo "== SSH successful logins (last 48h) =="
    for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
      awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Accepted (password|publickey)" || true
    done

    echo
    echo "== SSH failed logins (last 48h) =="
    for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
      awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Failed password" || true
    done

    echo
    echo "== SSH invalid users (last 48h) =="
    for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
      awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Invalid user" || true
    done

    echo
    echo "== SSH root or sudo session opens (last 48h) =="
    for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
      awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: (pam_unix|session opened|Accepted .+ for root)" || true
    done
  } > "$OUT_DIR/ssh-logins-$TS.txt" 2>&1

  # Compact summaries
  {
    echo "== Summary (counts, last 48h) =="
    echo -n "Accepted logins: "
    count=0
    for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
      count=$((count + $(awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Accepted (password|publickey)" | wc -l)))
    done
    echo "$count"
    echo -n "Failed passwords: "
    count=0
    for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
      count=$((count + $(awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Failed password" | wc -l)))
    done
    echo "$count"
    echo -n "Invalid users: "
    count=0
    for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
      count=$((count + $(awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Invalid user" | wc -l)))
    done
    echo "$count"
  } > "$OUT_DIR/ssh-summary-$TS.txt" 2>&1
fi

# -----------------------------
# Permissions
# -----------------------------
chmod 640 "$OUT_DIR"/*"$TS".txt || true

echo "Security scans written to: $OUT_DIR (timestamp $TS)"

Codice completo

#!/usr/bin/env python3

import os
import sys
import glob
import shutil
import socket
import datetime
import zipfile
import requests
import json
import logging
from dotenv import load_dotenv
from langchain_openai import AzureChatOpenAI
from langchain.schema import SystemMessage, HumanMessage
from pathlib import Path
import smtplib
from email.message import EmailMessage

# ============================
# Constants & Configuration
# ============================
ARCHIVE_RETENTION_DAYS = 30
LOG_DIR = Path(os.getenv('LOG_DIR', '/var/log/security'))
OUT_DIR = Path(os.getenv('OUT_DIR', '/var/log/security'))
ARCHIVE_DIR = Path(os.getenv('ARCHIVE_DIR', '/var/log/archived'))
ENV_FILE = os.environ.get('ENV_FILE', './.env')
SYSTEM_PROMPT_FILE = Path('system_prompt.txt')
JUNIOR_SYSTEM_PROMPT_FILE = Path('junior_system_prompt.txt')
TIMESTAMP = datetime.datetime.now().strftime('%Y%m%d-%H%M%S')
REPORT_JSON = OUT_DIR / f'security-report-{TIMESTAMP}.json'
REPORT_TXT = OUT_DIR / f'security-report-{TIMESTAMP}.txt'

MAIL_TO = os.getenv('MAIL_TO', 'info@email.com')
MAIL_SUBJECT = os.getenv('MAIL_SUBJECT', f'Security Report {TIMESTAMP}')
MAIL_FROM = os.getenv('MAIL_FROM', f"security-reports@{socket.getfqdn() if hasattr(socket, 'getfqdn') else socket.gethostname()}")

REQUIRED_ENV = [
    'AZURE_OPENAI_ENDPOINT',
    'AZURE_OPENAI_DEPLOYMENT',
    'API_VERSION',
    'AZURE_TENANT_ID',
    'AZURE_CLIENT_ID',
    'AZURE_CLIENT_SECRET',
]

# ============================
# Logging Setup
# ============================
logging.basicConfig(
    level=logging.INFO,
    format='%(asctime)s [%(levelname)s] %(message)s',
    handlers=[logging.StreamHandler(sys.stderr)]
)

# ============================
# Helper Functions
# ============================

def load_environment():
    """Load environment variables from dotenv file and check required ones."""
    if Path(ENV_FILE).exists():
        load_dotenv(dotenv_path=ENV_FILE)
        logging.info(f"Loaded environment from {ENV_FILE}")
    else:
        logging.warning(f".env file not found at {ENV_FILE}; relying on current environment.")

    missing_vars = [var for var in REQUIRED_ENV if not os.getenv(var)]
    if missing_vars:
        logging.error(f"Missing required environment variables: {', '.join(missing_vars)}")
        sys.exit(1)

def ensure_directories():
    """Ensure log, output, and archive directories exist."""
    if not LOG_DIR.is_dir():
        logging.error(f"Log directory not found: {LOG_DIR}")
        sys.exit(1)
    OUT_DIR.mkdir(parents=True, exist_ok=True)
    ARCHIVE_DIR.mkdir(parents=True, exist_ok=True)

def get_azure_ad_token():
    """Acquire Azure AD Token using client credentials."""
    token_url = f"https://login.microsoftonline.com/{os.environ['AZURE_TENANT_ID']}/oauth2/v2.0/token"
    token_data = {
        "grant_type": "client_credentials",
        "client_id": os.environ['AZURE_CLIENT_ID'],
        "client_secret": os.environ['AZURE_CLIENT_SECRET'],
        "scope": os.environ.get('AZURE_OPENAI_SCOPE', 'https://cognitiveservices.azure.com/.default'),
    }
    response = requests.post(token_url, data=token_data)
    response.raise_for_status()
    return response.json()["access_token"]

def load_prompt(prompt_file: Path) -> str:
    """Load a prompt from a file."""
    if not prompt_file.is_file():
        logging.error(f"{prompt_file} not found!")
        sys.exit(1)
    with prompt_file.open("r", encoding="utf-8") as f:
        return f.read()

def collect_log_files() -> list:
    """Collect log files (*.log, *.txt) from LOG_DIR."""
    log_files = list(LOG_DIR.glob("*.log")) + list(LOG_DIR.glob("*.txt"))
    return log_files

def send_email_report(report_path: Path):
    """Send the TXT report as the email body using smtplib."""
    SMTP_SERVER = os.getenv('SMTP_SERVER')
    SMTP_PORT = int(os.getenv('SMTP_PORT', '587'))
    SMTP_USERNAME = os.getenv('SMTP_USERNAME')  
    SMTP_PASSWORD = os.getenv('SMTP_PASSWORD')

    if not SMTP_SERVER or not MAIL_TO or not MAIL_FROM:
        logging.error("SMTP configuration or email addresses missing.")
        logging.info(f"Reports saved at:\n  JSON: {REPORT_JSON}\n  Text: {REPORT_TXT}")
        sys.exit(3)
    try:
        with report_path.open("r", encoding="utf-8") as f:
            report_content = f.read()

        msg = EmailMessage()
        msg['Subject'] = MAIL_SUBJECT
        msg['From'] = MAIL_FROM
        msg['To'] = MAIL_TO
        msg.set_content(report_content)

        with smtplib.SMTP(SMTP_SERVER, SMTP_PORT) as server:
            server.starttls()
            if SMTP_USERNAME and SMTP_PASSWORD:
                server.login(SMTP_USERNAME, SMTP_PASSWORD)
            server.send_message(msg)
        logging.info(f"Security report emailed to {MAIL_TO} via SMTP")
    except Exception as e:
        logging.error(f"Failed to email report via SMTP: {e}")

def archive_and_cleanup():
    """Archive logs and reports, and clean up old files."""
    zip_name = f"security-logs-{datetime.datetime.now().strftime('%Y%m%d-%H%M%S')}.zip"
    logs_to_archive = list(LOG_DIR.glob("*.txt")) + list(LOG_DIR.glob("*.log")) + [REPORT_JSON]
    files_to_archive = [f for f in logs_to_archive if f.is_file()]
    archive_path = ARCHIVE_DIR / zip_name

    if files_to_archive:
        try:
            with zipfile.ZipFile(archive_path, 'w', zipfile.ZIP_DEFLATED) as zipf:
                for file in files_to_archive:
                    zipf.write(file, arcname=file.name)
            for file in files_to_archive:
                file.unlink()
            logging.info(f"Archived and removed: {' '.join(str(f) for f in files_to_archive)}")
        except Exception as e:
            logging.error(f"Error archiving files: {e}")
    else:
        logging.info(f"No .txt, .log, or .json files to archive in {LOG_DIR}.")

    # Remove old archives
    now = datetime.datetime.now().timestamp()
    for f in ARCHIVE_DIR.glob("*.zip"):
        try:
            if f.stat().st_mtime < now - ARCHIVE_RETENTION_DAYS * 86400:
                f.unlink()
        except Exception as e:
            logging.error(f"Error deleting archive {f}: {e}")

    # Remove other files/directories in LOG_DIR except .log_offsets
    for f in LOG_DIR.iterdir():
        if f.name == '.log_offsets':
            continue
        try:
            if f.is_dir():
                shutil.rmtree(f)
            else:
                f.unlink()
        except Exception as e:
            logging.error(f"Error removing {f}: {e}")

def set_permissions():
    """Set secure permissions on report files."""
    try:
        REPORT_JSON.chmod(0o640)
        REPORT_TXT.chmod(0o640)
    except Exception as e:
        logging.error(f"Error setting permissions on report files: {e}")

# ============================
# Main Functionality
# ============================

def main():
    load_environment()
    ensure_directories()

    # Prepare LangChain AzureChatOpenAI
    llm = AzureChatOpenAI(
        azure_deployment=os.environ["AZURE_OPENAI_DEPLOYMENT"],
        api_version=os.environ["API_VERSION"],
        azure_endpoint=os.environ["AZURE_OPENAI_ENDPOINT"],
        azure_ad_token_provider=get_azure_ad_token,
        temperature=0.7,
    )

    # Load prompts
    final_system_prompt = load_prompt(SYSTEM_PROMPT_FILE)
    per_log_system_prompt = load_prompt(JUNIOR_SYSTEM_PROMPT_FILE)

    # Collect log files
    log_files = collect_log_files()
    individual_reports = []

    # Query OpenAI for each log file
    for log_file in log_files:
        try:
            with log_file.open("r", encoding="utf-8") as f:
                log_content = f.read()
            messages = [
                SystemMessage(content=per_log_system_prompt),
                HumanMessage(content=log_content)
            ]
            response = llm(messages)
            report_text = response.content if hasattr(response, 'content') else str(response)
            individual_reports.append({
                "log_file": log_file.name,
                "report": report_text
            })
        except Exception as e:
            logging.error(f"Error processing {log_file}: {e}")

    # Save individual reports
    individual_reports_txt = ""
    for report in individual_reports:
        individual_reports_txt += f"===== Report for {report['log_file']} =====\n{report['report']}\n\n"

    try:
        with REPORT_TXT.open("w", encoding="utf-8") as f:
            f.write(individual_reports_txt)
    except Exception as e:
        logging.error(f"Error writing individual reports: {e}")

    # Final aggregation OpenAI call
    reports_content = "\n".join(
        f"Report for {r['log_file']}:\n{r['report']}" for r in individual_reports
    )
    final_messages = [
        SystemMessage(content=final_system_prompt),
        HumanMessage(content=reports_content)
    ]
    try:
        final_response = llm(final_messages)
        final_report = final_response.content if hasattr(final_response, 'content') else str(final_response)
    except Exception as e:
        logging.error(f"Error during final aggregation OpenAI call: {e}")
        final_report = "Aggregation failed due to error."

    # Append final aggregated report
    try:
        with REPORT_TXT.open("a", encoding="utf-8") as f:
            f.write("\n===== FINAL AGGREGATED REPORT =====\n")
            f.write(final_report)
        with REPORT_JSON.open("w", encoding="utf-8") as f:
            json.dump({
                "individual_reports": individual_reports,
                "final_report": final_report
            }, f, indent=2)
    except Exception as e:
        logging.error(f"Error saving final report: {e}")

    set_permissions()
    send_email_report(REPORT_TXT)
    archive_and_cleanup()

if __name__ == "__main__":
    main()

Orchestrando strumenti open source e sfruttando un modello linguistico (LLM) per l’analisi e il reporting dei log, questo script trasforma un insieme di utility in una piattaforma di sicurezza potenziata dall’IA. L’approccio automatizza la revisione dei log, contestualizza i risultati e tiene informati gli amministratori, il tutto con un intervento manuale minimo.

Tuttavia ricorda:

  • L’IA può non rilevare minacce avanzate e deve essere un aiuto, non l’unico metodo di difesa.
  • Proteggi i file e le credenziali.
  • Convalida e rivedi i report.
  • Monitora il sistema e regola i parametri in base al carico e alla dimensione dei log.

Questa architettura è un esempio pratico di come l’IA possa potenziare la gestione dell’infrastruttura, portando chiarezza, efficienza e informazioni utilizzabili nel mondo rumoroso dei log di sicurezza.

Maximiliano Díaz Doglia

AI Platform Engineer & Full-Stack Developer
Building Enterprise Integrations & Automations