How to build an AI security suite from open source tools
Translated from the Spanish original. Read in Spanish
Our setup is a Linux server running several open source security tools: fail2ban, ClamAV, rkhunter, maldet and lynis. Each tool writes its log output to a shared directory, along with a bash script that aggregates events from fail2ban, SSH and other sources. To make sense of this data, we use a Python script that calls Azure OpenAI through the LangChain library.
Disclaimer:
This is only an example. If you want to use it in production, first analyse your specific requirements, test thoroughly and adapt the code to your organisation’s needs and policies. Don’t use this code as-is without reviewing it and adjusting it to your environment. Make sure you carefully protect configuration files, credentials and the generated reports.

What are these tools and why run them together?
In this example, we combine several widely recognised open source tools for Linux server security:
- fail2ban: Monitors the logs of services such as SSH to detect failed login attempts or brute-force attacks, and automatically blocks the offending IPs with firewall rules.
- ClamAV: An open source antivirus that scans files for known malware, viruses and trojans.
- rkhunter (Rootkit Hunter): Looks for rootkits, backdoors and local exploits by examining system files, binaries and suspicious configurations.
- maldet (Linux Malware Detect): Specialised in detecting known malware and web threats on Linux, especially useful for shared web servers.
- lynis: A security auditing and compliance tool that analyses configuration, looks for vulnerabilities and gives recommendations to harden the system.
Why run them together?
Each of these tools covers a different aspect of system security: from intrusion detection and blocking unauthorised access to malware and rootkit analysis and configuration audits. By running them all together and combining their results, we get a much more complete and in-depth view of the server’s security posture. On top of that, centralising the logs and analysing them with artificial intelligence lets us correlate events, uncover patterns that might go unnoticed and prioritise response actions more efficiently.
1. Loading the environment and configuration
The script starts by loading environment variables from a .env file (if there is one) or from the system environment itself. It then checks that all the required variables are present, stopping if any are missing.
Recommended practice:
Make sure the .env file and any configuration files are properly protected and accessible only to the user that needs them. Consider encrypting files that contain credentials.
def load_environment():
"""Load environment variables from dotenv file and check required ones."""
if Path(ENV_FILE).exists():
load_dotenv(dotenv_path=ENV_FILE)
logging.info(f"Loaded environment from {ENV_FILE}")
else:
logging.warning(f".env file not found at {ENV_FILE}; relying on current environment.")
missing_vars = [var for var in REQUIRED_ENV if not os.getenv(var)]
if missing_vars:
logging.error(f"Missing required environment variables: {', '.join(missing_vars)}")
sys.exit(1)
2. Checking and creating directories
It makes sure the log, output and archive directories exist, creating them if needed. If the log directory is missing, the script stops.
Important:
Check directory permissions and segment access so unauthorised users can’t tamper with the log files and reports.
def ensure_directories():
"""Ensure log, output, and archive directories exist."""
if not LOG_DIR.is_dir():
logging.error(f"Log directory not found: {LOG_DIR}")
sys.exit(1)
OUT_DIR.mkdir(parents=True, exist_ok=True)
ARCHIVE_DIR.mkdir(parents=True, exist_ok=True)
3. Authenticating with Azure
The get_azure_ad_token function gets an access token from Azure Active Directory using client credentials.
Note:
Make sure the Azure permissions are the minimum your use case needs. The credentials used must be stored and transmitted securely.
get_azure_ad_token():
"""Acquire Azure AD Token using client credentials."""
token_url = f"https://login.microsoftonline.com/{os.environ['AZURE_TENANT_ID']}/oauth2/v2.0/token"
token_data = {
"grant_type": "client_credentials",
"client_id": os.environ['AZURE_CLIENT_ID'],
"client_secret": os.environ['AZURE_CLIENT_SECRET'],
"scope": os.environ.get('AZURE_OPENAI_SCOPE', 'https://cognitiveservices.azure.com/.default'),
}
response = requests.post(token_url, data=token_data)
response.raise_for_status()
return response.json()["access_token"]
4. Loading the prompts for the AI
The script loads two prompts from text files: one for analysing each log individually and another for the final aggregated report.
Important:
The content of the prompts determines the quality and focus of the reports. Review and adjust the prompts to the technical level and needs of your organisation.
load_prompt(prompt_file: Path) -> str:
"""Load a prompt from a file."""
if not prompt_file.is_file():
logging.error(f"{prompt_file} not found!")
sys.exit(1)
with prompt_file.open("r", encoding="utf-8") as f:
return f.read()
5. Collecting the logs
It gathers all the .log and .txt files in the log directory for analysis.
Note:
Make sure only the expected logs are in the directory, so you don’t analyse irrelevant or potentially malicious files.
collect_log_files() -> list:
"""Collect log files (*.log, *.txt) from LOG_DIR."""
log_files = list(LOG_DIR.glob("*.log")) + list(LOG_DIR.glob("*.txt"))
return log_files
6. AI analysis of each log
It reads each log, sends it to the AI along with the individual prompt and stores the result.
Warning:
The AI can misinterpret things or miss advanced threats. It doesn’t replace human oversight and verification in critical contexts.
log_file in log_files:
try:
with log_file.open("r", encoding="utf-8") as f:
log_content = f.read()
messages = [
SystemMessage(content=per_log_system_prompt),
HumanMessage(content=log_content)
]
response = llm(messages)
report_text = response.content if hasattr(response, 'content') else str(response)
individual_reports.append({
"log_file": log_file.name,
"report": report_text
})
except Exception as e:
logging.error(f"Error processing {log_file}: {e}")
7. Saving the individual reports
It saves the individual reports generated by the AI to a text file.
Best practice:
Review the reports regularly and keep them under version control if your organisation requires it.
individual_reports_txt = ""
for report in individual_reports:
individual_reports_txt += f"===== Report for {report['log_file']} =====\n{report['report']}\n\n"
try:
with REPORT_TXT.open("w", encoding="utf-8") as f:
f.write(individual_reports_txt)
except Exception as e:
logging.error(f"Error writing individual reports: {e}")
8. Final AI-aggregated report
The AI receives all the individual reports and produces an overall analysis or summary.
Important:
Validate the AI’s results and keep its limitations in mind. Treat the output as support, not as a final automatic decision.
reports_content = "\n".join(
f"Report for {r['log_file']}:\n{r['report']}" for r in individual_reports
)
final_messages = [
SystemMessage(content=final_system_prompt),
HumanMessage(content=reports_content)
]
try:
final_response = llm(final_messages)
final_report = final_response.content if hasattr(final_response, 'content') else str(final_response)
except Exception as e:
logging.error(f"Error during final aggregation OpenAI call: {e}")
final_report = "Aggregation failed due to error."
9. Saving the reports and setting permissions
The final and individual reports are stored in .txt and .json files, with the appropriate permissions.
Recommendation:
Consider encrypting the reports if they contain sensitive information, and set up automatic backups in a secure environment.
try:
with REPORT_TXT.open("a", encoding="utf-8") as f:
f.write("\n===== FINAL AGGREGATED REPORT =====\n")
f.write(final_report)
with REPORT_JSON.open("w", encoding="utf-8") as f:
json.dump({
"individual_reports": individual_reports,
"final_report": final_report
}, f, indent=2)
except Exception as e:
logging.error(f"Error saving final report: {e}")
def set_permissions():
"""Set secure permissions on report files."""
try:
REPORT_JSON.chmod(0o640)
REPORT_TXT.chmod(0o640)
except Exception as e:
logging.error(f"Error setting permissions on report files: {e}")
10. Automatic email delivery (SMTP)
It sends the security report by email using SMTP.
Warning:
Make sure you use app passwords and TLS/SSL connections. Don’t use personal accounts or expose credentials in insecure files.
def send_email_report(report_path: Path):
"""Send the TXT report as the email body using smtplib."""
SMTP_SERVER = os.getenv('SMTP_SERVER')
SMTP_PORT = int(os.getenv('SMTP_PORT', '587'))
SMTP_USERNAME = os.getenv('SMTP_USERNAME')
SMTP_PASSWORD = os.getenv('SMTP_PASSWORD')
if not SMTP_SERVER or not MAIL_TO or not MAIL_FROM:
logging.error("SMTP configuration or email addresses missing.")
logging.info(f"Reports saved at:\n JSON: {REPORT_JSON}\n Text: {REPORT_TXT}")
sys.exit(3)
try:
with report_path.open("r", encoding="utf-8") as f:
report_content = f.read()
msg = EmailMessage()
msg['Subject'] = MAIL_SUBJECT
msg['From'] = MAIL_FROM
msg['To'] = MAIL_TO
msg.set_content(report_content)
with smtplib.SMTP(SMTP_SERVER, SMTP_PORT) as server:
server.starttls()
if SMTP_USERNAME and SMTP_PASSWORD:
server.login(SMTP_USERNAME, SMTP_PASSWORD)
server.send_message(msg)
logging.info(f"Security report emailed to {MAIL_TO} via SMTP")
except Exception as e:
logging.error(f"Failed to email report via SMTP: {e}")
11. Archiving and clean-up
All logs and reports are compressed into a ZIP file, and old files are deleted according to the retention policy.
Important:
Regularly check disk space and the state of the archived files, and monitor system performance, especially if the log volume is high.
def archive_and_cleanup():
"""Archive logs and reports, and clean up old files."""
zip_name = f"security-logs-{datetime.datetime.now().strftime('%Y%m%d-%H%M%S')}.zip"
logs_to_archive = list(LOG_DIR.glob("*.txt")) + list(LOG_DIR.glob("*.log")) + [REPORT_JSON]
files_to_archive = [f for f in logs_to_archive if f.is_file()]
archive_path = ARCHIVE_DIR / zip_name
if files_to_archive:
try:
with zipfile.ZipFile(archive_path, 'w', zipfile.ZIP_DEFLATED) as zipf:
for file in files_to_archive:
zipf.write(file, arcname=file.name)
for file in files_to_archive:
file.unlink()
logging.info(f"Archived and removed: {' '.join(str(f) for f in files_to_archive)}")
except Exception as e:
logging.error(f"Error archiving files: {e}")
else:
logging.info(f"No .txt, .log, or .json files to archive in {LOG_DIR}.")
# Remove old archives
now = datetime.datetime.now().timestamp()
for f in ARCHIVE_DIR.glob("*.zip"):
try:
if f.stat().st_mtime < now - ARCHIVE_RETENTION_DAYS * 86400:
f.unlink()
except Exception as e:
logging.error(f"Error deleting archive {f}: {e}")
# Remove other files/directories in LOG_DIR except .log_offsets
for f in LOG_DIR.iterdir():
if f.name == '.log_offsets':
continue
try:
if f.is_dir():
shutil.rmtree(f)
else:
f.unlink()
except Exception as e:
logging.error(f"Error removing {f}: {e}")
12. System prompts used
// MAIN LLM - This will handle the final report
/*
Act as a senior security analyst.
You are given output from an AI analyst.
Logs are from Fail2ban, ClamAV, RKHunter, Lynis, Maldet,
and SSH authentication logs from /var/log/security-scans.
Tasks:
Vulnerabilities: Identify misconfigurations, outdated practices,
open attack surfaces, or missing controls in the system.
Provide severity, verbatim evidence, and remediation steps for each finding.
Detections: Extract concrete security events, such as repeated failed SSH attempts,
or banned IP address activity.
Include first_seen, last_seen, counts, and impacted entities as evidence.
Patterns & Analysis: Summarize attacker behavior patterns,
such as targeted usernames and source IP ranges observed.
Include defenses triggered, like bans, lockouts, and possible evasion indicators.
Executive Summary: Provide a concise, non-technical summary of risk posture,
including urgent recommended actions if found.
Appendices:
a. List Indicators of Compromise (IPs, users, keys)
and note any data coverage limitations encountered in analysis.
b. Output a single JSON object containing:
{
“executive_summary”: “…”,
“detections”: [
{
“type”: “…”,
“details”: “…”,
“evidence”: [“…”],
“first_seen”: “…”,
“last_seen”: “…”,
“count”: 0,
“severity”: “low|medium|high|critical”
}
],
“vulnerabilities”: [
{
“issue”: “…”,
“severity”: “…”,
“evidence”: [“…”],
“remediation”: “…”
}
],
“patterns_analysis”: {
“attacker_behaviors”: “…”,
“defenses”: “…”,
“timeline”: “…”,
“notable_patterns”: “…”
},
“ioc”: {
“ips”: [“…”],
“users”: [“…”],
“other”: [“…”]
},
“limitations”: “…”
}
c. Use only concrete evidence from logs
and quote relevant lines in the evidence arrays.
*/
// JUNIOR LLM - This will handle individual log analysis
/*
You are a Linux security log analyst.
Review a single Linux security log entry for issues.
Identify security issues, anomalies, or suspicious activities present in the log.
Summarize findings concisely and clearly for further review.
Focus on:
All relevant security events or indicators in the log entry.
Potential impact or risk level identified from the analysis.
Actionable insights or next steps, if applicable, for remediating issues.
Do not include the full log entry in your response.
Avoid unnecessary verbosity to keep the summary focused and clear.
Your summary will be used by senior analysts for review and aggregation.
Format response as:
Findings:
Insights:
Recommendations (if any):
*/
Example: Bash script to create custom activity logs for analysis
- Change
/var/log/fail2ban.log, etc. to match your distribution’s log paths if needed. - For privacy, avoid running or sharing this on systems with sensitive or regulated data without proper review.
#!/usr/bin/env bash
set -euo pipefail
# Set output directory in a generic temp/logs path
OUT_DIR="/tmp/security-scans"
TS="$(date +'%Y%m%d-%H%M%S')"
mkdir -p "$OUT_DIR"
# Helper to safely run commands that might not exist
have() { command -v "$1" >/dev/null 2>&1; }
# -----------------------------
# Fail2ban: jails and bans (last 48h from fail2ban.log)
# -----------------------------
if have fail2ban-client && [ -r /var/log/fail2ban.log ]; then
{
echo "== Fail2ban bans (last 48 hours) =="
awk -v d="$(date --date='-48 hours' '+%Y-%m-%d %H:%M:%S')" \
'$0 >= d && /Ban /' /var/log/fail2ban.log
} > "$OUT_DIR/fail2ban-bans-activity-$TS.txt" 2>&1
fi
# -----------------------------
# SSH: successful and failed logins (last 48h)
# -----------------------------
AUTH_FILE=""
for cand in /var/log/auth.log /var/log/secure; do
if [ -r "$cand" ]; then AUTH_FILE="$cand"; break; fi
done
if [ -n "$AUTH_FILE" ]; then
{
echo "== SSH successful logins (last 48h) =="
for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Accepted (password|publickey)" || true
done
echo
echo "== SSH failed logins (last 48h) =="
for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Failed password" || true
done
echo
echo "== SSH invalid users (last 48h) =="
for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Invalid user" || true
done
echo
echo "== SSH root or sudo session opens (last 48h) =="
for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: (pam_unix|session opened|Accepted .+ for root)" || true
done
} > "$OUT_DIR/ssh-logins-$TS.txt" 2>&1
# Compact summaries
{
echo "== Summary (counts, last 48h) =="
echo -n "Accepted logins: "
count=0
for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
count=$((count + $(awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Accepted (password|publickey)" | wc -l)))
done
echo "$count"
echo -n "Failed passwords: "
count=0
for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
count=$((count + $(awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Failed password" | wc -l)))
done
echo "$count"
echo -n "Invalid users: "
count=0
for day in $(date --date='-0 days' '+%b %e') $(date --date='-1 days' '+%b %e'); do
count=$((count + $(awk -v d="$day" '$0 ~ d' "$AUTH_FILE" | grep -E "sshd\[.*\]: Invalid user" | wc -l)))
done
echo "$count"
} > "$OUT_DIR/ssh-summary-$TS.txt" 2>&1
fi
# -----------------------------
# Permissions
# -----------------------------
chmod 640 "$OUT_DIR"/*"$TS".txt || true
echo "Security scans written to: $OUT_DIR (timestamp $TS)"
Complete code
#!/usr/bin/env python3
import os
import sys
import glob
import shutil
import socket
import datetime
import zipfile
import requests
import json
import logging
from dotenv import load_dotenv
from langchain_openai import AzureChatOpenAI
from langchain.schema import SystemMessage, HumanMessage
from pathlib import Path
import smtplib
from email.message import EmailMessage
# ============================
# Constants & Configuration
# ============================
ARCHIVE_RETENTION_DAYS = 30
LOG_DIR = Path(os.getenv('LOG_DIR', '/var/log/security'))
OUT_DIR = Path(os.getenv('OUT_DIR', '/var/log/security'))
ARCHIVE_DIR = Path(os.getenv('ARCHIVE_DIR', '/var/log/archived'))
ENV_FILE = os.environ.get('ENV_FILE', './.env')
SYSTEM_PROMPT_FILE = Path('system_prompt.txt')
JUNIOR_SYSTEM_PROMPT_FILE = Path('junior_system_prompt.txt')
TIMESTAMP = datetime.datetime.now().strftime('%Y%m%d-%H%M%S')
REPORT_JSON = OUT_DIR / f'security-report-{TIMESTAMP}.json'
REPORT_TXT = OUT_DIR / f'security-report-{TIMESTAMP}.txt'
MAIL_TO = os.getenv('MAIL_TO', 'info@email.com')
MAIL_SUBJECT = os.getenv('MAIL_SUBJECT', f'Security Report {TIMESTAMP}')
MAIL_FROM = os.getenv('MAIL_FROM', f"security-reports@{socket.getfqdn() if hasattr(socket, 'getfqdn') else socket.gethostname()}")
REQUIRED_ENV = [
'AZURE_OPENAI_ENDPOINT',
'AZURE_OPENAI_DEPLOYMENT',
'API_VERSION',
'AZURE_TENANT_ID',
'AZURE_CLIENT_ID',
'AZURE_CLIENT_SECRET',
]
# ============================
# Logging Setup
# ============================
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s [%(levelname)s] %(message)s',
handlers=[logging.StreamHandler(sys.stderr)]
)
# ============================
# Helper Functions
# ============================
def load_environment():
"""Load environment variables from dotenv file and check required ones."""
if Path(ENV_FILE).exists():
load_dotenv(dotenv_path=ENV_FILE)
logging.info(f"Loaded environment from {ENV_FILE}")
else:
logging.warning(f".env file not found at {ENV_FILE}; relying on current environment.")
missing_vars = [var for var in REQUIRED_ENV if not os.getenv(var)]
if missing_vars:
logging.error(f"Missing required environment variables: {', '.join(missing_vars)}")
sys.exit(1)
def ensure_directories():
"""Ensure log, output, and archive directories exist."""
if not LOG_DIR.is_dir():
logging.error(f"Log directory not found: {LOG_DIR}")
sys.exit(1)
OUT_DIR.mkdir(parents=True, exist_ok=True)
ARCHIVE_DIR.mkdir(parents=True, exist_ok=True)
def get_azure_ad_token():
"""Acquire Azure AD Token using client credentials."""
token_url = f"https://login.microsoftonline.com/{os.environ['AZURE_TENANT_ID']}/oauth2/v2.0/token"
token_data = {
"grant_type": "client_credentials",
"client_id": os.environ['AZURE_CLIENT_ID'],
"client_secret": os.environ['AZURE_CLIENT_SECRET'],
"scope": os.environ.get('AZURE_OPENAI_SCOPE', 'https://cognitiveservices.azure.com/.default'),
}
response = requests.post(token_url, data=token_data)
response.raise_for_status()
return response.json()["access_token"]
def load_prompt(prompt_file: Path) -> str:
"""Load a prompt from a file."""
if not prompt_file.is_file():
logging.error(f"{prompt_file} not found!")
sys.exit(1)
with prompt_file.open("r", encoding="utf-8") as f:
return f.read()
def collect_log_files() -> list:
"""Collect log files (*.log, *.txt) from LOG_DIR."""
log_files = list(LOG_DIR.glob("*.log")) + list(LOG_DIR.glob("*.txt"))
return log_files
def send_email_report(report_path: Path):
"""Send the TXT report as the email body using smtplib."""
SMTP_SERVER = os.getenv('SMTP_SERVER')
SMTP_PORT = int(os.getenv('SMTP_PORT', '587'))
SMTP_USERNAME = os.getenv('SMTP_USERNAME')
SMTP_PASSWORD = os.getenv('SMTP_PASSWORD')
if not SMTP_SERVER or not MAIL_TO or not MAIL_FROM:
logging.error("SMTP configuration or email addresses missing.")
logging.info(f"Reports saved at:\n JSON: {REPORT_JSON}\n Text: {REPORT_TXT}")
sys.exit(3)
try:
with report_path.open("r", encoding="utf-8") as f:
report_content = f.read()
msg = EmailMessage()
msg['Subject'] = MAIL_SUBJECT
msg['From'] = MAIL_FROM
msg['To'] = MAIL_TO
msg.set_content(report_content)
with smtplib.SMTP(SMTP_SERVER, SMTP_PORT) as server:
server.starttls()
if SMTP_USERNAME and SMTP_PASSWORD:
server.login(SMTP_USERNAME, SMTP_PASSWORD)
server.send_message(msg)
logging.info(f"Security report emailed to {MAIL_TO} via SMTP")
except Exception as e:
logging.error(f"Failed to email report via SMTP: {e}")
def archive_and_cleanup():
"""Archive logs and reports, and clean up old files."""
zip_name = f"security-logs-{datetime.datetime.now().strftime('%Y%m%d-%H%M%S')}.zip"
logs_to_archive = list(LOG_DIR.glob("*.txt")) + list(LOG_DIR.glob("*.log")) + [REPORT_JSON]
files_to_archive = [f for f in logs_to_archive if f.is_file()]
archive_path = ARCHIVE_DIR / zip_name
if files_to_archive:
try:
with zipfile.ZipFile(archive_path, 'w', zipfile.ZIP_DEFLATED) as zipf:
for file in files_to_archive:
zipf.write(file, arcname=file.name)
for file in files_to_archive:
file.unlink()
logging.info(f"Archived and removed: {' '.join(str(f) for f in files_to_archive)}")
except Exception as e:
logging.error(f"Error archiving files: {e}")
else:
logging.info(f"No .txt, .log, or .json files to archive in {LOG_DIR}.")
# Remove old archives
now = datetime.datetime.now().timestamp()
for f in ARCHIVE_DIR.glob("*.zip"):
try:
if f.stat().st_mtime < now - ARCHIVE_RETENTION_DAYS * 86400:
f.unlink()
except Exception as e:
logging.error(f"Error deleting archive {f}: {e}")
# Remove other files/directories in LOG_DIR except .log_offsets
for f in LOG_DIR.iterdir():
if f.name == '.log_offsets':
continue
try:
if f.is_dir():
shutil.rmtree(f)
else:
f.unlink()
except Exception as e:
logging.error(f"Error removing {f}: {e}")
def set_permissions():
"""Set secure permissions on report files."""
try:
REPORT_JSON.chmod(0o640)
REPORT_TXT.chmod(0o640)
except Exception as e:
logging.error(f"Error setting permissions on report files: {e}")
# ============================
# Main Functionality
# ============================
def main():
load_environment()
ensure_directories()
# Prepare LangChain AzureChatOpenAI
llm = AzureChatOpenAI(
azure_deployment=os.environ["AZURE_OPENAI_DEPLOYMENT"],
api_version=os.environ["API_VERSION"],
azure_endpoint=os.environ["AZURE_OPENAI_ENDPOINT"],
azure_ad_token_provider=get_azure_ad_token,
temperature=0.7,
)
# Load prompts
final_system_prompt = load_prompt(SYSTEM_PROMPT_FILE)
per_log_system_prompt = load_prompt(JUNIOR_SYSTEM_PROMPT_FILE)
# Collect log files
log_files = collect_log_files()
individual_reports = []
# Query OpenAI for each log file
for log_file in log_files:
try:
with log_file.open("r", encoding="utf-8") as f:
log_content = f.read()
messages = [
SystemMessage(content=per_log_system_prompt),
HumanMessage(content=log_content)
]
response = llm(messages)
report_text = response.content if hasattr(response, 'content') else str(response)
individual_reports.append({
"log_file": log_file.name,
"report": report_text
})
except Exception as e:
logging.error(f"Error processing {log_file}: {e}")
# Save individual reports
individual_reports_txt = ""
for report in individual_reports:
individual_reports_txt += f"===== Report for {report['log_file']} =====\n{report['report']}\n\n"
try:
with REPORT_TXT.open("w", encoding="utf-8") as f:
f.write(individual_reports_txt)
except Exception as e:
logging.error(f"Error writing individual reports: {e}")
# Final aggregation OpenAI call
reports_content = "\n".join(
f"Report for {r['log_file']}:\n{r['report']}" for r in individual_reports
)
final_messages = [
SystemMessage(content=final_system_prompt),
HumanMessage(content=reports_content)
]
try:
final_response = llm(final_messages)
final_report = final_response.content if hasattr(final_response, 'content') else str(final_response)
except Exception as e:
logging.error(f"Error during final aggregation OpenAI call: {e}")
final_report = "Aggregation failed due to error."
# Append final aggregated report
try:
with REPORT_TXT.open("a", encoding="utf-8") as f:
f.write("\n===== FINAL AGGREGATED REPORT =====\n")
f.write(final_report)
with REPORT_JSON.open("w", encoding="utf-8") as f:
json.dump({
"individual_reports": individual_reports,
"final_report": final_report
}, f, indent=2)
except Exception as e:
logging.error(f"Error saving final report: {e}")
set_permissions()
send_email_report(REPORT_TXT)
archive_and_cleanup()
if __name__ == "__main__":
main()
Summary
By orchestrating open source tools and using a large language model (LLM) to analyse and report on logs, this script turns a set of utilities into an AI-powered security platform. The approach automates log review, puts findings in context and keeps administrators informed, all with minimal manual effort.
However, remember:
- AI may miss advanced threats and should be an aid, not the only line of defence.
- Protect the files and credentials.
- Validate and review the reports.
- Monitor the system and tune parameters to the load and log volume.
This architecture is a practical example of how AI can boost infrastructure management, bringing clarity, efficiency and actionable intelligence to the noisy world of security logs.
