Agent Architecture — Production Infrastructure
Translated from the Spanish original. Read in Spanish
This article describes one possible design for production-grade agent infrastructure on AWS.
There’s no single right architecture for running an AI agent in production. The right design depends on your scale, compliance requirements and how much operational complexity you’re willing to take on. What follows is one concrete option — not the only way, but a well-founded, production-grade approach that reflects real infrastructure decisions.
Compute — AWS ECS Fargate
Two clusters: production (always on, highly available) and development (50% uptime, scales to zero outside working hours). Each runs two services.
- The agent application handles the LLM calls, orchestration of the different nodes, tool coordination and an admin interface. For RAG, it queries
pgvector. For CAG, it loads the full-context payloads directly. Sized at1 vCPU / 2 GB. - The MCP proxy abstracts the orchestration of the different MCP servers, routing tool calls to external MCPs. It’s decoupled from the agent runtime — so you can add or swap MCPs without a new deployment of the main application. Also
1 vCPU / 2 GB.
High Availability in Production
A minimum of two tasks spread across multiple Availability Zones. Unhealthy tasks are deregistered and replaced automatically. Rolling deployments run at 100% minimum healthy / 200% maximum — zero downtime. Auto Scaling adjusts the task count based on CPU and memory targets via CloudWatch.
Development Cluster
The development cluster does more than serve as staging. Training, evaluation and integration tests run here within the same 50% uptime window. No separately scheduled tasks, no extra compute cost. It scales to zero outside working hours through ECS service scheduling.
CI/CD — GitLab
GitLab drives the whole deployment pipeline. Commits to the staging branch trigger the Eval Pipeline in the development cluster. Passing the evaluation unlocks the deployment stage, which promotes the changes to production through an ECS rolling update. Prompt changes, model updates and agent configuration changes all go through the same pipeline — nothing reaches production without first passing an automated evaluation. We assume GitLab is part of the existing platform, so there’s no additional infrastructure cost.
In a DR failover event, the GitLab pipeline is redirected to the DR region. Same pipeline definition, same checks, different deployment target.
KB Ingestion — AWS Lambda
The ingestion pipeline runs serverless. It turns raw source data into clean, structured content, writing embeddings to pgvector for RAG and preparing full-context payloads for CAG. Zero cost at rest, billed per invocation.
Observability — Arize Phoenix
Phoenix runs as a self-hosted Fargate service built specifically for LLM observability. It traces the whole agent chain — LLM calls, tool execution via MCP, retrieval paths and model responses. Sized at 0.5 vCPU / 1 GB. It stores all trace data in the existing RDS PostgreSQL instance — no separate storage layer needed.
All telemetry stays in-VPC. No data leaves for a third-party SaaS.
CloudWatch handles the infrastructure layer — ECS metrics, Auto Scaling target tracking and alarms for task failures or RDS thresholds. Phoenix handles the LLM chain. The two complement each other rather than overlap.
Storage — One RDS Instance, Four Roles
Storage Strategy: a single RDS instance
Consolidating app configuration, conversational memory, vector storage and Phoenix traces on a single db.m7g.large instance (2 vCPU, 8 GB RAM) is a decision that balances cost against real load capacity. Unlike burstable instances in the t4g family, the m7g.large delivers its full 2 vCPUs on a sustained basis — no CPU credits, no degradation under continuous load. That’s critical when pgvector’s vector searches compete with Phoenix’s telemetry writes, since both workloads are CPU-intensive and can coincide under real concurrency.
- Capacity: With 2 dedicated vCPUs and 8 GB of RAM, this instance comfortably supports up to 80 concurrent users running vector searches, telemetry writes and conversational memory queries at the same time. The 8 GB of RAM also lets pgvector keep a significant portion of the vector index in memory, reducing pressure on the disk.
- Scaling criterion: If sustained concurrent traffic consistently exceeds 80 users, the natural next step is to move Phoenix to its own instance (e.g.
db.t4g.small— telemetry isn’t on the request’s critical path, so a burstable instance is enough) and consider vertically scaling the main instance todb.m7g.xlarge. Alternatively, migrate pgvector to a dedicated vector database if the index grows beyond what a single PostgreSQL instance can serve efficiently.
The PostgreSQL instance (db.m7g.large, 50 GB gp3) has daily snapshots and a Multi-AZ standby replica. If it fails, RDS fails over automatically to the standby replica with no data loss.
Networking — ALB with SSO
A single ALB sits in front of the Admin Interface for agent configuration, prompt editing, tool management and human-in-the-loop review. TLS termination and OIDC authentication through Okta, or any other identity provider (IdP). The ALB’s health-check endpoint is monitored by Route 53 as the trigger for DR failover.
Disaster Recovery (DR) — Multi-Region Pilot Light
The primary stack runs in us-east-1. The DR region is us-west-2.
Asynchronous cross-region RDS replication feeds a read replica in us-west-2. Replication lag is usually under a minute, giving an RPO of under a minute. In a regional failover, the replica is promoted to an independent primary database and the DR stack points to it.
All Fargate clusters in the DR region run with a desired count of zero during normal operations — no tasks are running; application compute starts up during failover while the rest of the DR infrastructure stays pre-provisioned. This is the pilot light pattern. Container images are available through ECR cross-region replication. Route 53 health checks monitor the primary ALB. On failure, DNS switches to the DR ALB with a 60-second TTL. GitLab redirects to us-west-2. ECS scales up the DR clusters. Phoenix in the DR region comes up and points to the promoted replica. The KB Lambda pipeline in us-west-2 — already deployed identically — starts writing to the promoted database.
The RTO with this setup is about 15 minutes, depending on how quickly the failure is detected and the DR sequence is triggered. If the SLA demands an RTO under five minutes, the DR Fargate clusters should run with a desired count of one (warm standby) instead of zero, which adds about $72 a month.
An active-active multi-region design is a completely different architecture — it requires distributed state management and conflict resolution for RDS writes. For most agent workloads, pilot light DR with an RTO of about 15 minutes is the right balance.
The Continuous Improvement Loop
Without a feedback loop, the agent drifts. Prompt quality degrades, knowledge goes stale and nobody notices until users complain. This loop runs in the development cluster and gates every change behind automated evaluation and human approval.
- Interaction Analysis: Pull conversation logs from RDS and trace data from Phoenix. An LLM grades interaction quality and surfaces degradation patterns — low-confidence answers, fallbacks, failed tool calls, user corrections, abandoned sessions.
- Prompt and CAG Tuning: Propose changes to system instructions, few-shot examples, guardrails and prompt templates. For CAG, trigger the KB Lambda to regenerate stale context payloads. All changes are versioned and committed to a staging branch in GitLab.
- Automated Evaluation: GitLab detects the commit on staging and triggers the Eval Pipeline in the development cluster. Golden dataset regression, A/B comparison, latency benchmarks, tool-call correctness. If it passes — the deploy stage is unlocked. If it fails — the pipeline stops.
- Human-in-the-Loop: Evaluation results, diffs and before/after comparisons are shown in the Admin Interface via the ALB. Explicit approval is required. No exceptions.
- Promotion to Production: GitLab runs the ECS rolling update to the production cluster in
us-east-1. Phoenix monitors the first hours of traffic. Automatic rollback if a regression is detected.
Production interactions feed the next cycle. The loop sustains itself.
Monthly Infrastructure Cost
Primary Region (us-east-1):
- Fargate Production (2 services × 2 tasks, 1 vCPU / 2 GB) — $144.16
- Fargate Development (2 services × 1 task, 50% uptime) — $36.04
- Phoenix — $18.02
- Lambda — ~$1.00
- RDS Multi-AZ + pgvector (db.m7g.large, 50 GB gp3) — $256.78
- ALB + CloudWatch + ECR + Secrets — $28.03
- Primary Total — ~$484/month
DR Region (us-west-2, pilot light):
- RDS cross-region read replica (db.m7g.large, 50 GB gp3) — $128.39
- Pre-provisioned ALB — $16.00
- Route 53 health checks + failover routing — $4.00
- ECR cross-region replication — ~$1.00
- Fargate clusters (desired: 0, idle) — $0.00
- Phoenix DR + Lambda DR (idle) — $0.00
- Additional DR — ~$149/month
Total with DR — ~$633/month
The cross-region RDS replica accounts for most of the DR cost. Fargate costs nothing while no tasks are running. Activating DR compute in a real failover adds the Fargate and Phoenix costs for that month, but they’re one-off spikes rather than ongoing spend.
The cost difference compared with an equivalent burstable Multi-AZ instance (db.t4g.medium, ~$106/month including storage) is about $150/month in the primary region. That increase completely removes the risk of degradation from exhausted CPU credits — a justified trade-off when you need predictable performance under real concurrency.
A Note on Load Capacity
This base architecture, with its initial sizing (1 vCPU / 2 GB Fargate containers and a db.m7g.large RDS instance), is designed for a low to moderate load with the capacity to absorb real concurrency spikes.
- Normal Traffic: For 100 to 200 total users with occasional or spread-out use, this design is highly efficient, fault-tolerant and will keep costs in check.
- Concurrent Traffic: The
db.m7g.largeinstance supports up to 80 users interacting at the same time without performance degradation, thanks to its 2 dedicated (non-burstable) vCPUs and 8 GB of RAM. pgvector vector searches, Phoenix telemetry writes and conversational memory queries can run in parallel without competing for CPU credits. The Fargate compute layer will scale smoothly through ECS Auto Scaling to absorb the corresponding load. - Beyond 80 concurrent users: If sustained concurrency consistently exceeds this threshold, the first step is to move Phoenix to its own RDS instance and consider vertically scaling the main instance to
db.m7g.xlarge(4 vCPU, 16 GB RAM).
Stack Summary
Two Fargate clusters (Production HA, Development 50%). One Fargate service for Phoenix. One Lambda function. One RDS instance (Multi-AZ, db.m7g.large, pgvector, four roles). One ALB with SSO. GitLab for CI/CD and deployment gating. CloudWatch for infrastructure. Route 53 for DNS failover. Multi-region pilot light DR in us-west-2. Human approval validating every production change.

This is one way to build it. The components are interchangeable — pgvector could be replaced with a dedicated vector database if scale demands it, Phoenix could be swapped for another OTel-compatible observability tool, and the pilot light could be scaled up to warm standby or active-active if the SLA requires it. The principles behind the design — separation of concerns, controlled deployments, self-hosted telemetry, continuous improvement with human oversight — apply regardless of the specific tools chosen.
