Agent Observability Platform: Prerequisites (AWS)
What you need before deploying the Agent Observability data platform on AWS
This page walks through everything you need before running the installation: the right tooling, an AWS account with sufficient permissions, a deployment path, domains, and the platform artifacts. Work through the steps in order, then run the preflight check to confirm you are ready.
1. Install the required tooling
Install the following on the machine you will run Terraform from:
| Tool | Version | Used for |
|---|---|---|
| Terraform | >= 1.3 | Provisioning the platform |
| AWS CLI | latest | Authenticating to AWS and retrieving outputs |
| kubectl | latest | Verifying the cluster after deployment |
| Helm | 3.x | Only for the self-managed Helm path |
Configure the AWS CLI with credentials for your target account before continuing.
2. Prepare your AWS account
You deploy into your own AWS account. The principal running terraform apply needs permissions to create and manage:
- VPC and networking (subnets, NAT gateways, route tables, security groups) โ for the new-cluster path
- EKS cluster and managed node groups
- IAM roles, policies, and an OIDC provider (for IRSA)
- KMS keys and AWS Secrets Manager secrets
- ACM certificates and Route 53 records
The platform artifacts are public (Terraform Registry and Docker Hub), so no registry credentials or cross-account pull grants are needed โ only outbound access to Docker Hub: the helm provider pulls the ao-data-platform chart during apply, and the EKS nodes pull the ao-llm-worker image at runtime.
The principal that runs terraform apply becomes the cluster's administrator automatically (the module sets enable_cluster_creator_admin_permissions = true), so the same credentials you deploy with can administer the cluster afterward.
AWS services and actions the deploy uses
The table below lists the AWS services and representative actions the deploy exercises, by category. It reflects what the module and its upstream VPC/EKS modules create โ treat it as the breadth of access the deploying principal needs, not a hand-minimized least-privilege policy.
| Category | Service | Representative actions |
|---|---|---|
| VPC & networking (new-cluster path) | ec2 | Create*/Delete*/Describe*/Modify* on Vpc, Subnet, RouteTable, Route, InternetGateway, NatGateway, Address (EIP), SecurityGroup/SecurityGroupRules; CreateTags/DeleteTags |
| EKS cluster & node groups | eks | CreateCluster, DescribeCluster, CreateNodegroup, CreateAddon, CreateAccessEntry, AssociateAccessPolicy, TagResource; plus ec2 launch-template/ENI actions and iam:CreateServiceLinkedRole for EKS |
| IAM & OIDC (IRSA) | iam | CreateRole, GetRole, CreatePolicy, AttachRolePolicy, PutRolePolicy, CreateOpenIDConnectProvider, GetOpenIDConnectProvider, TagRole/TagPolicy, PassRole |
| KMS | kms | CreateKey, CreateAlias, DescribeKey, EnableKeyRotation, PutKeyPolicy, ScheduleKeyDeletion, TagResource |
| Secrets Manager | secretsmanager | CreateSecret, PutSecretValue, DescribeSecret, GetSecretValue, TagResource, DeleteSecret |
| ACM & Route 53 | acm, route53 | acm:RequestCertificate/DescribeCertificate/AddTagsToCertificate; route53:ListHostedZones/GetHostedZone/ChangeResourceRecordSets/GetChange (DNS-01 validation) |
| Identity | sts | GetCallerIdentity; eks get-token for the kubernetes/helm providers (which then act against the cluster's Kubernetes API using the cluster-admin access above โ not via IAM) |
For a tuned least-privilege policy, generate one from a captured terraform apply (for example with IAM Access Analyzer policy generation from CloudTrail).
3. Choose your deployment path
The Terraform module supports two paths. Decide which applies before configuring it:
| Path | When to use | Additional inputs |
|---|---|---|
| New cluster | You want Terraform to create the VPC and EKS cluster. | None โ uses defaults (cluster name monte-carlo; the created VPC spans the three AZs the default HA topology needs). |
| Existing cluster | You already run an EKS cluster and want to deploy the platform into it. | cluster.existing_cluster_name, networking.existing_vpc_id, and networking.existing_private_subnet_ids in different Availability Zones โ three AZs for the default HA topology (two minimum for a single-replica deployment via the self-managed Helm install), one subnet per AZ. |
Whichever path you choose, confirm the cluster has a NetworkPolicy engine enabled โ on EKS, the VPC CNI's network policy agent. The platform isolates ClickHouse Keeper's client port with a Kubernetes NetworkPolicy that's only enforced when such an agent is present; see High availability.
Existing cluster โ three things to check:
OIDC provider. If your cluster's OIDC provider was created outside Terraform, import it with an
importblock in your root module โ it is applied along with the rest of the configuration:import { to = module.ao_data_platform.aws_iam_openid_connect_provider.cluster[0] id = "<oidc-provider-arn>" }The
terraform importcommand does not work for this resource: it fails while evaluating the module's certificate-validation records, whose values are only known during apply.Existing controllers. If cert-manager, the External Secrets Operator, the AWS Load Balancer Controller, or external-dns are already installed, set the matching
helm.install_*flag tofalseto skip reinstalling them (see Installation). A pre-installed External Secrets Operator needs read access to the platform's secrets โ the module grants it only to an operator it installs.HA node groups. The module creates the per-AZ ClickHouse and Keeper node groups only on the new-cluster path. On an existing cluster you attach them yourself โ see High availability.
Bringing your own EKS cluster with the Generic agent
If you provision the EKS cluster yourself and run the Generic agent in it alongside the platform, the cluster must meet the requirements of both:
| Requirement | Needed by | Notes |
|---|---|---|
| VPC with private subnets in distinct AZs (three for the HA topology), outbound internet (NAT) or VPC endpoints, DNS hostnames enabled | Platform + Agent | The Agent needs S3, Secrets Manager, and STS โ via NAT or VPC endpoints |
| IAM OIDC provider for the cluster (IRSA) | Platform + Agent | Import it into the platform root if you created it outside Terraform (see above) |
| IRSA for the Agent, with a role you supply | Agent | Required when sharing the platform's cluster โ see Agent identity. Never mix IRSA and Pod Identity on one service account |
| NetworkPolicy engine (VPC CNI network policy agent) | Platform | Enforces the ClickHouse Keeper NetworkPolicy |
| Amazon EBS CSI driver, with an IRSA role for its controller | Platform | Provisions the ClickHouse and Keeper volumes. The module installs it only on clusters it creates โ without it, both volumes stay Pending |
| Dedicated, tainted node groups for ClickHouse and Keeper | Platform | dedicated=clickhouse:NoSchedule and dedicated=keeper:NoSchedule. On an existing cluster the module doesn't set node selectors or tolerations, so use the self-managed Helm install to place the pods |
| General-purpose node capacity | Platform + Agent | For the Collector, LLM worker, controllers, and the Agent pods |
| Controllers: External Secrets Operator, cert-manager, AWS Load Balancer Controller, external-dns | Platform | Pre-install them and set the helm.install_* flags to false, or let the module install them. A pre-installed ESO needs read access to the platform's secrets. The Agent reuses the platform's ESO (helm.install_external_secrets_operator = false in the Agent module) |
| Outbound access to Docker Hub and the Monte Carlo agent service | Platform + Agent | For the Agent, see AWS PrivateLink as an alternative to the internet |
Apply in this order: the cluster (with its OIDC provider, node groups, and the Agent's IRSA role), then the platform root, then the Agent root.
4. Configure domains and DNS
The OpenTelemetry Collector and ClickHouse are each exposed through a Network Load Balancer (NLB) with a DNS name and TLS. Decide on:
- A domain for the OpenTelemetry Collector endpoint (e.g.
otel.acme.com) โ passed asotel_collector_domain - A domain for the ClickHouse endpoint (e.g.
clickhouse.acme.com) โ passed asclickhouse_domain - A Route 53 hosted zone ID (
hosted_zone_id) covering those domains
When hosted_zone_id is set, Terraform creates the IRSA roles for cert-manager (ACME DNS-01 validation) and external-dns (automatic CNAME management). Leave it unset to manage DNS records manually.
Both domains must be names inside the zone that
hosted_zone_ididentifies โ for zoneacme.com,otel.acme.comandclickhouse.acme.comwork, but a zone forao.acme.comdoes not coverotel.acme.com. A domain outside the zone isn't rejected up front: its certificate-validation records are created with the zone name appended, the certificates never validate, and the apply fails late withmissing โฆ DNS validation recordโ see Troubleshooting.
5. Enable Bedrock model access
The LLM worker runs evaluations against Claude models on Amazon Bedrock. The module grants the worker IAM permission to invoke them, but each account must enable access to the Claude models once, per region, in the Bedrock console under Model access โ there is no Terraform resource or API for this step. Enable them in the region the worker calls (helm.llm_worker.bedrock_region, default your deployment region). Do it ahead of the deploy, or the worker's evaluation calls fail until access lands.
6. Get the platform artifacts
All three artifacts are public โ no Monte Carlo-issued credentials or registry login are required:
| Artifact | Location | Version |
|---|---|---|
| Terraform module | Terraform Registry โ monte-carlo-data/ao-data-platform/aws | 2.3.0 |
ao-data-platform Helm chart | Docker Hub (OCI) โ oci://registry-1.docker.io/montecarlodata/ao-data-platform | 4.0.0 |
ao-llm-worker image | Docker Hub โ montecarlodata/ao-llm-worker | 1.1.0-aws |
The worker image is published with per-cloud tags: pinned releases like 1.1.0-aws, plus a floating latest-aws tag. The module defaults to latest-aws; pin a released tag for production.
You don't fetch these by hand: terraform init pulls the module from the Registry, the helm provider pulls the chart during terraform apply, and the EKS nodes pull the image at runtime. You set the chart registry and version (and pin the image tag) as module inputs on the installation page.
Some optional features require a minimum chart version (noted in the configuration reference).
7. Run the preflight check
Confirm your environment is ready. Each command below should succeed before you proceed to installation:
# Terraform is installed and >= 1.3
terraform version
# AWS credentials resolve to the target account
aws sts get-caller-identity
# kubectl is installed (server check comes after the cluster exists)
kubectl version --client
# If you set hosted_zone_id: the hosted zone resolves
aws route53 get-hosted-zone --id <hosted_zone_id>
# Optional: confirm you can reach the public chart on Docker Hub (no login required)
helm pull oci://registry-1.docker.io/montecarlodata/ao-data-platform --version 4.0.0If aws sts get-caller-identity returns the wrong account, or terraform version reports below 1.3, fix those before continuing.
Next steps
Continue to Installation.
Updated 6 days ago
