Agent Observability Platform: Connect to Monte Carlo (AWS)
Deploy the Monte Carlo Agent and connect the platform to Monte Carlo
After installation, verify the deployment, deploy the Monte Carlo Agent, and create the ClickHouse integration in Monte Carlo. You carry out these steps yourself โ work through them in order.
1. Configure cluster access
If you ran terraform apply from this same machine, you can skip this step โ the module already ran aws eks update-kubeconfig during apply, so the cluster's context is in your ~/.kube/config and is current. Run the command below only when connecting from a different machine:
aws eks update-kubeconfig --name <eks_cluster_name> --region <region>This requires
eks:DescribeClusteron the cluster, plus access to authenticate to it. The principal that ranterraform applyis granted cluster-administrator access automatically (enable_cluster_creator_admin_permissions = true), so the same credentials you deployed with can run the verification below. To let a different principal runkubectl, add an EKS access entry for it.
2. Verify the deployment
All components run in the montecarlo namespace. Confirm each is healthy:
| Component | Command | Healthy when |
|---|---|---|
| ClickHouse operator | kubectl get pods -n montecarlo -l app.kubernetes.io/name=altinity-clickhouse-operator | Pod Running |
| ClickHouse instance | kubectl get chi -n montecarlo | Resource otel, status Completed |
| ClickHouse pods | kubectl get pods -n montecarlo -l clickhouse.altinity.com/chi=otel | One pod per replica, all Running |
| Keeper ensemble | kubectl get chk -n montecarlo | Resource otel, status Completed |
| Keeper pods | kubectl get pods -n montecarlo -l clickhouse-keeper.altinity.com/chk=otel | One pod per voter, all Running |
| Schema migration job | kubectl get jobs -n montecarlo | clickhouse-schema-<n> shows COMPLETIONS 1/1 |
| OpenTelemetry Collector | kubectl get pods -n montecarlo -l app.kubernetes.io/name=opentelemetry-collector | Pod Running |
| LLM worker | kubectl get pods -n montecarlo -l app.kubernetes.io/component=llm-worker | Pod Running |
| TLS certificates | kubectl get certificates -n montecarlo | clickhouse-server-tls and otel-collector-tls show READY True |
| External secrets | kubectl get externalsecret -n montecarlo | Each ao-clickhouse-*-credentials secret shows status SecretSynced |
The schema migration job must reach 1/1 before the Collector and LLM worker can write data โ both wait on it via an init container, so it is normal for them to sit in Init for a minute or two on a fresh install. <n> in the job name is the Helm release revision (e.g. clickhouse-schema-1).
A ClickHouse pod showing
RunningandReadyconfirms it can serve writes โ it does not confirm replication has caught up. To verify replication health on the HA topology, checksystem.replicasโ see Verifying cluster health.
If any check above doesn't reach its healthy state, see Troubleshooting & FAQ โ it covers ClickHouse pods stuck
Pending, certificates not becomingReady, external secrets not syncing, and schema-job failures.
3. Deploy the Monte Carlo Agent
All SQL from the Monte Carlo platform โ for both Trace Exploration and agent monitors โ flows through the Monte Carlo Agent, which queries ClickHouse over its internal endpoint as the monte_carlo user. The Agent is not deployed by the terraform-aws-ao-data-platform module; you deploy it separately, in one of two forms:
| Option | Agent | Network model | Integration credentials |
|---|---|---|---|
| A | Lambda function in the platform's VPC | Monte Carlo invokes the Agent | Entered in Monte Carlo, or self-hosted |
| B | Generic Agent pod in the platform's EKS cluster | The Agent connects out to Monte Carlo (egress-only) | Self-hosted |
You deploy and register the Agent yourself โ it is a separate deployment from the data platform, not something Monte Carlo runs in your account. Either way, the Agent must be able to reach the ClickHouse endpoint, which is on an internal load balancer: deploy it in the platform's VPC (or a network-attached one).
Option A โ Lambda agent
On AWS, the Lambda agent runs as a Lambda function in the same VPC as the EKS cluster. Deploy it with the terraform-aws-mcd-agent module. The Agent Observability-specific configuration is the ClickHouse connection: the endpoint (your clickhouse_domain) and the credentials retrieved below. Make sure the Agent's Lambda is in (or can route to) the platform's VPC. Your Monte Carlo representative can help you obtain the Agent's registration values.
Option B โ Generic (egress) agent
Instead of the Lambda agent, you can run Monte Carlo's Generic Agent as a pod inside the EKS cluster that hosts the platform. All connections are initiated from the Agent to Monte Carlo, so no inbound network path from Monte Carlo to your account is required.
Register the Agent
In Monte Carlo, go to Settings โ Deployments โ Add โ Generic, choose Key/Token authentication, and click Provision. Copy the key ID and secret โ they are shown only once. The backend_service_url is on the Account information page, in the agent-service section.
Deploy the Agent into the platform cluster
Use the monte-carlo-data/mcd-k8s-agent/aws module (version 0.1.10 or later) in its own Terraform root, applied after the platform, pointing at the cluster and VPC the platform uses:
locals {
platform_cluster_name = "<platform EKS cluster name>" # the platform's eks_cluster_name output
platform_region = "<platform region>"
}
data "aws_caller_identity" "current" {}
data "aws_eks_cluster" "platform" {
name = local.platform_cluster_name
}
data "aws_iam_openid_connect_provider" "platform" {
url = data.aws_eks_cluster.platform.identity[0].oidc[0].issuer
}
# The External Secrets Operator role the platform module created. It exists when the
# module installed ESO (helm.install_external_secrets_operator = true, the default);
# if you installed ESO yourself, look up that operator's IRSA role instead.
data "aws_iam_role" "platform_eso" {
name = "${local.platform_cluster_name}-${local.platform_region}-external-secrets"
}
# Your role for the Agent โ see "Agent identity" below for its trust and permissions.
resource "aws_iam_role" "mcd_agent" {
name = "${local.platform_cluster_name}-mcd-agent"
assume_role_policy = data.aws_iam_policy_document.mcd_agent_trust.json
}
module "mcd_agent" {
source = "monte-carlo-data/mcd-k8s-agent/aws"
version = "~> 0.1.10"
backend_service_url = "<backend_service_url>"
token_credentials = {
mcd_id = var.mcd_id
mcd_token = var.mcd_token
}
cluster = {
create = false
existing_cluster_name = local.platform_cluster_name
}
identity = {
mode = "irsa"
existing_eso_role_arn = data.aws_iam_role.platform_eso.arn
create_agent_role = false
existing_agent_role_arn = aws_iam_role.mcd_agent.arn
}
storage = {
create_bucket = false
existing_bucket_name = "<your agent bucket>"
}
networking = {
create_vpc = false
existing_vpc_id = data.aws_eks_cluster.platform.vpc_config[0].vpc_id
existing_private_subnet_ids = ["<platform private subnet IDs>"]
create_vpc_endpoints = false # the platform VPC has a NAT gateway
}
helm = {
chart_version = "<generic-agent-helm chart version>"
install_external_secrets_operator = false # the platform already runs ESO
}
}The Agent reaches the ClickHouse internal load balancer through the platform's private subnets; no peering or public endpoint is needed. After terraform apply, return to the registration screen in Monte Carlo and click Enable.
Agent identity
When the Agent shares the platform's cluster, use IRSA (identity.mode = "irsa"). The platform and its controllers already use IRSA through the cluster's OIDC provider, so no additional cluster add-on is needed, and the whole deployment stays on a single identity mechanism. Pass the platform's ESO role as existing_eso_role_arn so the Agent's token syncs through the operator the platform already runs.
Supplying the Agent's role yourself is recommended โ it lives in your Terraform and is reviewed like the rest of your IAM. Its trust policy must allow the cluster's OIDC provider for the Agent's service account:
data "aws_iam_policy_document" "mcd_agent_trust" {
statement {
effect = "Allow"
actions = ["sts:AssumeRoleWithWebIdentity"]
principals {
type = "Federated"
identifiers = [data.aws_iam_openid_connect_provider.platform.arn]
}
condition {
test = "StringEquals"
variable = "${trimprefix(data.aws_eks_cluster.platform.identity[0].oidc[0].issuer, "https://")}:sub"
values = ["system:serviceaccount:mcd-agent:mcd-agent-service-account"]
}
condition {
test = "StringEquals"
variable = "${trimprefix(data.aws_eks_cluster.platform.identity[0].oidc[0].issuer, "https://")}:aud"
values = ["sts.amazonaws.com"]
}
}
}Attach the S3 permissions from the object storage page for your bucket, plus secretsmanager:GetSecretValue on the Agent's token secret (mcd/agent/token by default) and on the integration secret you create in step 5.
create_agent_role = falseis required when the role is created in the same configuration: its ARN is unknown until apply, and without the flag the plan fails with Invalid count argument.- A role you supply requires
storage.existing_bucket_nameโ your role must be scoped to a bucket whose name is known in advance.- Do not create EKS Pod Identity associations for the Agent's service accounts on this cluster. An association's injected credentials take precedence over the IRSA annotation and silently revert the Agent to Pod Identity.
4. Retrieve credentials and endpoints
Monte Carlo connects to ClickHouse as the monte_carlo user โ a least-privilege identity that reads telemetry and appends only to the evaluation job queue. Its password is stored in AWS Secrets Manager; the module exposes the secret's ARN as the clickhouse_monte_carlo_credentials_secret_arn output. Retrieve it:
aws secretsmanager get-secret-value \
--secret-id "$(terraform output -raw clickhouse_monte_carlo_credentials_secret_arn)" \
--query SecretString --output textYou enter the following when you create the ClickHouse integration in the next step:
| Value to provide | Source |
|---|---|
| ClickHouse endpoint | The clickhouse_domain you configured (e.g. clickhouse.acme.com) |
| ClickHouse port | 8443 (HTTPS) |
| ClickHouse database | otel_traces |
| ClickHouse username | monte_carlo |
| ClickHouse password | The value retrieved above from clickhouse_monte_carlo_credentials_secret_arn |
Treat these credentials as secrets โ avoid printing them into shared logs or terminals, and if you need to share them (for example with a teammate or your Monte Carlo representative), use a secure channel.
monte_carlois a least-privilege user: it can read telemetry and append to the evaluation job queue, but cannot write telemetry, run DDL, or manage access.
5. Create the ClickHouse integration in Monte Carlo
With the Agent deployed (step 3) and the connection values in hand (step 4), create the ClickHouse integration in Monte Carlo. Monte Carlo reaches ClickHouse through the Agent you deployed, so this is the step that ties the platform to your account. Creating the integration runs a connection check โ when it succeeds, you have confirmation that the endpoint, port, credentials, and database are all wired correctly.
Monte Carlo's general ClickHouse integration guide covers connecting Monte Carlo to a ClickHouse warehouse. It does not strictly apply to this platform: you do not need to create a user for the connection โ the purpose-built
monte_carlouser is already provisioned with exactly the permissions the Agent needs.
Option B: Generic agent (customer-managed credentials)
Integrations that use the Generic agent use self-hosted credentials: the Agent reads the connection details from storage in your environment, not from values entered in Monte Carlo. The example below uses AWS Secrets Manager; the linked page covers the other options, such as environment variables. Store the values from step 4 as a secret โ for example mcd/integrations/clickhouse:
{
"connect_args": {
"host": "<your clickhouse_domain>",
"port": 8443,
"username": "monte_carlo",
"password": "<the monte_carlo password>",
"database": "otel_traces"
}
}The Agent's role needs secretsmanager:GetSecretValue on it (see Agent identity).
Then register the integration: Settings โ Integrations โ ClickHouse โ Customer managed credentials, select your Generic agent deployment, choose AWS Secrets Manager, and enter the secret. Or use the CLI โ the integration is created disabled; run the connection test, then enable it:
montecarlo integrations add-self-hosted-credentials-v2 \
--connection-type clickhouse \
--self-hosted-credentials-type AWS_SECRETS_MANAGER \
--aws-secret <secret ARN or name> \
--agent-id <your Generic agent ID>--agent-id selects the Generic agent when your account has more than one.
Declare the connection for Agent Observability. With customer-managed credentials, Monte Carlo can't tell from the connection alone that it is an Agent Observability trace store. On the add flow (or the connection's edit page), check Use this connection for Agent Observability โ without it, the connection never enters the Agent Observability read path and no agents appear.
6. Wait for your agents to appear
The Agents page is populated by a periodic sweep that runs about every 30 minutes, and it only discovers agents that already have spans in ClickHouse. Send traces from at least one instrumented agent first; it appears after the next sweep โ refreshing the page before then won't show it.
Next steps
- Tune retention, users, TLS, and sizing in the Configuration reference.
- Hit a problem during verification or onboarding? See Troubleshooting & FAQ.
Updated 6 days ago
