Agent Observability Platform: Connect to Monte Carlo (AWS)

Deploy the Monte Carlo Agent and connect the platform to Monte Carlo

โ˜๏ธ

This page covers AWS. Also available for Azure and GCP.

After installation, verify the deployment, deploy the Monte Carlo Agent, and create the ClickHouse integration in Monte Carlo. You carry out these steps yourself โ€” work through them in order.

1. Configure cluster access

If you ran terraform apply from this same machine, you can skip this step โ€” the module already ran aws eks update-kubeconfig during apply, so the cluster's context is in your ~/.kube/config and is current. Run the command below only when connecting from a different machine:

aws eks update-kubeconfig --name <eks_cluster_name> --region <region>
๐Ÿ“˜

This requires eks:DescribeCluster on the cluster, plus access to authenticate to it. The principal that ran terraform apply is granted cluster-administrator access automatically (enable_cluster_creator_admin_permissions = true), so the same credentials you deployed with can run the verification below. To let a different principal run kubectl, add an EKS access entry for it.

2. Verify the deployment

All components run in the montecarlo namespace. Confirm each is healthy:

ComponentCommandHealthy when
ClickHouse operatorkubectl get pods -n montecarlo -l app.kubernetes.io/name=altinity-clickhouse-operatorPod Running
ClickHouse instancekubectl get chi -n montecarloResource otel, status Completed
ClickHouse podskubectl get pods -n montecarlo -l clickhouse.altinity.com/chi=otelOne pod per replica, all Running
Keeper ensemblekubectl get chk -n montecarloResource otel, status Completed
Keeper podskubectl get pods -n montecarlo -l clickhouse-keeper.altinity.com/chk=otelOne pod per voter, all Running
Schema migration jobkubectl get jobs -n montecarloclickhouse-schema-<n> shows COMPLETIONS 1/1
OpenTelemetry Collectorkubectl get pods -n montecarlo -l app.kubernetes.io/name=opentelemetry-collectorPod Running
LLM workerkubectl get pods -n montecarlo -l app.kubernetes.io/component=llm-workerPod Running
TLS certificateskubectl get certificates -n montecarloclickhouse-server-tls and otel-collector-tls show READY True
External secretskubectl get externalsecret -n montecarloEach ao-clickhouse-*-credentials secret shows status SecretSynced

The schema migration job must reach 1/1 before the Collector and LLM worker can write data โ€” both wait on it via an init container, so it is normal for them to sit in Init for a minute or two on a fresh install. <n> in the job name is the Helm release revision (e.g. clickhouse-schema-1).

๐Ÿ“˜

A ClickHouse pod showing Running and Ready confirms it can serve writes โ€” it does not confirm replication has caught up. To verify replication health on the HA topology, check system.replicas โ€” see Verifying cluster health.

๐Ÿ“˜

If any check above doesn't reach its healthy state, see Troubleshooting & FAQ โ€” it covers ClickHouse pods stuck Pending, certificates not becoming Ready, external secrets not syncing, and schema-job failures.

3. Deploy the Monte Carlo Agent

All SQL from the Monte Carlo platform โ€” for both Trace Exploration and agent monitors โ€” flows through the Monte Carlo Agent, which queries ClickHouse over its internal endpoint as the monte_carlo user. The Agent is not deployed by the terraform-aws-ao-data-platform module; you deploy it separately, in one of two forms:

OptionAgentNetwork modelIntegration credentials
ALambda function in the platform's VPCMonte Carlo invokes the AgentEntered in Monte Carlo, or self-hosted
BGeneric Agent pod in the platform's EKS clusterThe Agent connects out to Monte Carlo (egress-only)Self-hosted
๐Ÿ“˜

You deploy and register the Agent yourself โ€” it is a separate deployment from the data platform, not something Monte Carlo runs in your account. Either way, the Agent must be able to reach the ClickHouse endpoint, which is on an internal load balancer: deploy it in the platform's VPC (or a network-attached one).

Option A โ€” Lambda agent

On AWS, the Lambda agent runs as a Lambda function in the same VPC as the EKS cluster. Deploy it with the terraform-aws-mcd-agent module. The Agent Observability-specific configuration is the ClickHouse connection: the endpoint (your clickhouse_domain) and the credentials retrieved below. Make sure the Agent's Lambda is in (or can route to) the platform's VPC. Your Monte Carlo representative can help you obtain the Agent's registration values.

Option B โ€” Generic (egress) agent

Instead of the Lambda agent, you can run Monte Carlo's Generic Agent as a pod inside the EKS cluster that hosts the platform. All connections are initiated from the Agent to Monte Carlo, so no inbound network path from Monte Carlo to your account is required.

Register the Agent

In Monte Carlo, go to Settings โ†’ Deployments โ†’ Add โ†’ Generic, choose Key/Token authentication, and click Provision. Copy the key ID and secret โ€” they are shown only once. The backend_service_url is on the Account information page, in the agent-service section.

Deploy the Agent into the platform cluster

Use the monte-carlo-data/mcd-k8s-agent/aws module (version 0.1.10 or later) in its own Terraform root, applied after the platform, pointing at the cluster and VPC the platform uses:

locals {
  platform_cluster_name = "<platform EKS cluster name>" # the platform's eks_cluster_name output
  platform_region       = "<platform region>"
}

data "aws_caller_identity" "current" {}

data "aws_eks_cluster" "platform" {
  name = local.platform_cluster_name
}

data "aws_iam_openid_connect_provider" "platform" {
  url = data.aws_eks_cluster.platform.identity[0].oidc[0].issuer
}

# The External Secrets Operator role the platform module created. It exists when the
# module installed ESO (helm.install_external_secrets_operator = true, the default);
# if you installed ESO yourself, look up that operator's IRSA role instead.
data "aws_iam_role" "platform_eso" {
  name = "${local.platform_cluster_name}-${local.platform_region}-external-secrets"
}

# Your role for the Agent โ€” see "Agent identity" below for its trust and permissions.
resource "aws_iam_role" "mcd_agent" {
  name               = "${local.platform_cluster_name}-mcd-agent"
  assume_role_policy = data.aws_iam_policy_document.mcd_agent_trust.json
}

module "mcd_agent" {
  source  = "monte-carlo-data/mcd-k8s-agent/aws"
  version = "~> 0.1.10"

  backend_service_url = "<backend_service_url>"

  token_credentials = {
    mcd_id    = var.mcd_id
    mcd_token = var.mcd_token
  }

  cluster = {
    create                = false
    existing_cluster_name = local.platform_cluster_name
  }

  identity = {
    mode                    = "irsa"
    existing_eso_role_arn   = data.aws_iam_role.platform_eso.arn
    create_agent_role       = false
    existing_agent_role_arn = aws_iam_role.mcd_agent.arn
  }

  storage = {
    create_bucket        = false
    existing_bucket_name = "<your agent bucket>"
  }

  networking = {
    create_vpc                  = false
    existing_vpc_id             = data.aws_eks_cluster.platform.vpc_config[0].vpc_id
    existing_private_subnet_ids = ["<platform private subnet IDs>"]
    create_vpc_endpoints        = false # the platform VPC has a NAT gateway
  }

  helm = {
    chart_version                     = "<generic-agent-helm chart version>"
    install_external_secrets_operator = false # the platform already runs ESO
  }
}

The Agent reaches the ClickHouse internal load balancer through the platform's private subnets; no peering or public endpoint is needed. After terraform apply, return to the registration screen in Monte Carlo and click Enable.

Agent identity

When the Agent shares the platform's cluster, use IRSA (identity.mode = "irsa"). The platform and its controllers already use IRSA through the cluster's OIDC provider, so no additional cluster add-on is needed, and the whole deployment stays on a single identity mechanism. Pass the platform's ESO role as existing_eso_role_arn so the Agent's token syncs through the operator the platform already runs.

Supplying the Agent's role yourself is recommended โ€” it lives in your Terraform and is reviewed like the rest of your IAM. Its trust policy must allow the cluster's OIDC provider for the Agent's service account:

data "aws_iam_policy_document" "mcd_agent_trust" {
  statement {
    effect  = "Allow"
    actions = ["sts:AssumeRoleWithWebIdentity"]
    principals {
      type        = "Federated"
      identifiers = [data.aws_iam_openid_connect_provider.platform.arn]
    }
    condition {
      test     = "StringEquals"
      variable = "${trimprefix(data.aws_eks_cluster.platform.identity[0].oidc[0].issuer, "https://")}:sub"
      values   = ["system:serviceaccount:mcd-agent:mcd-agent-service-account"]
    }
    condition {
      test     = "StringEquals"
      variable = "${trimprefix(data.aws_eks_cluster.platform.identity[0].oidc[0].issuer, "https://")}:aud"
      values   = ["sts.amazonaws.com"]
    }
  }
}

Attach the S3 permissions from the object storage page for your bucket, plus secretsmanager:GetSecretValue on the Agent's token secret (mcd/agent/token by default) and on the integration secret you create in step 5.

โš ๏ธ
  • create_agent_role = false is required when the role is created in the same configuration: its ARN is unknown until apply, and without the flag the plan fails with Invalid count argument.
  • A role you supply requires storage.existing_bucket_name โ€” your role must be scoped to a bucket whose name is known in advance.
  • Do not create EKS Pod Identity associations for the Agent's service accounts on this cluster. An association's injected credentials take precedence over the IRSA annotation and silently revert the Agent to Pod Identity.

4. Retrieve credentials and endpoints

Monte Carlo connects to ClickHouse as the monte_carlo user โ€” a least-privilege identity that reads telemetry and appends only to the evaluation job queue. Its password is stored in AWS Secrets Manager; the module exposes the secret's ARN as the clickhouse_monte_carlo_credentials_secret_arn output. Retrieve it:

aws secretsmanager get-secret-value \
  --secret-id "$(terraform output -raw clickhouse_monte_carlo_credentials_secret_arn)" \
  --query SecretString --output text

You enter the following when you create the ClickHouse integration in the next step:

Value to provideSource
ClickHouse endpointThe clickhouse_domain you configured (e.g. clickhouse.acme.com)
ClickHouse port8443 (HTTPS)
ClickHouse databaseotel_traces
ClickHouse usernamemonte_carlo
ClickHouse passwordThe value retrieved above from clickhouse_monte_carlo_credentials_secret_arn
๐Ÿ”’

Treat these credentials as secrets โ€” avoid printing them into shared logs or terminals, and if you need to share them (for example with a teammate or your Monte Carlo representative), use a secure channel. monte_carlo is a least-privilege user: it can read telemetry and append to the evaluation job queue, but cannot write telemetry, run DDL, or manage access.

5. Create the ClickHouse integration in Monte Carlo

With the Agent deployed (step 3) and the connection values in hand (step 4), create the ClickHouse integration in Monte Carlo. Monte Carlo reaches ClickHouse through the Agent you deployed, so this is the step that ties the platform to your account. Creating the integration runs a connection check โ€” when it succeeds, you have confirmation that the endpoint, port, credentials, and database are all wired correctly.

๐Ÿ“˜

Monte Carlo's general ClickHouse integration guide covers connecting Monte Carlo to a ClickHouse warehouse. It does not strictly apply to this platform: you do not need to create a user for the connection โ€” the purpose-built monte_carlo user is already provisioned with exactly the permissions the Agent needs.

Option B: Generic agent (customer-managed credentials)

Integrations that use the Generic agent use self-hosted credentials: the Agent reads the connection details from storage in your environment, not from values entered in Monte Carlo. The example below uses AWS Secrets Manager; the linked page covers the other options, such as environment variables. Store the values from step 4 as a secret โ€” for example mcd/integrations/clickhouse:

{
  "connect_args": {
    "host": "<your clickhouse_domain>",
    "port": 8443,
    "username": "monte_carlo",
    "password": "<the monte_carlo password>",
    "database": "otel_traces"
  }
}

The Agent's role needs secretsmanager:GetSecretValue on it (see Agent identity).

Then register the integration: Settings โ†’ Integrations โ†’ ClickHouse โ†’ Customer managed credentials, select your Generic agent deployment, choose AWS Secrets Manager, and enter the secret. Or use the CLI โ€” the integration is created disabled; run the connection test, then enable it:

montecarlo integrations add-self-hosted-credentials-v2 \
  --connection-type clickhouse \
  --self-hosted-credentials-type AWS_SECRETS_MANAGER \
  --aws-secret <secret ARN or name> \
  --agent-id <your Generic agent ID>

--agent-id selects the Generic agent when your account has more than one.

โš ๏ธ

Declare the connection for Agent Observability. With customer-managed credentials, Monte Carlo can't tell from the connection alone that it is an Agent Observability trace store. On the add flow (or the connection's edit page), check Use this connection for Agent Observability โ€” without it, the connection never enters the Agent Observability read path and no agents appear.

6. Wait for your agents to appear

The Agents page is populated by a periodic sweep that runs about every 30 minutes, and it only discovers agents that already have spans in ClickHouse. Send traces from at least one instrumented agent first; it appears after the next sweep โ€” refreshing the page before then won't show it.

Next steps


Did this page help you?