Presto

📘

Prerequisites

To complete this guide, you will need permissions to create a read-only user on Presto.

Before configuring your Presto query engine, make sure your metastore (e.g., Glue or Hive Metastore) is connected, as Monte Carlo requires it to map query engine activity to data lake assets.

To connect Monte Carlo to a Presto cluster to run data health SQL queries, follow these steps:

  1. Create a Presto cluster for Monte Carlo's data health queries using AWS EMR. Alternatively, you may use an existing cluster that you already have in your environment.
  2. Ensure that Monte Carlo's data collector has network connectivity to the cluster (VPC peering is required in most cases).
  3. Create a read-only service account on Presto if necessary.
  4. Provide service account credentials to Monte Carlo.

Creating a service account on Presto

Monte Carlo supports basic authentication or custom certificates for Presto connections. Please create a read-only user for Monte Carlo, or obtain a certificate file that enables authentication to Presto.

Providing account credentials to Monte Carlo

You will provide connection details for Presto using Monte Carlo's CLI:

  1. Please follow this guide to install and configure the CLI.
  2. Please use the command [montecarlo integrations add-presto](https://clidocs.getmontecarlo.com/#montecarlo-integrations-add-presto) to set up Presto connectivity. For reference, see help for this command below:
$ montecarlo integrations add-presto --help
Usage: montecarlo integrations add-presto [OPTIONS]

  Setup a Presto SQL integration. For health queries.

Options:
  --host TEXT                 Hostname.  [required]
  --port INTEGER              HTTP port.  [default: 8889]
  --user TEXT                 Username with access to catalog/schema.
  --password TEXT             User\'s password. If you prefer a prompt (with
                              hidden input) enter -1
  --catalog TEXT              Mount point to access data source.
  --metadata-catalog-id TEXT  Name of the catalog in the metadata store (e.g.
                              the AWS Glue federated catalog name, like
                              s3tablescatalog/<table-bucket>) when it differs
                              from the Trino mount name given in --catalog.
                              This option requires setting 'catalog'.
  --schema TEXT               Schema to access.
  --http-scheme [http|https]  Scheme for authentication.  [required]
  --cert-file FILE            Local SSL certificate file to upload to
                              collector. This option cannot be used with
                              'cert-s3'.
  --aws-profile TEXT          AWS profile to be used when uploading cert file.
                              This option requires setting 'cert-file'.
  --aws-region TEXT           AWS region to be used when uploading cert file.
                              This option requires setting 'cert-file'.
  --cert-s3 TEXT              Object path (key) to a certificate already
                              uploaded to the collector. This option cannot be
                              used with 'cert-file'.
  --skip-cert-verification    Skip SSL certificate verification.
  --name TEXT                 Friendly name of the warehouse which the
                              connection will belong to.
  --connection-name TEXT      Friendly name for the connection.
  --agent-id UUID             ID for the agent. To disambiguate accounts with
                              multiple agents. This option cannot be used with
                              'dc-id'.
  --collector-id UUID         ID for the data collector. To disambiguate
                              accounts with multiple collectors. This option
                              cannot be used with 'agent-id'.
  --skip-validation           Skip all connection tests. This option cannot be
                              used with 'validate-only'.
  --validate-only             Run connection tests without adding. This option
                              cannot be used with 'skip-validation'.
  --auto-yes                  Skip any interactive approval.
  --option-file FILE          Read configuration from FILE.
  --help                      Show this message and exit.

Using a Trino catalog name that differs from the metastore catalog name

If your Trino cluster mounts a catalog under a different name than your metastore uses (for example, AWS Glue calls the catalog awsdatacatalog but Trino mounts it as testcatalog), pass both names when adding the connection:

montecarlo integrations add-presto --host <host> --http-scheme https --catalog testcatalog --metadata-catalog-id awsdatacatalog
  • --catalog is the Trino mount name. Monte Carlo uses it to qualify table names in the queries it generates (for example, "testcatalog".schema.table).
  • --metadata-catalog-id is the catalog name as it appears in your metastore (for example, a Glue federated catalog such as s3tablescatalog/<table-bucket>). Queries on tables from this metastore catalog are qualified with the matching --catalog name.

Notes:

  • --metadata-catalog-id requires --catalog.
  • When --metadata-catalog-id is set, --catalog may only contain lowercase letters, digits, underscores and dashes.
  • Each metastore catalog can be mapped by only one active Presto connection per warehouse. To map several metastore catalogs, add one connection per catalog.

Did this page help you?