Convalesce Handbook
Data lakes and object storage

S3

Connect an S3 bucket step by step: the path spec that turns files into datasets, the read-only user, and what is read.

Connect it

  1. Name it: what to call this connection, and the deployment it belongs to.
  2. Map files to datasets: which files make up which dataset.
  3. Create the read-only user: a read-only IAM user in your account, and its key.
  4. Add the access key: an access key for the reading identity.
  5. Test the connection: check Convalesce can reach it with what you entered.
  6. Choose how often: how often Convalesce reads it.
  7. Review and connect: check everything, then save the connection.

Convalesce reads the files in your S3 bucket and turns them into datasets. You tell it which files make up which dataset with a path spec, and it reads through one read-only IAM user that you create.

The connect screen writes that user's policy for the bucket your path spec names.

Convalesce is a hosted service, so it connects to AWS or your storage over the internet. Nothing is installed on your side.

Before you start

Have these ready and the rest takes a few minutes:

  • The bucket and the layout of the files in it, to write the path spec.
  • The region the bucket is in.
  • The AWS CLI signed in as an IAM admin, to run one block of commands. Or an existing AWS connection in Convalesce whose key you want to share.

Connect it

In Convalesce, open Integrations, choose S3, and follow the steps. Each one is shown below as it looks on screen, with what it asks for and anything to copy and run.

Step 1 of 7: Name it

What to call this connection, and the deployment it belongs to.

The "Name it" step of the connect screen
What it asks forNeededWhat to enter
NameYesHow it is listed in Convalesce. Something that says which one it is, if there will be more than one. For example, Orders database.
DeploymentYesWhich environment this is. Choose the same one as the pipelines that write to it, so both name its tables alike. Choose one of: Production, Staging, Development, Test, Quality assurance, User acceptance, Pre-production, Sandbox.
Instance nameOptionalOnly when you connect two of these in the same deployment, such as two production servers: it keeps their tables apart. Leave it empty otherwise. For example, eu1.

Step 2 of 7: Map files to datasets

Which files make up which dataset.

The "Map files to datasets" step of the connect screen

S3 has no native concept of a "table". A path spec is what turns a bucket prefix into one: a glob pattern that groups one or more files into a single dataset.

s3://bucket/*.csv maps each matching file to its own dataset. s3://bucket/{table}/*.avro maps each folder under the bucket to one dataset named by the {table} placeholder, made of every file inside it.

Keep the fixed prefix as long as possible, and avoid /*/ wildcards you don't need. A vaguer pattern means more of the bucket gets listed, which costs time and, on S3, money.

What it asks forNeededWhat to enter
Path specYesIt starts with s3://, then the bucket. For example, s3://my-bucket/{table}/*.parquet.
AWS regionYesThe region the bucket is in. For example, us-east-1.
Patterns to skipOptionalAdd a glob for each set of files or folders under the path spec that should not be read, such as temporary or archived data. For example, s3://my-bucket/**/_tmp/**.
Which partition folders to readOptionalFor a dataset split into dated folders. Reading every folder lists the whole prefix, which takes longer and costs more on S3. Choose one of: The newest only, The oldest and the newest, Every one.

Step 3 of 7: Create the read-only user

A read-only IAM user in your account, and its key.

The "Create the read-only user" step of the connect screen

Make one IAM user for Convalesce in your own account with only the policy below, and issue it an access key. The commands do all three.

s3:ListBucket walks the bucket to find files matching your path specs, s3:GetBucketLocation resolves the bucket's region, and s3:GetObject reads object content, which schema inference and profiling both need.

There is no network rule to add: Convalesce calls AWS's own API, which is public. The exception is a bucket policy, a key policy or an IAM condition that limits source addresses (aws:SourceIp): it has to allow 34.66.85.47, the address Convalesce connects from.

Create the user, its policy and its key, with the AWS CLI
cat > convalesce-read.json <<'EOF'
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ConvalesceRead",
      "Effect": "Allow",
      "Action": ["s3:ListBucket", "s3:GetBucketLocation", "s3:GetObject"],
      "Resource": ["arn:aws:s3:::my-lake", "arn:aws:s3:::my-lake/*"]
    }
  ]
}
EOF
aws iam get-user --user-name convalesce-reader >/dev/null 2>&1 || aws iam create-user --user-name convalesce-reader
aws iam put-user-policy --user-name convalesce-reader --policy-name ConvalesceS3Read \
  --policy-document file://convalesce-read.json
aws iam create-access-key --user-name convalesce-reader

Run it as an IAM admin in your own account, where the AWS CLI is signed in (set AWS_PROFILE first if you use a named profile). It is safe to run again: the user is made once, and each tool's policy goes on under its own name. The last command prints AccessKeyId and SecretAccessKey once: paste them straight into the next step, nowhere else. AWS allows two keys per user, so delete an old one before issuing a third. In the IAM console instead, paste the JSON between the EOF lines as an inline policy.

Step 4 of 7: Add the access key

An access key for the reading identity.

The "Add the access key" step of the connect screen

An access key for an IAM user in your account that has only the policy above. Convalesce has no AWS identity of its own, so it reads only what that user is granted.

What it asks forNeededWhat to enter
Access key IDYes
Secret access keyYes Stored encrypted the moment you enter it, and shown to no one afterwards.
Session tokenOptionalOnly for temporary credentials. Stored encrypted the moment you enter it, and shown to no one afterwards.

Step 5 of 7: Test the connection

Check Convalesce can reach it with what you entered.

The "Test the connection" step of the connect screen

The test runs on the same worker a real run would, with the recipe exactly as it will be saved, so it fails the way a run would.

Step 6 of 7: Choose how often

How often Convalesce reads it.

The "Choose how often" step of the connect screen

Step 7 of 7: Review and connect

Check everything, then save the connection.

The "Review and connect" step of the connect screen

Network

There is no network rule to add: Convalesce calls AWS's own API, which is public. The exception is a bucket policy, a key policy or an IAM condition that limits source addresses (aws:SourceIp): it has to allow 34.66.85.47, the address Convalesce connects from.

Writing a path spec

A path spec is a pattern that starts with s3:// and the bucket. s3://my-bucket/*.csv makes each matching file its own dataset. s3://my-bucket/{table}/*.parquet makes each folder one dataset, named by the folder, made of every file inside it.

Use * for one level and ** for any depth. Keep the fixed part at the front as long as you can: the longer it is, the less of the bucket is listed.

Patterns to skip takes one pattern for each set of files or folders to leave out, such as s3://my-bucket/**/_tmp/**.

Settings

What the connect screen asks for

InputOn the stepNeeded
NameName itYes
DeploymentName itYes
Instance nameName itOptional
Path specMap files to datasetsYes
AWS regionMap files to datasetsYes
Patterns to skipMap files to datasetsOptional
Which partition folders to readMap files to datasetsOptional
Access key IDAdd the access keyYes
Secret access keyAdd the access keyYes
Session tokenAdd the access keyOptional

Set for you

These are the same on every connection. The connect screen does not ask for them.

What it meansSetting
A path spec can use ** for any depth.path_specs.0.allow_double_stars: true
Row counts and column statistics are not read.profiling.enabled: false
Something that is no longer there is marked as removed.stateful_ingestion.enabled: true

Troubleshooting

  • AccessDenied. The policy has to be attached to the user whose key you entered, for the bucket in the path spec. Run the commands on Create the read-only user again.
  • No datasets appear. The path spec matches no files: check the bucket name, the folder levels, and the file extension.
  • Too many small datasets. Use {table} on the folder that should be one dataset, so its files are read together.

On this page