Read from Amazon S3 Using Environment Credentials

This example shows how to read a CSV file from Amazon S3 without putting the access keys in your code. The keys are resolved from environment variables by an EnvironmentCredentialsResolver each time the file system opens, so credentials rotated by your container platform, CI runner or AWS STS are picked up on the next connection without a restart or a redeploy.

Set AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY and, for temporary credentials, AWS_SESSION_TOKEN in the environment before running the job. Variables that are not set are left out of the resolved credentials, so the session token is optional. See Read from Amazon S3 Using a Rotating Credentials File for the same job reading its keys from a mounted file.

Java Code Listing

package com.northconcepts.datapipeline.examples.security;

import java.io.InputStreamReader;

import com.northconcepts.datapipeline.amazons3.AmazonS3FileSystem;
import com.northconcepts.datapipeline.core.DataReader;
import com.northconcepts.datapipeline.core.DataWriter;
import com.northconcepts.datapipeline.core.StreamWriter;
import com.northconcepts.datapipeline.csv.CSVReader;
import com.northconcepts.datapipeline.job.Job;
import com.northconcepts.datapipeline.security.Credentials;
import com.northconcepts.datapipeline.security.EnvironmentCredentialsResolver;

public class ReadFromAmazonS3UsingEnvironmentCredentials {

    private static final String BUCKET = "YOUR BUCKET";
    private static final String KEY = "output/trades.csv";

    public static void main(String[] args) throws Throwable {
        AmazonS3FileSystem s3 = new AmazonS3FileSystem()
                .setCredentialsResolver(new EnvironmentCredentialsResolver()
                        .map(Credentials.ACCESS_KEY, "AWS_ACCESS_KEY_ID")
                        .map(Credentials.SECRET_KEY, "AWS_SECRET_ACCESS_KEY")
                        .map(Credentials.SESSION_TOKEN, "AWS_SESSION_TOKEN"));
        s3.open();
        try {
            DataReader reader = new CSVReader(new InputStreamReader(s3.readFile(BUCKET, KEY)))
                    .setFieldNamesInFirstRow(true);
            DataWriter writer = StreamWriter.newSystemOutWriter();

            Job.run(reader, writer);
        } finally {
            s3.close();
        }
    }

}

Code Walkthrough

  1. BUCKET and KEY name the bucket and the object key (output/trades.csv) to read. Replace YOUR BUCKET with your own bucket name.
  2. An AmazonS3FileSystem is created and given an EnvironmentCredentialsResolver through setCredentialsResolver(). Each map() call pairs a credential key from Credentials (ACCESS_KEY, SECRET_KEY, SESSION_TOKEN) with the environment variable that supplies it.
  3. s3.open() asks the resolver for the credentials and connects to S3. A CredentialsResolver is consulted on every open, never when it is constructed, which is what makes rotation work.
  4. s3.readFile(BUCKET, KEY) returns an InputStream for the object, which is wrapped in a CSVReader that takes its field names from the first row.
  5. Job.run() transfers the records to a StreamWriter that prints them to the console.
  6. s3.close() in the finally block disconnects and clears the credentials snapshot held by the file system.

Console Output

Each record in trades.csv is printed to the console, followed by the record count.

Mobile Analytics