Credentials Resolver
Connectors that reach remote systems such as Amazon S3, FTP, SFTP, Dropbox, Google Drive, and HDFS need secrets, and secrets rotate. The security package, new in DataPipeline 11.0, gives every connector one way to obtain them: a CredentialsResolver that is asked for Credentials each time a connection is opened, never when the connector is constructed. Rotated passwords, keys, and tokens are picked up on the next connection without rebuilding the pipeline; secrets stay out of your code and out of serialized pipeline definitions; and one resolver can serve every file system in your application.
The security package is part of the core DataPipeline jar. The file systems come from their own add-ons
(northconcepts-datapipeline-filesystems-amazons3, -ftp, -sftp, -dropbox, -googledrive, and -hadoop);
see the Integrations page for how add-ons are added to a build.
Credentials
Credentials is an immutable bundle of named secret values with an optional expiry time. The well-known keys
(USERNAME, PASSWORD, ACCESS_KEY, SECRET_KEY, SESSION_TOKEN, and TOKEN)
let resolvers and connectors agree on names without a type hierarchy, and any other key can be added for connectors of your own.
Build credentials with Credentials.of(...) from key-value pairs, or with the builder when you also know when they expire.
Read values back with get(key), which returns null when the key is absent, or require(key), which throws a
DataException naming the missing key. Values are held as char[] so that clear() can wipe them,
copy() makes an independent snapshot, and toString() prints key names only, so secret values never appear in logs or exception messages.
Built-in Resolvers
A resolver implements one method, getCredentials(), plus the optional refresh() and isRotating() hints. These ship with the core library:
| Resolver | Resolves from | Typical use |
|---|---|---|
| StaticCredentialsResolver | Fixed values given in code. | Development and tests. |
| EnvironmentCredentialsResolver | Environment variables, mapped key by key and read on every resolve. | Containers and CI, where the platform injects secrets. |
| SystemPropertyCredentialsResolver | JVM system properties (-D), read on every resolve. |
Values set on the command line or changed at runtime. |
| FileCredentialsResolver | A Java properties file, re-read whenever its last-modified time changes. | Mounted secret files that an orchestrator such as Kubernetes rotates in place. |
| SuppliedCredentialsResolver | A Supplier<Credentials> lambda. |
Bridging to Vault, AWS Secrets Manager, STS, KMS, or your own store in one line. |
| CachingCredentialsResolver | Another resolver, cached until the credentials expire, a time-to-live elapses, or refresh() is called. |
Wrapping any resolver that is expensive to call. |
A credentials file is an ordinary properties file whose keys are the credential names, so a file for Amazon S3 looks like this:
Connecting Resolvers to File Systems
Every remote file system has a setCredentialsResolver(...) method and a constructor that takes no credentials at all.
The resolver is consulted at the top of open(). Keys it does not supply fall back to whatever the constructor or the legacy setters provided,
so you can resolve only the password and keep a fixed user name. Without a resolver, each connector behaves exactly as it did before.
| File system | Credential keys used | Resolver-only constructor |
|---|---|---|
| AmazonS3FileSystem (also AmazonS3File, AmazonS3FileSource, and AmazonS3FileSink) | ACCESS_KEY and SECRET_KEY, plus SESSION_TOKEN for temporary credentials | new AmazonS3FileSystem() |
| FtpFileSystem | USERNAME and PASSWORD | new FtpFileSystem(host, port) |
| SftpFileSystem | USERNAME and PASSWORD | new SftpFileSystem(host, port) |
| DropBoxFileSystem | TOKEN, the access token | new DropBoxFileSystem() |
| GoogleDriveFileSystem | TOKEN, an OAuth access token; the interactive sign-in flow is skipped when one is resolved | new GoogleDriveFileSystem(appName) |
| HdfsFileSystem | USERNAME, the Hadoop simple-authentication user | new HdfsFileSystem(hdfsUri) |
Credential Rotation
Rotation needs no extra code in the common case. Every call to a file system's open() resolves credentials afresh: sources and sinks open the
file system when they start a transfer, and a file system you open yourself can be closed and reopened to pick up new values. Amazon S3 goes further.
The AWS SDK asks for credentials when it signs each request, and AmazonS3FileSystem re-resolves automatically once the current credentials
report that they expire within the next minute. That is why resolvers that hand out temporary STS or OAuth credentials should always set expiresOn.
Each open() keeps an independent copy of the resolved credentials and wipes that copy on close(). A resolver that returns the same
Credentials instance every time, such as StaticCredentialsResolver, is therefore safe to share across file systems and reopen cycles.
Caching and the Resolver Registry
getCredentials() can be called often (once per request signing in the S3 case), so a resolver that talks to a remote store should be
wrapped in a CachingCredentialsResolver. It returns the cached credentials until they report expiry (five minutes before expiresOn by default;
see setExpirySkewMilliseconds), until an optional time-to-live elapses, or until refresh() is called.
CredentialsResolverRegistry maps ids to resolvers so that pipeline definitions can refer to a resolver by name instead of carrying secrets;
resolvers and resolved credentials are never written by toRecord(), toXmlElement(), or JSON serialization. Register resolvers once at
startup, in the JVM-wide getSystemRegistry() or in a registry you own, then look them up wherever file systems are built.
Writing Your Own Resolver
SuppliedCredentialsResolver covers most secret stores, so implement CredentialsResolver directly only when the store has its own
lease or renewal lifecycle. Either way, follow the contract:
- Return credentials or throw a
DataException; never return null. - Set
expiresOn(in epoch milliseconds) whenever the store reports an expiry; it is what drives automatic re-resolution. - Keep
getCredentials()cheap or wrap the resolver inCachingCredentialsResolver; userefresh()to drop any internal cache andisRotating()to report that values can change. - Put store ids and key names in exceptions and logs, never the values.
- Do not
clear()the credentials you hand out; each connector wipes its own copy.
Credentials Resolver Examples
- Read from Amazon S3 Using Environment Credentials
- Read from Amazon S3 Using a Rotating Credentials File
- Read from SFTP Using a Credentials Resolver
- Read from FTP Using a Credentials Resolver
- Cache and Register a Credentials Resolver
See the security Javadocs for the complete API.
