DataPipeline 11.0 Released

Welcome to the 11.0 release of DataPipeline.

This release moves DataPipeline to its new home at DataPipeline.io and raises the minimum Java version to JDK 21. It adds JSONata queries, LaTeX and Microsoft Word output, Open Financial Exchange (OFX) input, Google Sheets, rotating credentials for every file system, a pluggable record detection framework for XML and JSON, and an Amazon S3 connector rebuilt on the AWS SDK for Java 2.x. It also comes with 26 new examples and five new user guide chapters.

A New Home: DataPipeline.io

DataPipeline now lives at DataPipeline.io. The user guide, examples, Javadocs, changelog, and this blog have all moved here, and links to the old northconcepts.com documentation, examples, Javadocs, and blog posts redirect to their new addresses, so existing bookmarks keep working. DataPipeline is still built and supported by North Concepts Inc.; only its address has changed.

A few practical details for the move:

  • Maven repository: artifacts are served from https://maven.datapipeline.io/public/repositories/datapipeline. The previous maven.northconcepts.com address continues to work, so existing builds do not need to change.
  • No more zip downloads: the library comes from the Maven repository, the examples are on GitHub, and the Javadocs are on this site. The Getting Started page has the dependency snippets for every edition, and your license file is emailed to you after a quick sign-up.
  • Support: reach us at support@datapipeline.io.

JDK 21 Minimum

DataPipeline 11.0 requires Java 21 or later and is tested on JDK 21 through 26. DataPipeline 10.1 was the last release to run on Java 8; it stays available in the Maven repository if you are not ready to move yet. Raising the minimum lets DataPipeline use current Java APIs and keep its dependencies current.

Two other breaking changes in the core library come with the recompile:

  • The public static log fields (DataObject.log/DataEndpoint.log, RecordList.log, EventBus.log, DeMux.log, SplitWriter.log, and DataReaderServer.log) are now com.northconcepts.datapipeline.logger.Logger instead of org.apache.log4j.Logger. Code that declares them with the Log4j type must switch imports and recompile. Log4j itself remains a dependency.
  • RemoveDuplicateFields.DuplicateFieldsPolicy.RENAME now renames duplicate fields in place instead of moving them to the end of the record.

Core Changes

  1. Credentials resolution (new feature): the new com.northconcepts.datapipeline.security package resolves credentials on demand so rotated secrets are picked up on the next connection. A CredentialsResolver returns a Credentials object (username, password, access key, secret key, session token, and token, plus an optional expiry), with built-in resolvers for static values, environment variables, system properties, properties files, suppliers, and caching, and a CredentialsResolverRegistry to share them by name. FileSystem.setCredentialsResolver() wires a resolver into any file system (see the file system changes below). Read the Credentials Resolver chapter in the user guide.
  2. TextTableWriter (new feature): TextTableWriter writes records as an easy-to-read text table with padded, aligned columns (columnSeparator and newLine properties). See Write a CSV File as a Text Table.
  3. ValueNode conversions: ValueNode.from(Object value) wraps any value, and the new asSingleValue(), asRecord(), asArrayValue(), toJson(), and toXml() methods convert between node types and serialize them. See Convert Values and Records to JSON and XML.
  4. Duplicate field renaming: RemoveDuplicateFields adds DuplicateFieldsPolicy.RENAME_UNDERSCORE and RenamePolicy(String separator) to control the separator used when renaming duplicate fields. See Rename Duplicate Fields with a Separator.
  5. Query ordering: added Select.orderBy(FieldList fields) and QueryOrder.add(FieldList fields), so a query can be sorted by the fields of a record. Read the SQL Generators chapter in the user guide.
  6. Larger JSON strings: JsonReader, SimpleJsonReader, and JsonRecordReader now accept string values up to 100,000,000 characters (up from Jackson's 20,000,000 default). See Read a JSON Stream.
  7. More Javadoc: Record, Field, ArrayValue, SingleValue, ValueNode, Node, Session, and the DataReader, DataWriter, ProxyReader, ProxyWriter, Filter, Transformer, and Lookup subclass contracts are now documented.
  8. Bug fixes: MeteredInputStream and ThrottledInputStream no longer change the byte count when a read reaches the end of the stream (see Meter and Throttle DataPipeline Jobs); the streaming Excel reader (ExcelDocument.ProviderType.POI_XSSF_SAX) now opens .xlsx files read-only, so reading no longer rewrites the file on disk when the document is closed (see Streaming Large Workbooks); SqlPart.getSqlFragment() and toString() now remove all trailing whitespace from generated SQL (see How the Builders Work).

Foundations Changes

  1. Pluggable record detection (new feature): the new pipeline.tree.detect package controls how record breaks and fields are detected when XML or JSON is loaded into a Tree. A TreeDetectionStrategy holds the node name rules applied while the tree is built and the transformer passes applied after loading. Opt-in rules collapse date and time keyed nodes (DateTimeNodeNameRule), regex-matched names (RegexNodeNameRule), and numbered or structurally similar siblings (CollapseSiblingPatterns) into a single wildcard node. Six new examples cover the options, starting with Detect Records in JSON with Header and Details and Collapse Numbered Siblings in XML.
  2. Default detection change (breaking): attributes no longer count toward their parent's score and single-instance wrapper nodes are no longer chosen over a repeating structured child, so header-plus-details documents now yield one record per detail element. TreeDetectionStrategy.standard(boolean optimizeTreeStructure) reproduces the previous behavior.
  3. SQL from pipeline actions: PipelineAction now implements SqlSelectGenerator; generateSqlSelect(QuerySource... querySources) returns an equivalent Select for the add, aggregate, copy, rename, select, sort, and remove-duplicates actions, so the database can do the work instead of the JVM. See Generate SQL SELECT from Pipeline Actions.
  4. Column type inference (breaking): multi-digit whole numbers with a leading zero (for example "007") are now inferred as String instead of numeric; NumberMatch.isLeadingZeroWholeNumber() exposes the check.
  5. Datasets: the new protected LocalFileDataset.newInputStream(File) and newOutputStream(File) let subclasses control how dataset files are stored, for example encrypted at rest.
  6. Smaller additions: SortFieldsAction.add(String fieldName, boolean ascending); the FileType constants DOCX, WEBP, OFX, QFX, and QBO; TreeNode.namePattern; and Tree.createFieldName() and Tree.normalizeFieldName() are now public.
  7. Bug fixes: JdbcPipelineOutput.generateJavaCode() now emits the commitBatch setting instead of repeating the autoCloseConnection value; JdbcResultPage.hasNext() no longer returns true on the last page.

Integration Changes

  1. JSONata (new module): Jsonata.jsonata(String expression) compiles a JSONata expression and evaluate(ValueNode<?> input) runs it against a Record, ArrayValue, or SingleValue, returning a ValueNode<?>. Variable bindings (assign) and custom functions (registerFunction) are supported. Read the JSONata chapter or see Query JSON Using JSONata.
  2. LaTeX (new module): LatexWriter and LatexPipelineOutput write records as a table in a LaTeX document. Read the LaTeX chapter or see Write a LaTeX File.
  3. Microsoft Word (new module): MicrosoftWordWriter and MicrosoftWordPipelineOutput write records as a table in a Word (.docx) document, with pageSize and pageOrientation properties. See Write a Microsoft Word File.
  4. Open Financial Exchange (new module): OpenFinancialExchangeReader and OpenFinancialExchangePipelineInput read OFX, QFX, and QBO files (v1 SGML and v2 XML) as one record per file that mirrors the OFX hierarchy. See Read an OFX File and Convert OFX Transactions to CSV.
  5. Google Sheets (new feature): GoogleSheetReader, GoogleSheetWriter, and GoogleSheetDocument read and write spreadsheets in Google Sheets (google module). See Read a Google Sheet, Write to a Google Sheet, and Create and Write a Google Sheet.
  6. Upsert statements to a Writer: MySqlUpsertWriter(String table, Writer writer) and PostgreSqlUpsertWriter(String table, Writer writer, String... keyFieldNames) write upsert statements to any Writer, such as a script file, instead of executing them. See Write MySQL Upsert Statements to a File and Write PostgreSQL Upsert Statements to a File.
  7. Improvements: PdfPipelineOutput now writes through its FileSink's output stream, so it works with non-local sinks; JiraClient.Proxy.get(String domain, String username, String apiKey) now fails with a DataException when any argument is null or empty.
  8. Bug fixes: ParquetDataWriter's write-failure exception now names the field and includes the cause instead of the misleading "Expected value of DOUBLE in record, but found DOUBLE" message; PostgreSqlUpsertWriter no longer throws a NullPointerException when keyFieldNames is null.

File System Changes

  1. Amazon S3 on the AWS SDK for Java 2.x (breaking): the amazon-s3 module moved from AWS SDK for Java 1.x (1.12.794) to 2.x (2.55.9), and the v1 SDK is no longer a dependency. AmazonS3FileSystem (along with AmazonS3File, AmazonS3FileSource, and AmazonS3FileSink) now exposes the v2 S3Client and AwsCredentialsProvider types; setCredentials(AWSCredentials) is replaced by setBasicAWSCredentials(String accessKey, String secretKey) or setCredentialsProvider(); setProfileCredentialsProvider() becomes useProfileCredentialsProvider(String profileName); the endpoint configuration becomes setEndpointOverride(String endpointOverride) plus setRegion(String region); and the listing methods return the SDK-neutral S3Bucket, S3ObjectListing, and S3ObjectSummary types. With no credentials or region configured, the file system now falls back to the default AWS credentials and region chains. A file system can be shared across threads by sources, sinks, and getFileSize() lookups, and streams no longer close a file system the caller opened. See Read from Amazon S3 Using an AWS Profile, List an Amazon S3 Folder One Level at a Time, and Get the Size of an Amazon S3 Object.
  2. Credentials resolvers for every file system (new feature): setCredentialsResolver(CredentialsResolver credentialsResolver) was added to AmazonS3FileSystem, FtpFileSystem, SftpFileSystem, DropBoxFileSystem, GoogleDriveFileSystem, and HdfsFileSystem, so credentials are resolved on each open(), and keys missing from the resolved credentials fall back to the constructor-supplied values. Each file system gained a credential-less constructor for use with a resolver, and AmazonS3FileSystem re-resolves expiring credentials mid-session. See Read from Amazon S3 Using a Rotating Credentials File and Read from SFTP Using a Credentials Resolver.
  3. Google Drive (breaking): GoogleDriveFileSystem now uses GsonFactory instead of JacksonFactory; code using JacksonFactory must add google-http-client-jackson2 or switch to GsonFactory. Its Javadoc was expanded, and listFolder() and readFile() no longer throw a NullPointerException for files without parents or for a missing folder or file.
  4. Other S3 additions and fixes: setDelimiter(String delimiter), getFileSize(String bucket, String key), writeMultipartFile(String bucket, String filePath, String contentType), and AmazonS3File.refreshFileSize(); AmazonS3File.getFileSize() now caches the size and reports the object's real size instead of always returning -1; sources and sinks keep their name, bucket, path, region, endpoint override, and debug settings through toRecord()/fromRecord() and their XML equivalents (credentials are intentionally not serialized); a file system can be reopened after close(); and AmazonS3Util.getExpirationDays() now skips enabled lifecycle rules that have no expiration.
  5. Clearer failures: FtpFileSystem, SftpFileSystem, DropBoxFileSystem, and GoogleDriveFileSystem now fail open() with a clear DataException when credentials are missing.

Documentation and Examples

Upgrading to 11.0

Point your build at the Maven repository and set the DataPipeline version to 11.0.0. For Maven:

<repositories>
  <repository>
    <id>datapipeline</id>
    <url>https://maven.datapipeline.io/public/repositories/datapipeline</url>
    <releases/>
  </repository>
</repositories>

<dependencies>
  <dependency>
    <groupId>com.northconcepts</groupId>
    <artifactId>northconcepts-datapipeline-small-business</artifactId>
    <version>11.0.0</version>
    <scope>compile</scope>
  </dependency>
</dependencies>

For Gradle:

repositories {
    mavenCentral()
    maven {
        url = 'https://maven.datapipeline.io/public/repositories/datapipeline'
    }
}

dependencies {
    implementation 'com.northconcepts:northconcepts-datapipeline-small-business:11.0.0'
}

Swap small-business for your edition (express, team, or enterprise) and add the integration modules you use, such as northconcepts-datapipeline-integrations-jsonata. The Getting Started page has the complete snippets for each edition. Then review the breaking changes above, especially if you use the Amazon S3 file system, the static log fields, or the default tree detection in DataPipeline Foundations.

Get started with DataPipeline 11.0 today to take advantage of the new modules, rotating credentials, and the many enhancements across the core, foundations, integration, and file system layers.

See the CHANGELOG for the full set of updates in DataPipeline 11.0.0.

Also see the Javadocs and examples for more info.

Happy coding!

About The DataPipeline Team

We make Data Pipeline — a lightweight ETL framework for Java. Use it to filter, transform, and aggregate data on-the-fly in your web, mobile, and desktop apps. Learn more about it at datapipeline.io.