Welcome to the 11.0 release of DataPipeline.
This release moves DataPipeline to its new home at DataPipeline.io and raises the minimum Java version to JDK 21. It adds JSONata queries, LaTeX and Microsoft Word output, Open Financial Exchange (OFX) input, Google Sheets, rotating credentials for every file system, a pluggable record detection framework for XML and JSON, and an Amazon S3 connector rebuilt on the AWS SDK for Java 2.x. It also comes with 26 new examples and five new user guide chapters.
A New Home: DataPipeline.io
DataPipeline now lives at DataPipeline.io. The user guide, examples, Javadocs, changelog, and this blog have all moved here, and links to the old northconcepts.com documentation, examples, Javadocs, and blog posts redirect to their new addresses, so existing bookmarks keep working. DataPipeline is still built and supported by North Concepts Inc.; only its address has changed.
A few practical details for the move:
- Maven repository: artifacts are served from
https://maven.datapipeline.io/public/repositories/datapipeline. The previousmaven.northconcepts.comaddress continues to work, so existing builds do not need to change. - No more zip downloads: the library comes from the Maven repository, the examples are on GitHub, and the Javadocs are on this site. The Getting Started page has the dependency snippets for every edition, and your license file is emailed to you after a quick sign-up.
- Support: reach us at support@datapipeline.io.
JDK 21 Minimum
DataPipeline 11.0 requires Java 21 or later and is tested on JDK 21 through 26. DataPipeline 10.1 was the last release to run on Java 8; it stays available in the Maven repository if you are not ready to move yet. Raising the minimum lets DataPipeline use current Java APIs and keep its dependencies current.
Two other breaking changes in the core library come with the recompile:
- The public static
logfields (DataObject.log/DataEndpoint.log,RecordList.log,EventBus.log,DeMux.log,SplitWriter.log, andDataReaderServer.log) are nowcom.northconcepts.datapipeline.logger.Loggerinstead oforg.apache.log4j.Logger. Code that declares them with the Log4j type must switch imports and recompile. Log4j itself remains a dependency. RemoveDuplicateFields.DuplicateFieldsPolicy.RENAMEnow renames duplicate fields in place instead of moving them to the end of the record.
Core Changes
- Credentials resolution (new feature): the new
com.northconcepts.datapipeline.securitypackage resolves credentials on demand so rotated secrets are picked up on the next connection. ACredentialsResolverreturns aCredentialsobject (username, password, access key, secret key, session token, and token, plus an optional expiry), with built-in resolvers for static values, environment variables, system properties, properties files, suppliers, and caching, and aCredentialsResolverRegistryto share them by name.FileSystem.setCredentialsResolver()wires a resolver into any file system (see the file system changes below). Read the Credentials Resolver chapter in the user guide. - TextTableWriter (new feature):
TextTableWriterwrites records as an easy-to-read text table with padded, aligned columns (columnSeparatorandnewLineproperties). See Write a CSV File as a Text Table. - ValueNode conversions:
ValueNode.from(Object value)wraps any value, and the newasSingleValue(),asRecord(),asArrayValue(),toJson(), andtoXml()methods convert between node types and serialize them. See Convert Values and Records to JSON and XML. - Duplicate field renaming:
RemoveDuplicateFieldsaddsDuplicateFieldsPolicy.RENAME_UNDERSCOREandRenamePolicy(String separator)to control the separator used when renaming duplicate fields. See Rename Duplicate Fields with a Separator. - Query ordering: added
Select.orderBy(FieldList fields)andQueryOrder.add(FieldList fields), so a query can be sorted by the fields of a record. Read the SQL Generators chapter in the user guide. - Larger JSON strings:
JsonReader,SimpleJsonReader, andJsonRecordReadernow accept string values up to 100,000,000 characters (up from Jackson's 20,000,000 default). See Read a JSON Stream. - More Javadoc:
Record,Field,ArrayValue,SingleValue,ValueNode,Node,Session, and theDataReader,DataWriter,ProxyReader,ProxyWriter,Filter,Transformer, andLookupsubclass contracts are now documented. - Bug fixes:
MeteredInputStreamandThrottledInputStreamno longer change the byte count when a read reaches the end of the stream (see Meter and Throttle DataPipeline Jobs); the streaming Excel reader (ExcelDocument.ProviderType.POI_XSSF_SAX) now opens .xlsx files read-only, so reading no longer rewrites the file on disk when the document is closed (see Streaming Large Workbooks);SqlPart.getSqlFragment()andtoString()now remove all trailing whitespace from generated SQL (see How the Builders Work).
Foundations Changes
- Pluggable record detection (new feature): the new
pipeline.tree.detectpackage controls how record breaks and fields are detected when XML or JSON is loaded into aTree. ATreeDetectionStrategyholds the node name rules applied while the tree is built and the transformer passes applied after loading. Opt-in rules collapse date and time keyed nodes (DateTimeNodeNameRule), regex-matched names (RegexNodeNameRule), and numbered or structurally similar siblings (CollapseSiblingPatterns) into a single wildcard node. Six new examples cover the options, starting with Detect Records in JSON with Header and Details and Collapse Numbered Siblings in XML. - Default detection change (breaking): attributes no longer count toward their parent's score and single-instance wrapper nodes are no longer chosen over a repeating structured child, so header-plus-details documents now yield one record per detail element.
TreeDetectionStrategy.standard(boolean optimizeTreeStructure)reproduces the previous behavior. - SQL from pipeline actions:
PipelineActionnow implementsSqlSelectGenerator;generateSqlSelect(QuerySource... querySources)returns an equivalentSelectfor the add, aggregate, copy, rename, select, sort, and remove-duplicates actions, so the database can do the work instead of the JVM. See Generate SQL SELECT from Pipeline Actions. - Column type inference (breaking): multi-digit whole numbers with a leading zero (for example "007") are now inferred as String instead of numeric;
NumberMatch.isLeadingZeroWholeNumber()exposes the check. - Datasets: the new protected
LocalFileDataset.newInputStream(File)andnewOutputStream(File)let subclasses control how dataset files are stored, for example encrypted at rest. - Smaller additions:
SortFieldsAction.add(String fieldName, boolean ascending); theFileTypeconstantsDOCX,WEBP,OFX,QFX, andQBO;TreeNode.namePattern; andTree.createFieldName()andTree.normalizeFieldName()are now public. - Bug fixes:
JdbcPipelineOutput.generateJavaCode()now emits thecommitBatchsetting instead of repeating theautoCloseConnectionvalue;JdbcResultPage.hasNext()no longer returns true on the last page.
Integration Changes
- JSONata (new module):
Jsonata.jsonata(String expression)compiles a JSONata expression andevaluate(ValueNode<?> input)runs it against aRecord,ArrayValue, orSingleValue, returning aValueNode<?>. Variable bindings (assign) and custom functions (registerFunction) are supported. Read the JSONata chapter or see Query JSON Using JSONata. - LaTeX (new module):
LatexWriterandLatexPipelineOutputwrite records as a table in a LaTeX document. Read the LaTeX chapter or see Write a LaTeX File. - Microsoft Word (new module):
MicrosoftWordWriterandMicrosoftWordPipelineOutputwrite records as a table in a Word (.docx) document, withpageSizeandpageOrientationproperties. See Write a Microsoft Word File. - Open Financial Exchange (new module):
OpenFinancialExchangeReaderandOpenFinancialExchangePipelineInputread OFX, QFX, and QBO files (v1 SGML and v2 XML) as one record per file that mirrors the OFX hierarchy. See Read an OFX File and Convert OFX Transactions to CSV. - Google Sheets (new feature):
GoogleSheetReader,GoogleSheetWriter, andGoogleSheetDocumentread and write spreadsheets in Google Sheets (google module). See Read a Google Sheet, Write to a Google Sheet, and Create and Write a Google Sheet. - Upsert statements to a Writer:
MySqlUpsertWriter(String table, Writer writer)andPostgreSqlUpsertWriter(String table, Writer writer, String... keyFieldNames)write upsert statements to anyWriter, such as a script file, instead of executing them. See Write MySQL Upsert Statements to a File and Write PostgreSQL Upsert Statements to a File. - Improvements:
PdfPipelineOutputnow writes through itsFileSink's output stream, so it works with non-local sinks;JiraClient.Proxy.get(String domain, String username, String apiKey)now fails with aDataExceptionwhen any argument is null or empty. - Bug fixes:
ParquetDataWriter's write-failure exception now names the field and includes the cause instead of the misleading "Expected value of DOUBLE in record, but found DOUBLE" message;PostgreSqlUpsertWriterno longer throws aNullPointerExceptionwhenkeyFieldNamesis null.
File System Changes
- Amazon S3 on the AWS SDK for Java 2.x (breaking): the amazon-s3 module moved from AWS SDK for Java 1.x (1.12.794) to 2.x (2.55.9), and the v1 SDK is no longer a dependency.
AmazonS3FileSystem(along withAmazonS3File,AmazonS3FileSource, andAmazonS3FileSink) now exposes the v2S3ClientandAwsCredentialsProvidertypes;setCredentials(AWSCredentials)is replaced bysetBasicAWSCredentials(String accessKey, String secretKey)orsetCredentialsProvider();setProfileCredentialsProvider()becomesuseProfileCredentialsProvider(String profileName); the endpoint configuration becomessetEndpointOverride(String endpointOverride)plussetRegion(String region); and the listing methods return the SDK-neutralS3Bucket,S3ObjectListing, andS3ObjectSummarytypes. With no credentials or region configured, the file system now falls back to the default AWS credentials and region chains. A file system can be shared across threads by sources, sinks, andgetFileSize()lookups, and streams no longer close a file system the caller opened. See Read from Amazon S3 Using an AWS Profile, List an Amazon S3 Folder One Level at a Time, and Get the Size of an Amazon S3 Object. - Credentials resolvers for every file system (new feature):
setCredentialsResolver(CredentialsResolver credentialsResolver)was added toAmazonS3FileSystem,FtpFileSystem,SftpFileSystem,DropBoxFileSystem,GoogleDriveFileSystem, andHdfsFileSystem, so credentials are resolved on eachopen(), and keys missing from the resolved credentials fall back to the constructor-supplied values. Each file system gained a credential-less constructor for use with a resolver, andAmazonS3FileSystemre-resolves expiring credentials mid-session. See Read from Amazon S3 Using a Rotating Credentials File and Read from SFTP Using a Credentials Resolver. - Google Drive (breaking):
GoogleDriveFileSystemnow usesGsonFactoryinstead ofJacksonFactory; code usingJacksonFactorymust addgoogle-http-client-jackson2or switch toGsonFactory. Its Javadoc was expanded, andlistFolder()andreadFile()no longer throw aNullPointerExceptionfor files without parents or for a missing folder or file. - Other S3 additions and fixes:
setDelimiter(String delimiter),getFileSize(String bucket, String key),writeMultipartFile(String bucket, String filePath, String contentType), andAmazonS3File.refreshFileSize();AmazonS3File.getFileSize()now caches the size and reports the object's real size instead of always returning -1; sources and sinks keep their name, bucket, path, region, endpoint override, and debug settings throughtoRecord()/fromRecord()and their XML equivalents (credentials are intentionally not serialized); a file system can be reopened afterclose(); andAmazonS3Util.getExpirationDays()now skips enabled lifecycle rules that have no expiration. - Clearer failures:
FtpFileSystem,SftpFileSystem,DropBoxFileSystem, andGoogleDriveFileSystemnow failopen()with a clearDataExceptionwhen credentials are missing.
Documentation and Examples
- Five new user guide chapters: Excel, LaTeX, JSONata, Credentials Resolver, and SQL Generators.
- 26 new examples covering the features above, all in the DataPipeline-Examples project on GitHub, which now builds with Java 21 against 11.0.0.
- The Integrations page lists every add-on module with its Maven and Gradle coordinates.
Upgrading to 11.0
Point your build at the Maven repository and set the DataPipeline version to 11.0.0. For Maven:
<repositories>
<repository>
<id>datapipeline</id>
<url>https://maven.datapipeline.io/public/repositories/datapipeline</url>
<releases/>
</repository>
</repositories>
<dependencies>
<dependency>
<groupId>com.northconcepts</groupId>
<artifactId>northconcepts-datapipeline-small-business</artifactId>
<version>11.0.0</version>
<scope>compile</scope>
</dependency>
</dependencies>
For Gradle:
repositories {
mavenCentral()
maven {
url = 'https://maven.datapipeline.io/public/repositories/datapipeline'
}
}
dependencies {
implementation 'com.northconcepts:northconcepts-datapipeline-small-business:11.0.0'
}
Swap small-business for your edition (express, team, or enterprise) and add the integration modules you use, such as northconcepts-datapipeline-integrations-jsonata. The Getting Started page has the complete snippets for each edition. Then review the breaking changes above, especially if you use the Amazon S3 file system, the static log fields, or the default tree detection in DataPipeline Foundations.
Get started with DataPipeline 11.0 today to take advantage of the new modules, rotating credentials, and the many enhancements across the core, foundations, integration, and file system layers.
See the CHANGELOG for the full set of updates in DataPipeline 11.0.0.
Also see the Javadocs and examples for more info.
Happy coding!