All Classes and Interfaces
Class
Description
Base class for field mappings that set one target field, with an optional condition, default value expression and
type conversion.
Used to run, pause, and cancel Runnables.
A default, abstract implementation of JobCallback.
Base class for declarative pipelines that read records from a
PipelineInput, process them and write them to a
PipelineOutput.Abstract super-class, with some common logic, for reading records.
Base class for Twitter API v2 readers; each page from
AbstractTwitterReader.search() yields one record per data element.Abstract super-class, with some common logic, for writing records.
Categories of built-in
PipelineActions (transform, convert, aggregate, filter and validate), each holding a default
instance of the actions it lists.A
PipelineAction that sets fields on every record, replacing existing ones, to a typed constant or the result of an
expression.The type a field's value is converted to, or
EXPRESSION to evaluate it as an expression.A field's value and the type it is converted to.
Derives a
SiblingSetPattern from siblings sharing a fixed prefix/suffix around a varying
dynamic segment — for example item-1, item-2, item-3
collapse to item-*.A
PipelineAction that groups records by the fields marked GROUPBY and aggregates the others (count, sum,
average, first, last, minimum or maximum) using a GroupByReader.A source field and the operator applied to it: either a grouping key (
GROUPBY) or an aggregate written to the
target field.The operator applied to a
AggregateGroupFieldsAction.GroupByField: GROUPBY marks a grouping key and the others are aggregates.Deprecated.
Base class for the set functions an
AggregateReader applies to each record passing through it.Base class for a file in Amazon S3, located by bucket name and path and accessed through an
AmazonS3FileSystem that other files may share.A
FileSink that writes an Amazon S3 object with a multipart upload through an AmazonS3FileSystem,
connecting on demand if the file system is not open.A
FileSource that reads an Amazon S3 object through an AmazonS3FileSystem, connecting on demand if
the file system is not open.Class for accessing an Amazon S3 file system.
Adapts the backend-neutral
so resolved credentials flow through the native client natively.
CredentialsResolver seam to the AWS SDK's
invalid reference
AwsCredentialsProvider
Static helpers that write and read object expiration rules in an S3 bucket's lifecycle configuration through an
S3ControlClient.Matches all filers -- filer A invalid input: '&'invalid input: '&' filer B invalid input: '&'invalid input: '&' filer C.
A
ConjunctivePart that joins its parts with AND.Indicates the action to take when the Twitter API call limits are reached.
Exposes an
ArrayValue to FreeMarker templates as a sequence whose elements are wrapped by the object wrapper.ArrayValue holds an ordered collection (also known as a sequence) of persistent data.
The public API for a node in the expression language's abstract syntax tree.
A
DecryptingReader for fields written by an AsymmetricEncryptingReader; it recovers the symmetric key
from the first value it decrypts, using the private key, and uses it for all values.An
EncryptingReader that encrypts fields with a symmetric key and appends that key, encrypted with a public
key, to every value after a -; read the values back with an AsymmetricDecryptingReader.AsyncMultiReader reads from one or more source DataReaders asynchronously using a separate thread for each one.
A proxy that reads data asynchronously using a separate thread.
A proxy that uses multiple threads to process incoming data, where the work applied to the incoming data is specified by the a
DataReaderDecorator.A proxy that writes data asynchronously using a separate thread.
User-defined name-value pairs that attach arbitrary metadata to schema, entity, relationship and field definitions;
kept sorted by name.
A pipeline input that reads Apache Avro data from a
FileSource using an AvroReader.A pipeline output that writes Apache Avro data to a
FileSink using an AvroWriter.Read an Apache Avro file and convert the contents into
Records .Write
Records to an Apache Avro file.See
FieldPath for supported field name expressions.A named step that a
BasicFieldTransformer applies to each of its fields.An
BasicFieldTransformer.Operation applied to each single value in a field, including array elements and values in nested records.A
BasicFieldTransformer.SingleValueOperation that transforms the string form of each value into a new string.An in-memory
Lookup filled with add calls, where each key maps to one or more result records.Base class for Java beans wishing to participate in Java, Record, XML, and JSON serialization;
adding their properties to DataException; and having a dynamic toString().
A
FieldDef for binary (byte[]) values with optional minimum and maximum lengths; allowed values are
not supported.Abstract super-class for writing records to a binary stream.
A Bloomberg message parsed by
BloombergMessage.parse(Reader) into its header, field list, data records and footer.The data section of a
BloombergMessage, holding one record per data line.The
name=value properties that follow the data section of a BloombergMessage; a name can have several values.The
name=value properties that precede the data section of a BloombergMessage; a name can have several values.Reads the data records of a Bloomberg message; the whole message is parsed into memory on open and stays available from
BloombergMessageReader.getMessage().Selects which boards
TrelloBoardsReader reads; toString() gives the lowercase value sent to Trello.Selects which of each board's lists
TrelloBoardsReader requests; toString() gives the lowercase value sent to Trello.A
FieldDef for boolean values.A proxy that organizes incoming data by collecting records of the same type (using values in a subset of fields) to release them downstream together.
Determines when
BufferedReader closes an open sliding buffer.A
Lookup that lazily caches the results of successful lookups performed by another instance.A
CredentialsResolver decorator that caches its delegate's credentials, re-fetching when:
the cached credentials report expiry within CachingCredentialsResolver.getExpirySkewMilliseconds() (see
Credentials.isExpired(long)),
the optional time-to-live has elapsed, or
CachingCredentialsResolver.refresh() is called.
Gives rotation with backpressure — remote secret stores are only called when needed.A variable computed from an expression before a decision table or tree evaluates its rules;
the result is stored in the expression context under the variable's name.
Selects which cards
TrelloBoardCardsReader and TrelloBoardListsReader request; toString() gives the
lowercase value sent to Trello.Determines when
GroupByReader closes an open sliding window.A programming-language agnostic class for emitting formatted code using indentation and line breaks.
Collapses down (merges) sibling tree-node-paths whose names share a derivable naming/wildcard
pattern (see
SiblingSetRule) into a single wildcard node, mirroring what the build-time
NodeNameRules produce for individually dynamic names: the collapsed node is named
*, its XPath expression ends in node(), member names become its values,
and descendant XPaths collapse the merged segment to *.Copies the values in specific fields to an in-memory
Collection.Holds the metadata and statistics for each field in a dataset.
Obtains records from a webserver's access log file using the Combined Log format.
A list of values parsed from one comma-separated line;
toString() joins them back with commas.A collection of comma delimited values.
A
PipelineAction that applies a list of nested actions in order, as a single step.A fixed-size list of values used as a composite key (for example in lookups, grouping and duplicate removal);
a single-value instance hashes like, and equals, that value itself.
A
RecordList that guards many of its methods with a read-write lock; others, such as size() and
clear(), are unguarded, as are the iterators, streams and lists it hands out.A single SQL condition, such as
"price > ?", with the values for its ? parameters.Base class for parts that join their child parts with an operator such as AND or OR, parenthesizing each child when
more than one applies.
Constants required by DateTimeUtils
A
FieldFilterRule that allows fields whose value, as a string, contains the given text; null values are rejected.A
PipelineAction that converts the selected date/time fields to strings, optionally using a SimpleDateFormat
pattern.A
PipelineAction that replaces null values in the selected fields with a typed constant or the result of an
expression.The type a replacement value is converted to, or
EXPRESSION to evaluate it; DATETIME, DATE and
TIME values use the yyyy-MM-dd HH:mm:ss, yyyy-MM-dd and HH:mm:ss formats.A replacement value and the type it is converted to.
A
PipelineAction that converts the selected numeric fields to strings, optionally using a DecimalFormat
pattern.Renames schema definitions in place from database-style snake_case names to Java-style camel case.
A
PipelineAction that converts the selected string fields to booleans; only "true", ignoring case, becomes
true.A
PipelineAction that parses the selected string fields into date-time, date or time values using a
SimpleDateFormat pattern.The temporal type strings are parsed into: a date-time (
java.util.Date), a date (java.sql.Date) or a time
(java.sql.Time).A
PipelineAction that parses the selected string fields into numbers of the chosen type, optionally using a
DecimalFormat pattern.The number type strings are parsed into;
SHORT, BYTE and FLOAT leave values unconverted when a
pattern is set.A
PipelineAction that converts the values of the selected fields, whatever their type, to strings.Creates a duplicate field with the specified target name.
A
PipelineAction that copies the values of source fields to target fields.A copy of one source field to a target field, optionally overwriting an existing target.
Generates an
ALTER TABLE...ADD CONSTRAINT statement to add a foreign key constraint in H2.Generates an
ALTER TABLE...ADD CONSTRAINT statement to add a foreign key constraint in MySQL.Generates an
ALTER TABLE...ADD CONSTRAINT statement to add a foreign key constraint in PostgreSQL.Generates H2 DDL from
SchemaDefGenerates a
CREATE INDEX statement for H2.Generates a
CREATE INDEX statement for MySQL.Generates a
CREATE INDEX statement for PostgreSQL.Generates MySQL DDL from
SchemaDefGenerates PostgreSQL DDL from
SchemaDefGenerates a
CREATE TABLE statement for H2.Generates a
CREATE TABLE statement for MySQL.Generates a
CREATE TABLE statement for PostgreSQL.Generates a column definition for a
CREATE TABLE statement for H2.Generates a column definition for a
CREATE TABLE statement for MySQL.Generates a column definition for a
CREATE TABLE statement for PostgreSQL.Determines when
GroupByReader opens a new sliding window.An immutable bundle of resolved secret values keyed by name, with an optional expiry.
Assembles
Credentials from key-value pairs and an optional expiry; see Credentials.builder().Resolves
Credentials on demand.A thread-safe mapping of logical ids to
CredentialsResolvers, enabling
"reference, not secret" late binding: serialized pipelines and sources/sinks carry only a
resolver id and are re-wired to the live resolver on load.Reads a pipeline's source data from a CSV or other delimited
FileSource using a CSVReader.Obtains records from a Comma Separated Value (CSV) or delimited stream.
Writes records to a Comma Separated Value (CSV) stream.
The interface used by null and empty string writers.
The default implementation used to write null and empty string.
Selects whether an
FtpFileSystem uses active or passive mode for its data connections.Abstract super-class for reading and writing records.
Lifecycle states of an endpoint:
NEW until opened, then OPENED, then CLOSED.Helper class--makes it easy to work with multiple readers and writers as a single unit.
Indicates an error has occurred during the course of execution.
An interface for classes capable of adding context/information to exceptions usually to aid in identifying and troubleshooting issues.
The top-level abstraction to map a set of values from source to target using the DataPipeline Expression Language or custom implementation.
A
PipelineAction that maps each record to a new record using a DataMapping.Applies a
DataMapping to records from a DataReaderFactory and pages through the mapped results in memory,
keeping a sample of the records that failed.The expression context used by a
DataMapping, exposing the source, target,
previousSource and previousTarget contexts to expressions by name.Base class for the data mapping classes:
DataMappingParts, problems and results.Base class for the elements of a data mapping -- the
DataMapping itself and its field mappings -- that
DataMappingProblems refer to.The kinds of
DataMappingPart that DataMappingProblems can refer to.A pipeline that maps each input record to a new record using a
DataMapping; records failing the mapping go to the
discard writer, if one is set, instead of stopping the pipeline.Describes an issue found when checking a data mapping and the
DataMappingPart it was found in.Implemented by data mapping elements that can check themselves, and optionally their children, for
DataMappingProblems.Maps each record read from a nested reader with a
DataMapping and returns the target record instead; records
failing a mapping condition are skipped.The outcome of mapping one record with a
DataMapping: the target record plus any failed condition, field
mapping exceptions and schema validation results.Static checks that data mapping parts use to add
DataMappingProblems to a list.Maps each record with a
DataMapping and writes the target record to a nested writer instead; records failing a
mapping condition are not written.Abstract super-class for classes that work with data and require a consistent error handling mechanism.
Abstract super-class for reading records.
A processing step, such as a conversion, filter or transformation, applied to records by wrapping their
DataReader.A Decorator that allows a DataReader to be wrapped in new behavior.
A factory to create new DataReaders.
A DataReader that sequentially reads from an Iterator of other DataReaders.
A
Lookup that loads all records from a DataReader into memory, keyed by the parameter fields;
the reader is read and closed in the constructor.The base class for caching records produced by a
Pipeline or DataMappingPipeline.Reads one summary record per column, with its name, counts, inferred type, lengths and sample value.
Obtains records from those cached in a
Dataset.Obtains records from those cached in a
Dataset.Traverses records from those cached in a
Dataset.Abstract super-class for writing records.
A Decorator that allows a DataWriter to be wrapped in new behavior.
A factory to create new DataWriters.
Writes a pipeline's records to the
DataWriter created by a DataWriterFactory.A
FieldFilterRule that allows fields whose date-time value equals the given date to the millisecond.A
FieldFilterRule that allows fields whose date-time value is after the given date; null values are rejected.A
FieldFilterRule that allows fields whose date-time value is before the given date; null values are also allowed.A date parameter that can be created from a
yyyy-MM-dd string or a Date.Classifies date/timestamp names (for example
2024-01-15 map keys) as dynamic using a
DateTimePatternDetector.A serializable, comparable wrapper for
DateTimeFormatter that can be used as keys in Map and values in Set.A utility for parsing date, time, and date-time values from string using a set of possible patterns.
An immutable value parsed from string using
DateTimePattern.A proxy that prints records passing through to a stream in a human-readable format
A proxy that prints records passing through to a stream in a human-readable format
Evaluates rules in order against a record and returns the outcomes of the first rule whose conditions all hold, or the
default outcomes if none does.
A
PipelineAction that evaluates a DecisionTable for each record and copies the resulting outcome fields into
it.A condition of a
DecisionTableRule: a logical expression in which ? stands for the condition's
variable.Base class for the decision table classes: the table, its rules, conditions and outcomes, and its results.
An outcome of a
DecisionTableRule, or a default outcome of a DecisionTable: sets the result field named
by its variable to the value of its expression.Evaluates a
DecisionTable for each record read from a nested reader and adds the outcome fields to the record.The result of evaluating a
DecisionTable: the rule that fired and the outcome fields.A rule of a
DecisionTable that fires when all its conditions hold, setting the result fields with its outcomes.Evaluates a
DecisionTable for each record and adds the outcome fields to it before writing it to a nested
writer.Evaluates a record by walking from the root node, at each node entering the first child whose condition holds (or else
the default node), and combines the outcomes of every node visited.
A
PipelineAction that evaluates a DecisionTree for each record and copies the resulting outcome fields into
it.A node of a
DecisionTree: the condition for entering it, the outcomes it sets when visited, its child nodes and
the default node entered when no child's condition holds.Base class for the decision tree classes: the tree, its nodes and outcomes, and its results.
An outcome of a
DecisionTreeNode: sets the result field named by its variable to the value of its expression.Evaluates a
DecisionTree for each record read from a nested reader and adds the outcome fields to the record.The result of evaluating a
DecisionTree: the path of nodes entered and the combined outcome fields.Callback for
DecisionTree.visit(DecisionTreeVisitor), called once for each node of the tree.Evaluates a
DecisionTree for each record and adds the outcome fields to it before writing it to a nested
writer.Base class for proxy readers that decrypt the selected fields of each record in place, or every field when none are selected.
A default implementation of JobCallback.
Builds a
DELETE FROM table WHERE ... statement with its parameter values.Converts a single source DataReader into many using one of the provided strategies.
Strategy for distributing the source records among the readers created by
DeMux.createReader().The built-in strategies: send every record to all readers or each record to one reader in turn.
Marks every
TreeNode that held at least one text value as a field, assigns its initial
field name, and flags values that should cascade across multiple output records.Finds the fields, or smallest field combinations, whose values are unique across all records of a
RecordList.Marks the repeating
TreeNodes that represent one record per instance.Helps diagnose environmental issues by collecting properties and logging them to the console.
A node in a tree of differences: whether a named value was added, removed, changed or left unchanged,
with its old and new values and the diffs of its children.
The kind of difference recorded by a
Diff.Class for accessing the DropBox file system.
Generates a
DROP TABLE statement for H2.Generates a
DROP TABLE statement for MySQL.Generates a
DROP TABLE statement for PostgreSQL.Generates a
ALTER TABLE...DROP CONSTRAINT statement for H2.Generates a
ALTER TABLE...DROP CONSTRAINT statement for MySQL.Generates a
ALTER TABLE...DROP CONSTRAINT statement for PostgreSQL.The order in which an
EmailReader reads the messages it found: first to last or last to first.Reads emails (and their attachments) from IMAP mailboxes.
Email format (HTML or text) of a MailChimp
ListMember.Converts values to and from an encoded form, such as the password stored by a
JdbcConnection.An
Encoder that converts strings to and from the hexadecimal form of their UTF-8 bytes;
null and empty strings are returned unchanged.An
Encoder that returns values unchanged.Base class for proxy readers that encrypt the selected fields of each record in place, or every field when none are selected.
Base class for anything that is opened and closed, such as readers, writers and file systems; tracks its state,
open and close times, and elapsed time.
The metadata structure defining a file, JSON object, database table, or class for validation and mapping.
The differences between two versions of an
EntityDef, with child diffs for its properties, fields and indexes.Indicates how multi-valued values should be returned.
The metadata model for relating two entities in a schema.
The property differences between two versions of an
EntityRelationshipDef.A
CredentialsResolver that reads credential values from environment variables
at each resolve.A method call made on an
EventBus publisher, captured with its arguments for delivery to the bus's listeners.An asynchronous, in-memory, event delivery service.
The lifecycle states of a bus; listeners, publishers and events are only accepted while
ALIVE.Notified when any
EventBus is created or begins shutting down; register with
EventBus.addEventBusLifecycleListener(EventBusLifecycleListener).Reads the records published to an
EventBus on the given topics (all topics if none are given), for example
by an EventBusWriter.Writes records to an event bus under a specific topic.
Decides which events a listener receives; supplied when the listener is added to an
EventBus.Resource types for
ShopifyEventCriteria.setFilters(EventFilter...); EventFilter.toString() returns the lower-case name.Matches events published through any one of a set of listener interfaces (all events if none are given).
Matches events whose source is one of a set of objects, compared by identity (all events if none are given).
Backs the publisher proxies returned by
EventBus.getPublisher(Object, Class, Object), publishing each
method call to the bus as an Event.Event verbs for
ShopifyEventCriteria.setVerb(EventVerb); EventVerb.toString() returns the lower-case name.Represents the styling properties that can be applied to an Excel cell.
A functional interface for applying styling to Excel cells based on their location.
Represents a color that can be used in Excel cells.
Excel-compatible indexed color palette (BIFF legacy)
The in-memory abstraction for an Excel workbook.
Selects the Excel library (Apache POI or JXL) and workbook format an
ExcelDocument uses.Provides access to Excel-specific metadata attached to a
Field.Represents font styling options for Excel cells.
Represents a hyperlink that can be attached to an Excel cell.
A functional interface that creates
ExcelHyperlink instances based on a FieldLocation.Enumeration of the types of hyperlinks supported in Excel.
Reads a pipeline's source data from a sheet of an Excel
FileSource using an ExcelReader.A pipeline output that writes to an Excel file using the
ExcelWriter.Obtains records from a Microsoft Excel document.
Indicates how to handle failures when evaluating cell formulas/expressions.
Writes records to a Microsoft Excel document.
Notified when a listener on an
EventBus throws an exception while receiving an event; register with
EventBus.addExceptionListener(ExceptionListener).Deprecated.
use
RemoveFieldsParses expression-language text into a syntax tree (
ASTNode).An performant alternative to
RenameField that assumes
1) a flat table structure (no nested fields or arrays),
2) columns are always in the same position from row to row
3) no missing fields (nulls field values are okay), and
4) the target name does not already exist (otherwise there will be two fields with the same name).Field holds persistent key-value data as part of a record.
See
FieldPath for supported field name expressions.A filter that checks for the total number of fields
Declares the Java type of named fields for expressions such as
FilterExpression and SetCalculatedField;
declared types take precedence over those inferred from each record.The metadata structure of one column or property in a file, JSON object, database table, or class for validation and mapping.
The property differences between two versions of a
FieldDef.A filter that checks if all the specified field exists.
See
FieldPath for supported field name expressions.Base class for the per-field conditions of a
FieldFilter; a record is kept only if every rule allows each checked field.A FreeMarker key/value pair holding a
Field's name and value, both wrapped by the object wrapper.Iterates over a
Record's fields in order as FreeMarker key/value pairs (FieldKeyValuePairs).Convenience wrapper to read and write field-level data lineage properties.
See
FieldPath for supported field name expressions.Represents the location of a field within a data pipeline context.
A functional interface for predicates that test
FieldLocation instances.A field mapping that sets one target field to the result of an expression evaluated in a
DataMappingExpressionContext.A filter that checks if all the specified field does not exist.
An abstract representation for the location of a field within a record.
See
FieldPath for supported field name expressions.FieldType lists of all field data types used in records.
A
CredentialsResolver that reads credential values from a Java properties file,
re-reading it whenever the file's last-modified time changes.Detects the
FileType of a file or stream from its name's extension or, failing that, by sampling its first
bytes.Obtains records from a binary stream previously written using a
FileWriter.A local or remote file to write, supplying an
OutputStream for its content.Base class for pipeline outputs that write their data to a
FileSink.A local or remote file to read, supplying an
InputStream for its content.Base class for pipeline inputs that read their data from a
FileSource.Abstract class for reading and writing to virtual file systems.
Creates records from the metadata of files within a directory.
Selects how an
FtpFileSystem transfers file data: as a stream, in blocks or compressed.Selects whether an
FtpFileSystem transfers files as ASCII text or as binary.File formats known to DataPipeline, each with a display name, MIME type, binary flag and default file extension.
Writes records to a binary stream that can be later read using a
FileReader.Base class for conditions that decide whether a record is kept or discarded, for example by a
FilteringReader.A
Filter that keeps records for which a logical expression-language condition evaluates to true.A
PipelineAction that keeps only the records containing all the selected fields.A
PipelineAction that keeps only the records whose fields each fully match their regular expression.A
PipelineAction that keeps only the records whose selected fields are all non-null.A proxy that chooses records using a filter criteria.
A
PipelineAction that keeps only the records for which a boolean expression evaluates to true.Order financial statuses for
ShopifyOrder.setFinancialStatus(FinancialStatus); FinancialStatus.toString() returns the lower-case name.Alignment of a value within its fixed-width column; the fill character pads (when writing) or is trimmed from (when reading)
the opposite side.
Describes one column (name, width, alignment and fill character) of the records handled by
FixedWidthReader
and FixedWidthWriter.Reads a pipeline's source data from a fixed-width text
FileSource using a FixedWidthReader.Writes a pipeline's records to a fixed-width text
FileSink using a FixedWidthWriter.Obtains records from a fixed width stream.
Writes records to a fixed width stream.
Flattens an array field by creating one record for each value.
A foreign key's ON UPDATE or ON DELETE referential action in an
EntityRelationshipDef.Represents foreign key constraint actions for H2.
Formats the error messages used across DataPipeline Foundations; each static method returns one message built from its arguments.
Abstract super-class for DP Foundation classes that require a consistent error handling mechanism.
Class for accessing the FTP file system.
Order fulfillment statuses for
ShopifyOrder.setFulfillmentStatus(FulfillmentStatus); FulfillmentStatus.toString() returns the lower-case name.Generates a full join clause.
Functions is the entry point for adding new method aliases to the dynamic expression language.
A utility class for generating EntityDef code from a DataReader.
Enumeration of field type selection modes for determining how to select field types from column statistics.
Prints generated
hashCode() and equals() methods covering the non-static, non-transient fields declared
by the given classes.Generates a class for each query registered with a
JdbcConnection, with a field per result column and static
executeQuery methods.A utility class for generating toRecord() and fromRecord(Record source) methods from existing Java beans.
Builds a
SchemaDef from JDBC metadata: an entity per table, relationships from foreign keys, and the indexes
other than primary keys.A utility class for generating SchemaDef code from a Record or JSON object.
A utility class for generating Spring Data JPA entities and repositories using the metadata from a live database.
A utility class for generating data access Java beans using the metadata from a live database.
The class name, package and output file used to generate the class for one table.
An upsert strategy that attempts to either:
Retrieves all query parameter values in a URL for the given name and stores them in the target field.
Names of the standard Gmail folders, for use with
EmailReader.setFolderName(String).Reads Gmail messages, optionally selected by a search query, one record per message with its addresses, headers, flags,
content parts and Gmail metadata such as label and thread IDs.
Reads a Google Analytics (Reporting API v4) report for a view, one record per report row holding its dimension values
followed by its metric values for each date range.
Base class for readers that call Google APIs, holding the OAuth credential, HTTP transport and application name used to
build the API client.
Reads the events of a Google Calendar, one record per event with its details, attendees, attachments, conference data and
calendar properties.
Reads a Google account's contacts through the People API, one record per person with their names, email addresses, phone
numbers, addresses and other details.
A
FileSystem that reads and writes files stored in
Google Drive.A Google Sheets spreadsheet shared by
GoogleSheetReader and GoogleSheetWriter.Reads records from one sheet, or one range, of a Google Sheets spreadsheet opened as a
GoogleSheetDocument; the
sheet is the one the range names, else the named sheet, else the sheet at the index.Writes records to one sheet of a Google Sheets spreadsheet opened as a
GoogleSheetDocument, creating the sheet
when its name is new and clearing the target block first.A
GroupOperation that averages the source field in each group as a BigDecimal; unless
excludeNulls is set, null values count as zero.A proxy that divides records into groups and applies summary operations to
each group; similar to "group by" in SQL, but applied to streaming data.
Collects the individual values into an array.
A
GroupOperation that counts the records in each group.A
GroupOperation that returns the first value of the source field in each group (the first non-null value when
excludeNulls is set).A
GroupOperation that returns the last value of the source field in each group (the last non-null value when
excludeNulls is set).A
GroupOperation that returns the largest non-null value of the source field in each group.A
GroupOperation that returns the smallest value of the source field in each group; unless excludeNulls
is set, a null value discards the minimum found so far.See
FieldPath for supported field name expressions.Accumulates one
GroupOperation's result for a single group within a Window.A
GroupOperation that totals the non-null values of the source field in each group as a BigDecimal.Generates H2 insert statement from records.
Base class for H2 SQL components.
Class for accessing the Hadoop Distributed File System.
Represents a schedule of fixed minutes past each hour (for example: 0, 15, 30, and 45 minutes past each hour).
Indicates the action to take when the Twitter API call limits are reached.
Indicates how multi-valued values should be returned.
A mapping that sets one field of a
DataMapping's target record.The strategy used to write records to a database.
A unit of work that can be run synchronously or asynchronously, paused, resumed, cancelled and monitored.
Renders the USING clause of a
Merge statement.Deprecated.
use
SelectFieldsThe metadata model for an entity's index.
The differences between two versions of an
IndexDef, with child diffs for its properties and index fields.The metadata model for a field in an entity's index.
The property differences between two versions of an
IndexFieldDef.Generates an INNER JOIN clause.
Receives callbacks while
NodeVisitor.visit(Node, INodeVisitor) walks a tree of records, fields, arrays and values depth-first.Opens input streams on demand; name an implementation (with a public no-argument constructor) in the
datapipeline.license.factory system property to load the license from it.Builds an
INSERT INTO table (columns) VALUES (?, ...), ... statement with one or more rows; the column list
comes from the first row.Generates an
INSERT INTO statement for H2.Generates an
INSERT INTO statement for MySQL.Generates an
INSERT INTO statement for PostgreSQL.One row of an
Insert, rendered as a parenthesized list of ? placeholders.Represents a row of values for H2 insert statements.
A column and its parameter value in an
InsertRow.Represents a column-value pair for H2 insert statements.
A column name and value pair in an
InsertRow.A column name and value pair in an
InsertRow.Search Instagram for media given a specific location (latitude, longitude).
Search for a location by geographic coordinate in Instagram.
Base class for all InstagramXXXMediaReader.
Get the list of users this user is followed by.
Get the list of users this user follows.
Get your most recent media published.
Base class for all InstagramSelfXXXReader.
Search Instagram for media given a specific tag.
Search for matching tags in Instagram by name - results are ordered first as an exact match, then by popularity.
Get the most recent media published by a user.
Base class for proxy readers in the license-gated integration modules;
IntegrationProxyReader.open() leaves the reader unopened
unless the product edition and license allow integrations.Base class for proxy writers in the license-gated integration modules;
IntegrationProxyWriter.open() leaves the writer unopened
unless the product edition and license allow integrations.Base class for readers in the license-gated integration modules;
IntegrationReader.open() leaves the reader unopened
unless the product edition and license allow integrations.Base class for writers in the license-gated integration modules;
IntegrationWriter.open() leaves the writer unopened
unless the product edition and license allow integrations.Character-level parser with lookahead, used by the text readers to peek at, match and consume their input.
Strategy for computing how long to wait before each retry of a failed read, write or operation.
Field filter rule that returns true if a field value is empty.
A
FieldFilterRule that allows fields of the given FieldType.A
FieldFilterRule that allows fields whose value is an instance of the given Java type or a subtype; null values
are rejected.A
FieldFilterRule that allows fields whose value is exactly of the given class (not a subclass); null values
are also allowed.Field filter rule that returns true if a field value is not empty.
A
FieldFilterRule that allows fields with non-null values.A
FieldFilterRule that allows fields with null values.Copies Twitter tweets, users and user lists into records;
TwitterConverter is the default implementation.The strategy used to upsert records to a database.
Obtains records from an object using bean introspection and reflection.
A simple model containing generated Java code, imports, and variable names.
Interface for classes capable of participating in Java code generation.
Describes a database (driver, URL, credentials and driver properties), creates JDBC connections to it
and loads its table, schema and query metadata.
A consistent interface for creating JDBC connections.
Callbacks fired while
JdbcConnection.loadTables(LoadTablesRequest) runs: a table... method after each
table's step and a plural one once the step is done for all tables.Caches the dataset's records on database.
Enumeration of supported database providers for JDBC datasets.
One column of a foreign key that references the owning table's primary key, as reported by JDBC
getExportedKeys.One column of a
JdbcIndex: its name, position in the index and sort direction.One column of a foreign key: the referencing (foreign key) column and the primary key column it points to, as reported
by JDBC
getImportedKeys; a multi-column key has one entry per column.Metadata for a table index and its columns, as reported by JDBC
DatabaseMetaData.getIndexInfo.A
Lookup that runs a SQL query for each lookup, binding the key values to its parameters in order.Writes records to a database table using the upsert idiom (attempt to insert, but update if duplicate key exists).
Writes records to a database table using 1 or more connections; each connection writing asynchronously using a separate thread.
Reads a pipeline's source data from a database query using a
JdbcReader.Writes a pipeline's records to a database table using a
JdbcWriter.A named SQL query with its parameters and result columns;
JdbcConnection.loadQueries() fills in the columns by running it.Whether a query returns many rows or a single row; generated DAO code reads only the first row for
ONE.Metadata for one result column of a
JdbcQuery, read from JDBC ResultSetMetaData by
JdbcConnection.loadQueries().A
? parameter of a JdbcQuery: its name, Java and SQL types, and an expression giving the example value
bound when JdbcConnection.loadQueries() runs the query.Obtains records from a database query.
One page of query results: its rows plus the page number, page size and total row count.
A database schema and the catalog it belongs to, as loaded by
JdbcConnection.loadCatalogAndSchemas().A named database destination, such as a
JdbcTable.A named table or query whose rows can be read with the SQL from
JdbcSource.getQuery().Whether a
JdbcSource is a table or a query.Metadata for a database table or view (name, schema, type, columns, keys and indexes), as loaded by
JdbcConnection.loadTables(LoadTablesRequest).Metadata for a column of a
JdbcTable, read from JDBC DatabaseMetaData by
JdbcConnection.loadTables(LoadTablesRequest).Writes records to a database table using the upsert idiom (attempt to insert, but update if duplicate key exists).
The interface used to read column values from a JDBC ResultSet.
Writes records to a database table.
RESTEasy client proxy for the Jira REST endpoints used by the Jira readers and
JiraService; methods return the raw
JSON response.Holds a RESTEasy client and the
JiraClient proxy created by JiraClient.Proxy.get(String, String, String);
call JiraClient.JiraStub.close() to release the client.Creates Jira clients that authenticate with a username and API key.
Reads Jira epics from a board based on the API user's permission.
The request body for creating, updating or transitioning a Jira issue: its key, its
fields record and an optional transition.Reads Jira issues using JQL.
Reads Jira projects from based on the API user's permission.
The fields a
JiraProjectReader can order projects by.The direction a
JiraProjectReader sorts projects in; DEFAULT adds no direction prefix.Base class for readers that page through a Jira REST resource, returning each item of every page as a record.
A Jira issue search request: the JQL query, page size, fields to return and the token of the next page.
Creates, updates, searches, transitions, comments on and deletes Jira issues through
JiraClient, returning JSON
responses as records.Reads Jira sprints from a board based on the API user's permission.
The sprint states a
JiraSprintReader can filter on; ALL sends no state filter.The JMS session acknowledge modes for
JmsSettings; in CLIENT mode, JmsReader acknowledges
each message once it has been converted to a record.A connection to Java Message Service (JMS) provider.
Constants for the type of messaging to use.
Read
Records from a Java Message Service (JMS) provider.Write
Records to a Java Message Service (JMS) provider.Used to run, manage, and track pipelines.
Receives the progress and outcome of a
Job; the default onFailure rethrows the exception.Factory for default callbacks such as
JobCallback.NULL.Receives start, pause, resume and finish notifications for jobs; register with
Job.addJobLifecycleListener(JobLifecycleListener).JobTemplate is the base type for classes that transfer records from DataReaders to DataWriters.
The default
JobTemplate, which runs each transfer as a Job.Base class for the JOIN clauses of a
Select, each joining a table on a condition.Vanilla JSON parser
Uses classes JsonParser + JsonHandler from:
https://github.com/ralfstx/minimal-json
A JSON array in a
JsonTemplate; elements are added with JsonArray.object(), JsonArray.array() and
JsonArray.value(String).A lightweight
List adapter over a DataPipeline ArrayValue.JFunction definition class
JFunction callable Lambda interface
A fragment of a
JsonTemplate whose nodes are written only when its logical expression is true; created by
when(String).Marks where records go in a
JsonTemplate; the nodes added to it are written once for each record.An expression for a field name or value in a
JsonTemplate, parsed once and evaluated on each write.A field of a JSON object in a
JsonTemplate: an expression for its name plus a JsonValue.Base class for template nodes that write only their children, not JSON of their own, and take their type from
the parent node.
A handler for parser events.
Writes records in JSON Lines format: each record as a JSON object on its own line.
A boolean expression that decides whether a
JsonConditionalNode's nodes are written.Base class for the nodes of a
JsonTemplate; each writes its opening and closing JSON to a Jackson generator.The kinds of nodes in a
JsonTemplate: scalar values, arrays, objects and object fields.Walks a
JsonTemplate's node tree depth-first, writing one node to a Jackson generator per step.A JSON object in a
JsonTemplate; fields are added with JsonObject.field(String, String), JsonObject.object(String)
and JsonObject.array(String).A streaming parser for JSON text.
Reads a pipeline's source data from a JSON
FileSource using a JsonReader,
which selects fields and record breaks by location path.A scalar value in a
JsonTemplate computed from an expression; dates, times and datetimes are written as
yyyy-MM-dd, HH:mm:ss.SSS and yyyy-MM-dd'T'HH:mm:ss.SSS'Z' strings.Obtains records from a JSON stream using
XmlReader location paths, where JSON objects and arrays appear
as object and array elements (for example //array/object/name).Reads a pipeline's source data from a JSON
FileSource using a JsonRecordReader,
which turns each node matching a record break into one nested record.Writes a pipeline's records to a JSON
FileSink using a JsonRecordWriter, one object per record.Obtains records from a JSON stream, turning each object, array or field matched by a record break into one record
that keeps nested objects and arrays as nested records and arrays.
Writes JSON using the structure of each record's natural representation.
The interface implemented by classes to participate in serialization to and from JSON.
Describes the JSON a
JsonWriter produces: nodes before the JsonFragment.detail() marker form the header, the
marker's nodes are written once per record and the remaining nodes form the footer.Base class for the objects, arrays and scalar values of a
JsonTemplate.Writes records as JSON using a
JsonTemplate that defines the header, each record's detail and the footer.Read
Records from an Apache Kafka distributed messaging system.Write
Records to an Apache Kafka distributed messaging system.Strategy for text values longer than
ExcelWriter.MAX_CELL_LENGTH characters, the most an Excel cell can hold.A pipeline output that writes records as a table in a LaTeX document to a
FileSink using a LatexWriter.Writes records to a standalone LaTeX document containing a table.
Generates a LEFT JOIN clause.
A proxy that can skip a number of upstream records, limit the number of records sent downstream, or both.
Abstract super-class for writing records to a text stream.
An
IParser that reads its input one line at a time; lookahead and matching only see the current line until
LineParser.cacheNextLine() moves to the next one.Selects which lists
TrelloBoardListsReader reads; toString() gives the lowercase value sent to Trello.A member of a MailChimp list, as read from or sent to the API's list member endpoints.
A page of members returned by the MailChimp list members endpoint.
Statuses of a MailChimp
ListMember, also used to filter the members requested from a list.The options for
JdbcConnection.loadTables(LoadTablesRequest): which tables to load and which of their details
(columns, keys, indexes) to load with them.Base class for the
FileSource and FileSink implementations backed by a file on the local file system.Caches the dataset's records on disk as binary data.
A
FileSink that writes a file on the local file system, replacing or appending to its content.A
FileSource that reads a file on the local file system.An immutable object that represents a location in the parsed text.
Location details of a MailChimp
ListMember.Searches for data in another source to merge with each record passing through this lookup's
transformer -- the streaming version of a join.
A
Transformer that merges into each record the fields of the record a Lookup returns for the values of
the key fields.Client for the MailChimp API 3.0 list member endpoints; create one with
MailChimpClient.Proxy.get(String).Creates
MailChimpClient proxies that authenticate with an API key.Reads the list of subscribed, unsubscribed, and cleaned members from a MailChimp list.
The mail protocols an
EmailReader can connect with, each with its standard port.Final naming pass: reassigns field names that still collide after the tree structure was
transformed, keeping every field name unique.
Copies the values in specific key-value fields to an in-memory
Map.Caches the dataset's records in RAM.
Obtains records from an in-memory
RecordList.Writes records to an in-memory
RecordList.Builds a MERGE statement that updates the target row when the ON condition matches and inserts one otherwise;
the target table is aliased
target.A batch-able upsert strategy that relies on the SQL Merge statement.
The standard
IMergeUsing clauses, which supply the source row as parameters:
USING (VALUES (?, ...)) source(...) or USING (SELECT ?, ...) source(...).A column in the INSERT or UPDATE SET clause of a
Merge, with its parameter value.A diagnostic message with a severity level, an optional exception or stack trace, and the time and thread it was created
on; collected by
Messages.The severity of a
Message.Thread-local container for messages and exceptions originating in asynchronous operations.
Counts units (such as bytes or records) and reports their average rate per second since the meter was created or last reset.
Implemented by streams, readers and writers that measure the data passing through them with a
Meter.An input stream that counts the bytes read or skipped in a
Meter; MeteredInputStream.reset() also resets the meter.An output stream that counts the bytes written in a
Meter.A proxy that measures the rate (bytes/second) at which data is read.
A proxy that measures the rate (bytes/second) at which data is written.
A pipeline output that writes records as a table in a Microsoft Word document to a
FileSink using a
MicrosoftWordWriter.Writes records to a Microsoft Word (OOXML
.docx) document stream containing a single table.The page orientations available to a Word document.
The page sizes available to a Word document, with their portrait dimensions in twips (1/1440 of an inch).
Reads documents from a MongoDB database collection.
Converts DataPipeline records and values to and from MongoDB BSON documents and values.
Writes records to a MongoDB database collection.
See
FieldPath for supported field name expressions.See
FieldPath for supported field name expressions.A transformer used to move a named field to a new position.
Sample Generated SQL:
INSERT INTO tableName (col1, col2, col3) VALUES
(?, ?, ?),
(?, ?, ?),
(?, ?, ?),...Sample Generated SQL:
INSERT INTO tableName (col1, col2, col3) VALUES
(val1, val2, val3),
(val1, val2, val3),
(val1, val2, val3),..Writes records to multiple
DataWriter.Writes each record to a single writer by choosing the one with the
highest available capacity (
DataWriter.available() If all writers
have identical capacities, this strategy behaves like
MultiWriter.RoundRobinWriteStrategy and writes to each writer in turn.Writes a clone (to prevent side effects) of each record to all writers.
Writes each record to all writers.
Divides records evenly between all writers by cycling through the list
and writing each record to a single writer in turn.
Strategy for distributing each record among a
MultiWriter's target writers.Caches the dataset's records on disk using MVStore.
Generates MySQL insert statement from records.
Base class for the MySQL statement builders; identifiers are quoted with backticks.
Sample Generated SQL:
INSERT IGNORE INTO tableName (col1, col2, col3) VALUES (?, ?, ?)A batch-able upsert strategy where when a row to be inserted would cause a duplicate value in a UNIQUE index or PRIMARY KEY, an UPDATE of the old row occurs.
Generates MySQL upsert statement from records.
Extracts each run/sequence of N words from the source field and stores them in a target field as an array.
Node is the base class for all persistent data in Data Pipeline.
DuplicateNodeAction lists all actions that can be taken automatically when a target field already exists during a record copy or lookup.
NodeType lists all concrete data node types used in Data Pipeline.
Classifies a single XML element or JSON field name as dynamic data (an identifier, date, or other
generated value) rather than a fixed structural name.
Base
INodeVisitor whose callbacks do nothing; override only the ones you need and walk a tree with NodeVisitor.visit(Node, INodeVisitor).Matches only when none of a set of filters match -- !(filter A || filter B || filter C).
Negates another criteria part, rendering it as
not (...).A
FieldFilterRule that inverts another rule, allowing the fields it rejects.Matches every event; used when a listener is added without a filter.
A pipeline with no processing whose input defaults to a
NullReader (no records) and output to a NullWriter.A data source that produces no records.
Discards records.
A utility for parsing simple numeric values from string.
The result of
NumberDetector.match(String): the matched text split into whole and fraction parts, plus a
NumberDescriptor for it.A
FieldDef for numeric values with optional range, precision and scale limits and a parsing pattern for mapping.Builds the
UPDATE SET column=?, ... action that follows ON CONFLICT (...) DO in an upsert statement.A pipeline input that reads Open Financial Exchange (OFX) data from a
FileSource using an
OpenFinancialExchangeReader.Read an OFX / QFX (Quicken) / QBO (QuickBooks) file as a single nested
Record that
mirrors the document's aggregate hierarchy, using the OFX tag names (BANKMSGSRSV1,
STMTTRN, TRNAMT, ...) as field names.Cleans up a loaded
Tree by merging non-wildcard siblings into an existing wildcard node,
adding a synthetic value field under childless wildcard nodes, and marking JSON array elements
as record breaks.Sample Generated SQL:
OracleMultiRowInsertAllStatementInsert
OracleMultiRowInsertAllStatementInsert
INSERT ALL
INTO tableName (col1, col2, col3) VALUES ('val1', 'val2', 'val3')
INTO tableName (col1, col2, col3) VALUES ('val1', 'val2', 'val3')
INTO tableName (col1, col2, col3) VALUES ('val1', 'val2', 'val3')
...
SELECT 1 FROM dualSample Generated SQL:
INSERT INTO tableName (col1, col2, col3)
SELECT 'val1', 'val2', 'val3' FROM dual UNION ALL
SELECT 'val1', 'val2', 'val3' FROM dual UNION ALL
...
SELECT 'val1', 'val2', 'val3' FROM dualBase class for the Oracle-specific multi-row insert strategies, which write the values into the SQL as literals.
A batch-able upsert strategy that relies on the SQL Merge statement.
Read records from Apache ORC columnar files.
Writes records to Apache ORC columnar files.
A pipeline input that reads the Apache ORC file at a
FileSource's path using an OrcDataReader.A pipeline output that writes an Apache ORC file to a
FileSink's path using an OrcDataWriter.Order statuses for
ShopifyOrder.setStatus(OrderStatus); OrderStatus.toString() returns the lower-case name.Matches any one of a set of filers -- filer A || filer B || filer C.
A
ConjunctivePart that joins its parts with OR.A
FieldFilterRule that allows a field when any of its rules does; with no rules it rejects every field.Read records from Apache Parquet columnar files.
Writes records to Apache Parquet columnar files.
A Parquet
that writes to a Hadoop
, with checksum writing turned off on checksum file systems.
invalid reference
OutputFile
invalid reference
Path
A pipeline input that reads the Apache Parquet file at a
FileSource's path using a ParquetDataReader.A pipeline output that writes an Apache Parquet file to a
FileSink's path using a ParquetDataWriter.Utility methods to read Parquet schemas and to rewrite schema fields identified by column path.
Thrown when expression-language text cannot be parsed.
An unchecked exception to indicate that an input does not qualify as valid JSON.
Abstract super-class for obtaining records from a text stream using a parser.
A
FieldFilterRule that allows fields whose whole value, as a string, matches any of its regular expressions;
null values are rejected.A PDF file or stream whose tables are extracted from every page when it is opened, for reading with a
PdfReader.Page, image and text statistics of a
PdfDocument.A pipeline input that reads a table from a PDF in a
FileSource using a PdfReader.Obtains records from the rows of one table extracted from a
PdfDocument, with one field per cell.Writes records to a PDF document stream.
A PDF page size, expressed as a CSS
@page size value.The page orientations, each mapped to its CSS
@page orientation keyword.Common page sizes, each mapped to its CSS
@page size keyword.Represents a fixed, cyclical schedule (for example: every 30 minutes starting now).
Reads the records written to a connected
PipedWriter, usually on another thread; reads block until a record arrives
and the stream ends when the writer is closed.Writes records to a connected
PipedReader, blocking while its queue is full; closing this writer ends the reader's stream.A declarative pipeline that applies an ordered list of
PipelineActions to its input records, with undo and redo of
the most recent actions.Base class for the steps of a declarative
Pipeline; each action wraps the reader produced by the previous step.Base class for the sources a pipeline reads its records from, such as CSV, Excel, fixed-width, JSON or XML files, JDBC
queries and datasets.
Base class for pipelines and their parts (inputs, outputs and actions), which can be serialized and can generate Java code.
Base class for the destinations a pipeline writes its records to, such as CSV, Excel, fixed-width, JSON or XML files and
JDBC tables.
Provider for
ExcelDocument.ProviderType.POI_SXSSF: writes .xlsx files with Apache POI's streaming SXSSF API,
keeping the last 1000 rows in memory and the rest in compressed temporary files.Read-only provider for
ExcelDocument.ProviderType.POI_XSSFB: reads binary .xlsb sheets from a File,
returning each cell as its formatted string.Provider for
ExcelDocument.ProviderType.POI_XSSF_SAX: opens .xlsx files read-only and only from a
File; new workbooks are written as with ExcelDocument.ProviderType.POI_XSSF.Generates PostgreSQL insert statement from records.
A batch-able upsert strategy that relies on the SQL Merge statement.
The action taken when an insert conflicts on the key fields: update the existing row or do nothing.
Generates PostgreSQL upsert statement from records.
Base class for the PostgreSQL statement builders; identifiers are quoted with double quotes.
Sample Generated SQL:
INSERT INTO tableName (col1, col2, col3) VALUES (?, ?, ?)See the definitions in JPA's javax.persistence.GenerationType
Product statuses for
ShopifyProductCriteria.setStatus(ProductStatus...); ProductStatus.toString() returns the lower-case name.A name-value pair of strings, such as a JDBC driver property of a
JdbcConnection.A leaf
Diff for a single value, such as a record field or an array element.A
FileSource that delegates its file methods and record serialization to a nested file source; subclasses
override selected methods.Abstract super-class for obtaining records from another
Lookup,
possibly transforming them along the way.Base class for inputs that wrap another
PipelineInput and decorate each reader it creates.Base class for outputs that wrap another
PipelineOutput and decorate each writer it creates.Abstract super-class for obtaining records from another
DataReader,
possibly transforming them along the way.Abstract super-class for writing records to another
DataWriter,
possibly transforming them along the way.Product published states for
ShopifyProductCriteria.setPublishedStatus(PublishedStatus); PublishedStatus.toString() returns the lower-case name.The WHERE or HAVING clause of a query, built from conditions combined with AND and OR; renders nothing when empty.
Whether a
QueryCriteria renders as a WHERE or a HAVING clause.Base class for the nodes of a
QueryCriteria condition tree.The GROUP BY clause of a
Select; renders nothing when it has no fields.The ORDER BY clause of a
Select; renders nothing when it has no fields.One ORDER BY entry: a field, its direction and optional null placement.
Base class for the clauses and conditions that make up a
Select query.The SELECT column list of a
Select; selects * when no fields are added.A Parquet
backed by a
invalid reference
InputFile
RandomAccessFile; all streams it opens share the file and its
read position, and closing them does not close the file.A Parquet
that reads from a
invalid reference
SeekableInputStream
RandomAccessFile; closing it does not close the file.Record holds persistent data in key-value fields as it flows through the pipeline.
Lists the lifecycle states of a record as it flows through the pipeline.
Collects the records that share one key in a
BufferedReader; they can be read back only after the buffer is closed.Orders records by a list of fields, each ascending or descending with an optional
Collator; when no fields are
added, records are compared on all their fields by name.A general interface for objects that are associated with a
Record.A
Diff between two records; its children describe the differences in their fields.A
FieldDef for nested records, validated and mapped using another entity in the same schema.Jackson component annotated on Record to load it from JSON.
Jackson component annotated on Record to save it to JSON.
Convenience wrapper to read and write record-level data lineage properties.
An in-memory list of records that can be sorted, searched with a
Filter, and converted to binary, arrays or a reader.Listener interface for publishing and receiving records on an
EventBus, as EventBusWriter and
EventBusReader do.An in-memory lookup implementation that performs joins using a
RecordList.A
Meter that counts records or their estimated size in bytes.The unit a
RecordMeter counts in: bytes (each record's estimated size) or records.The interface implemented by classes to participate in serialization to and from Record.
Exposes a
Record to FreeMarker templates as a hash keyed by field name, with values wrapped by the object wrapper.Classifies names matching a caller-supplied regular expression as dynamic.
A
ZipEntryFilter that accepts entries whose entire name, including any folder path, matches a
regular expression.Whether one primary entity record relates to one or to many foreign entity records in an
EntityRelationshipDef.Removes or collapses multiple fields with the same name in each record.
Built-in policies that keep all duplicates, keep only the first or last, merge their values into an array in the
first field, or rename the later ones with a numeric suffix.
Strategy for resolving fields that share a name in a record, used by
RemoveDuplicateFields.An alternative to
RemoveDuplicateFields.DuplicateFieldsPolicy.RENAME that uses a configurable separator between
the original field name and the numeric suffix, preventing ambiguous names like 20101 when
the field name itself ends in digits (e.g.A
PipelineAction that drops records whose values in the selected fields (all fields if none are selected) match an
earlier record.A proxy that removes duplicate records.
See
FieldPath for supported field name expressions.See
FieldPath for supported field name expressions.A
PipelineAction that renames fields; mappings with a blank target or the same name as the source are ignored.A rename of one source field to a target name.
Describes a failed attempt of a
RetryingOperation; passed to its retry predicate to decide whether to retry.A proxy that attempts to re-run on failure.
A proxy that attempts to continue reading on failure.
A proxy that attempts to continue writing on failure.
Standard
IRetryStrategy backoff policies; every policy waits the initial delay before the first retry.Generates a right join clause.
Rounds numbers to a number of decimal places using a
Rounder.RoundingPolicy.Rounding modes supported by
Rounder.Writes records to an RTF document stream using Apache POI.
Runs a
Runnable as an AbstractJob; pausing has no effect and cancelling only interrupts its thread.A backend-neutral description of an S3 bucket, returned by
AmazonS3FileSystem.listBuckets().A backend-neutral, single page of an S3 object listing, returned by
AmazonS3FileSystem.listRootFolder(String),
AmazonS3FileSystem.listFolder(String, String) and
AmazonS3FileSystem.nextBatch(S3ObjectListing).A backend-neutral summary of a single object returned when listing an S3 bucket or folder.
Determines the next scheduled event time from a given point in time.
The metadata model containing a set of related entities.
The differences between two versions of a
SchemaDef, with child diffs for its properties, entities and
relationships.A filter that matches a record against an entity definition.
Formats the error messages for invalid schema definitions and for values and records that fail schema validation or mapping.
Base class for the schema model (
SchemaParts, SchemaProblems and validation results), adding Java code
generation to FoundationObject.Base class for the elements of a schema definition -- the
SchemaDef itself and its entities, relationships,
fields, indexes and index fields.The kinds of
SchemaPart: schema, entity, entity relationship, field, index and index field.Describes an issue found when checking a schema definition and the
SchemaPart it was found in.Severity levels for schema problems; currently unused.
Implemented by schema elements that can check themselves, and optionally their children, for
SchemaProblems.Converts and validates records using the specified entity schema.
Keeps a per-thread stack of the names being validated, used to build qualified field names such as
order.items.[0] for validation messages.Static checks used by the schema definitions to add a
SchemaProblem to a list for each issue found.Builds a SELECT statement with its parameter values; as a
QuerySource it can also be nested in another query.A
PipelineAction that keeps only the selected fields, in the order they were added, like an SQL SELECT.Select and arrange fields according to the supplied field names -- like an SQL SELECT statement.
Combines one or more DataReaders into a single stream by reading from each until empty then moving to the next.
Writes to sequence of data writers created by a factory in turn and rolled based on a ISequenceStrategy policy.
Creates the nested writer for each new sequence of a
SequenceWriter.Ends a sequence once it has been running for a given time; checked only when the next record arrives.
Decides when a
SequenceWriter ends the current sequence and starts a new one.Ends the current sequence when any of the given strategies does.
Ends each sequence once it has received the given number of records.
Ends the current sequence at each time produced by a
Scheduler; checked only when the next record arrives.One segment of a
SequenceWriter's output, with its 0-based index, start time and record counts.Session is the base type for objects that can contain temporary, non-persistent data in DataPipeline.
Stores session properties in a lazily created map, keyed by property name, class name, or both.
Adds a sequence (auto increment) field that is incremented with the step value when the value(s) in the specified
watch fields differs from the previous record.
See
FieldPath for supported field name expressions.A
Transformer that sets a field on every record to a fixed value, creating the field if needed; when
overwrite is false, records that already have the field are left unchanged.Adds a sequence (auto increment) field that is incremented with the step value as long as the value(s) in the specified
watch fields remain the same as the previous record.
Adds a sequence (auto increment) field with the specified increment (step).
Adds a field with a randomly generated UUID as its value.
Class for accessing the SFTP file system.
Base class for readers that page through a Shopify Admin API resource using
page_info cursors and retry
requests that fail with a 429 (rate limit) error.JAX-RS client for the Shopify Admin REST API (version 2023-07); each method returns the JSON response body.
Request filter that sends a token in the
X-Shopify-Access-Token header.Response filter that copies the
page_info cursor of the Link header's next-page URL into a
next field of JSON responses with status 200.Factory for
ShopifyClient.ShopifyStub instances.Holds a
ShopifyClient proxy with the RESTEasy client and web target behind it.Filters for listing Shopify customers, as used by
ShopifyCustomerReader; properties left null are not sent.Reads customers from a Shopify store, optionally filtered by a
ShopifyCustomerCriteria.Filters for listing Shopify events, as used by
ShopifyEventReader; properties left null are not sent.Reads events from a Shopify store, optionally filtered by a
ShopifyEventCriteria.The inventory item fields sent by
ShopifyClient.updateInventoryItem(String, ShopifyInventoryItem); properties left null are not sent.Filters for listing Shopify inventory items, as used by
ShopifyInventoryItemReader; properties left null are not sent.Reads inventory items from a Shopify store, optionally filtered by a
ShopifyInventoryItemCriteria.Filters for listing Shopify inventory levels, as used by
ShopifyInventoryLevelReader; properties left null are not sent.Reads inventory levels from a Shopify store, optionally filtered by a
ShopifyInventoryLevelCriteria.Reads the locations of a Shopify store.
Filters for listing Shopify orders, as used by
ShopifyOrderReader; properties left null are not sent.Reads orders from a Shopify store, optionally filtered by
ShopifyOrder criteria.Filters for listing Shopify products, as used by
ShopifyProductReader; properties left null are not sent.Reads products from a Shopify store, optionally filtered by a
ShopifyProductCriteria.Calls the Shopify Admin REST API and returns each JSON response as a
Record, rethrowing failures as DataExceptions.Shipping address of the order sent by
ShopifyClient.updateOrder(String, ShopifyOrder); properties left null are not sent.Notified when an
EventBus begins shutting down; register with
EventBus.addShutdownListener(EventFilter, ShutdownListener).An immutable naming pattern derived from a set of sibling names — a fixed prefix and suffix
around a dynamic middle segment (for example
item-* derived from
item-1, item-2, item-3).Examines a parent's candidate children collectively and derives a
SiblingSetPattern when
their names share a naming/wildcard pattern, or returns null when no pattern can be
derived.Manages signature related functions
Reads a JSON stream created by
SimpleJsonWriter.Writes JSON in the following simple format.
Reads an XML stream created by
SimpleXmlWriter.Writes XML in the following simple format.
SingleValue is an immutable
ValueNode that holds a single scalar value (or null).A
PipelineAction that sorts the records by one or more fields, each in ascending or descending order.A proxy that sorts records.
Flattens an array field by creating one record for each value.
Used by SplitWriter as a consumer when converting a single source DataReader into many downstream sources.
Flattens an array field by creating one record for each value.
Converts a single source DataReader into many downstream sources using one of the provided strategies.
Strategy for distributing the records written to a
SplitWriter among the readers it created.The built-in strategies: send every record to all readers or each record to one reader in turn.
Base class for the SQL builders: renders a SQL fragment and collects the values for its
? parameters.Interface for classes that can generate SQL SELECT statements.
The default
.
NodeNameRule: treats numeric names, UUID names, and names containing special
characters as dynamic, matching
invalid reference
XmlNode#isWildcardCandidate(String)
A
CredentialsResolver that always returns the same fixed Credentials.Writes records to a stream in a human-readable format.
How a
StreamWriter prints each record (after its 0-based number): plain text, JSON, XML, or text plus session properties.An
IParser over an in-memory string.Derives a
SiblingSetPattern from siblings whose child structures overlap, instead
of their names: each candidate's structural signature is the set of its child element and
attribute names (looking through JSON object/array container nodes),
and siblings whose signatures agree above minimumSimilarity
(Jaccard overlap) form the collapse group.A
CredentialsResolver that delegates each resolve to a Supplier.A batch-able upsert strategy that relies on the SQL Merge statement.
A
DecryptingReader that decrypts fields written by a SymmetricEncryptingReader, restoring each field's
original type and value.An
EncryptingReader that encrypts fields with a symmetric key, storing each as Base64 text from which a
SymmetricDecryptingReader restores the original type and value.A
CredentialsResolver that reads credential values from JVM system properties
at each resolve.Represents a table column definition for H2.
A table column's name and
TableColumnType.Represents data types for H2 table columns.
MySQL column data types for
CreateTableColumn, noting which accept a length, a scale or AUTO_INCREMENT.PostgreSQL column data types for
CreateTableColumn, noting which accept a length and a scale.A single table in the FROM clause of a
Select.A tag assigned to a MailChimp
ListMember.A sorted set of unique labels used to organize, filter and classify schema, entity, relationship and field definitions.
Keeps the most recent n records in an in-memory
RecordList which can be retrieved
asynchronously while data is flowing through.Writes records to an in-memory
RecordList.A proxy reader that also writes every record passing through it to a DataWriter.
A FreeMarker
DefaultObjectWrapper that also wraps Records as RecordTemplateModels and
ArrayValues as ArrayTemplateModels.Writes records to a text stream using FreeMarker template.
A
FieldDef for date, time and datetime values with optional minimum and maximum limits and a parsing pattern
for mapping.A
FieldDef for text values with optional length limits, blank check, regular expression and trimming.Abstract super-class for obtaining records from a text stream.
Abstract super-class for writing records to a text stream.
Writes records as an easy-to-read, plain-text table with padded, aligned columns — a header row
of field names, a dashed separator line, then one row per record:
Abstract super-class for writing records to a text stream.
Limits the average rate measured by a
Meter by making the calling thread sleep while the count is ahead of the allowed rate.Implemented by streams, readers and writers whose rate is limited by a
Throttle.A
MeteredInputStream that pauses after reads when needed to keep its average rate at or below a set number of bytes per second.A
MeteredOutputStream that pauses before writes when needed to keep its average rate at or below a set number of bytes per second.A proxy that limits the rate (bytes/second) at which data is read.
A proxy that limits the rate (bytes/second) at which data is written.
Configure max runtime / max recursion depth.
Reads an underlying DataReader for a maximum period of time or until the source DataReader is finished.
Matches any one of a set of topics -- topic A || topic B || topic C.
Base class for transformations that modify each record passing through a
TransformingReader or TransformingWriter.A proxy that applies transformations to records passing through.
A proxy that applies transformations to records passing through.
Summarizes the structure of an XML or JSON document as a tree of
TreeNodes, one per distinct path, and detects which
nodes are record breaks and fields.Defines the algorithms used to detect record breaks and fields in a
Tree: an ordered list
of NodeNameRules that classify dynamic element/field names while the tree is being built,
followed by an ordered list of TreeTransformer passes that run over the loaded tree.One distinct element, attribute or JSON path in a
Tree, with its instance count, value statistics and record-break and
field flags.A single pass in a
TreeDetectionStrategy that analyzes or restructures a Tree
after it has been loaded from an XML or JSON document.Superclass for Trello readers supporting pagination.
Superclass for Trello writers.
Reads the list of Trello cards belonging to a board.
Reads the list of Trello lists belonging to a board.
Reads the list of Trello boards belonging to a member (or the current member if non are specified).
Writes each record to Trello as a new card in a list, reading the list ID, card name and description from configurable
record fields that must all be present.
Reads the list of Trello cards belonging to a list.
Default
ITwitterConverter, available as TwitterConverter.DEFAULT, that maps tweets, users and lists to record fields.The API key and secret plus an OAuth 1.0a access token or a bearer token, used to authorize Twitter API v2 requests.
Known Twitter API error codes, each with its error message and a longer description.
Continuously reads tweets matching several criteria (hashtag, user, language, etc.)
Continuously reads filtered Tweets in real-time based on a set of filter rules.
Reads the IDs of accounts following the specified account.
Reads the details of accounts following the specified account.
Obtains records of Twitter users that follow the specified Twitter account.
Reads the IDs of accounts a specified account follows.
Reads the details of accounts a specified account follows.
Obtains records of Twitter users that a specified Twitter account is following.
Follows accounts written to it.
Reads the users of list.
Reads the list for the specified account.
Base class for Twitter readers that page through cursor-based API results, returning one record per element.
Paging direction: follow each page's next cursor (
FORWARD) or its previous cursor (BACKWARD).Non-streaming Twitter service provider.
Base class for all Twitter ProxyReaders.
A snapshot of a Twitter API rate limit: the calls allowed, the calls remaining and when the limit resets.
Base class for all Twitter DataReaders.
Continuously reads a small random sample of all tweets.
Obtains records by searching Twitter for the most recent tweets matching a search criteria.
Obtains records of Tweets posted over the last 7 days that match a search criteria.
Streaming Twitter service provider.
Obtains records of Tweets mentioning the specified Twitter account.
Base class for the Twitter API v2 readers of a user's timeline, optionally limited by date and tweet ID range.
Obtains records of Tweets, Retweets, replies, and Quote Tweets published by a specific Twitter account.
Unfollows accounts written to it.
Adds Twitter user details to records from a nested reader by looking up the user ID or screen name in a source field.
Minimal Twitter API v2 client (tweet search, followers, timelines, filtered stream) used by this package's readers.
Base class for all Twitter DataWriters.
Holds the listeners subscribed to an
EventBus through one listener interface.Combines several
Select queries with UNION or UNION ALL; can be nested in another query as a source.Receives events published on an
EventBus whatever their listener interface; register with
EventBus.addUntypedEventListener(UntypedEventListener).Builds an
UPDATE table SET column=?, ... WHERE ... statement with its parameter values.A
column=? assignment and its parameter value in an UPDATE SET list.Generates an upsert statement for H2.
Generates an upsert statement for MySQL.
Generates an upsert statement for PostgreSQL.
List representing an int range [a,b]
Both sides are included.
A
PipelineAction whose generated Java code checks that records contain the selected fields;
ValidateFieldsExistsAction.apply(DataReader) returns the reader unchanged.A
PipelineAction whose generated Java code checks that fields fully match their regular expressions;
ValidateFieldsMatchPatternAction.apply(DataReader) returns the reader unchanged.A
PipelineAction whose generated Java code checks that the selected fields are not null;
ValidateFieldsNotNullAction.apply(DataReader) returns the reader unchanged.A
PipelineAction whose generated Java code checks that records satisfy a boolean expression;
ValidateMatchExpression.apply(DataReader) returns the reader unchanged.A proxy that validates records by applying a set of filters.
One mapping or validation error, with the entity, field or validation rule it concerns.
Collects the
ValidationMessage errors produced when mapping or validating data against a schema definition.A
FieldFilterRule that allows fields whose value equals one of its values; values are compared with
equals, so their Java types must match (for example 1 and 1L differ).ValueNode is the base class for all persistent values held in fields and arrays.
A comparator for
ValueNode instances that supports sorting and comparing records, arrays, and single values.A general interface for nodes that allow values to be added to itself.
Converts values between DataPipeline
ValueNodes and the maps, lists and plain Java values used by the JSONata engine.A upsert strategy that wraps another strategy to support a variable set of fields.
Creates the strategy used for each distinct sequence of field names.
Pairs a strategy with the number of records it has upserted.
A
ZipEntryFilter that accepts entries whose name, including any folder path, matches a wildcard pattern
(? for one character, * for zero or more).Windows accept detail data for aggregation while open and only yield its summary after being closed.
An attribute of an
XmlElement whose name and value are expressions, optionally written only when a
condition is true.A group of template nodes written only when its logical expression is true; created by
XmlNodeContainer.when(String).Marks where records go in an
XmlTemplate; the nodes added to it are written once for each record.An element in an
XmlTemplate whose name is an expression; holds attributes and child nodes.An expression for an element name, attribute or text in an
XmlTemplate, written XML-escaped.Maps the values found at a location path to a named record field for
XmlReader, optionally cascading the
last value into later records.A group of template nodes that writes only its children, with no markup of its own.
A boolean expression that decides whether conditional template nodes or attributes are written.
Base class for the nodes of an
XmlTemplate; each writes its opening and closing markup to a CodeWriter.Base class for template nodes with children; provides the
XmlNodeContainer.element(String), XmlNodeContainer.text(String),
XmlNodeContainer.when(String) and XmlNodeContainer.detail() builder methods.Reads a pipeline's source data from an XML
FileSource using an XmlReader,
which selects fields and record breaks by location path (a subset of XPath).Obtains records from an XML stream.
Strategy for a field whose location path matches more than once within the same record.
A location path, such as
//book, identifying the elements that mark records for the XML and JSON readers.Reads a pipeline's source data from an XML
FileSource using an XmlRecordReader,
which turns each element matching a record break into one nested record.Writes a pipeline's records as XML elements to a
FileSink using an XmlRecordWriter.Obtains records from an XML stream, turning each element matched by a record break into a record of its child
elements (as nested records) and attributes (as
@name fields).Writes XML using the structure of each record's natural representation.
The interface implemented by classes to participate in serialization to and from XML.
Describes the XML an
XmlWriter produces: nodes before the XmlNodeContainer.detail() marker form the header, the
marker's nodes are written once per record and the remaining nodes form the footer.Text in an
XmlTemplate computed from an expression and written XML-escaped.Writes records to an XML stream using a template.
Creates a
DataReader for an entry of a zip file.Decides which entries of a zip file
listEntries(ZipEntryFilter) returns.Class for accessing zips as file systems.
GroupByReaderinstead.