Class ParquetPipelineInput
java.lang.Object
com.northconcepts.datapipeline.foundations.core.Bean
com.northconcepts.datapipeline.foundations.core.FoundationObject
com.northconcepts.datapipeline.foundations.pipeline.PipelineObject
com.northconcepts.datapipeline.foundations.pipeline.PipelineInput
com.northconcepts.datapipeline.foundations.pipeline.input.FileSourcePipelineInput
com.northconcepts.datapipeline.parquet.input.ParquetPipelineInput
- All Implemented Interfaces:
DataExceptionContributor,JsonSerializable,RecordSerializable,XmlSerializable,DataReaderFactory,JavaCodeGenerator,Serializable
A pipeline input that reads the Apache Parquet file at a
FileSource's path using a ParquetDataReader.- See Also:
-
Field Summary
Fields inherited from class com.northconcepts.datapipeline.foundations.core.FoundationObject
internalId, internalName, log, TIMESTAMP_FORMATFields inherited from interface com.northconcepts.datapipeline.core.RecordSerializable
SERIALIZED_CLASS_NAME, TYPEFields inherited from interface com.northconcepts.datapipeline.core.XmlSerializable
XML_SERIALIZED_CLASS_NAME -
Constructor Summary
Constructors -
Method Summary
Modifier and TypeMethodDescriptionfromRecord(Record source) Loads this instance's state from a record and returnsthis(for fluid API call chaining).fromXmlElement(Element pipelineInputElement) voidAppends Java source code representing this object to the given builder.Returns the Parquet configuration parameters.MetadataFilterReturns the metadata filter the reader applies when it opens the file (default isNO_FILTER).getName()Returns a display name for this input, such as the source's name followed by its format.MessageTypeIndicates the schema used to read the file.booleanisDebug()Indicates if debugging is enabled to print log statements.booleanIndicates if this input can save data lineage on the records it reads (false unless a subclass overrides it).booleanIndicates if top-level optional columns with no nulls and equal min/max statistics may be read as required (default is false).booleanIndicates if top-level required columns whose statistics report nulls are read as optional (default is false).booleanIndicates if schema fields with no column metadata in any row group are dropped from the schema used for reading (default is false).booleanIndicates if top-level columns that have no values in a row group are dropped from the schema used for reading (default is false).setConfiguration(String key, String value) Sets a single Parquet configuration parameter, keeping the others.setConfiguration(Map<String, String> configuration) Sets the Parquet configuration parameters.setDebug(boolean debug) Indicates if debugging is enabled to print log statements.setFileSource(FileSource fileSource) setFilter(MetadataFilter filter) Sets the metadata filter the reader applies when it opens the file (default isNO_FILTER).setMakeOptionalFieldsRequired(boolean makeOptionalFieldsRequired) Indicates if top-level optional columns with no nulls and equal min/max statistics may be read as required (default is false).setMakeRequiredFieldsOptional(boolean makeRequiredFieldsOptional) Indicates if top-level required columns whose statistics report nulls are read as optional (default is false).setRemoveFieldsWithoutColumnMetadata(boolean removeFieldsWithoutColumnMetadata) Indicates if schema fields with no column metadata in any row group are dropped from the schema used for reading (default is false).setRemoveFieldsWithoutValues(boolean removeFieldsWithoutValues) Indicates if top-level columns that have no values in a row group are dropped from the schema used for reading (default is false).setSchema(MessageType schema) Indicates the schema used to read the file.toRecord()Converts this object's state to a record thatRecordSerializable.fromRecord(Record)can load.toXmlElement(Document document) Returns an element, created withdocument, describing this object; the default implementation throws aDataException.Methods inherited from class com.northconcepts.datapipeline.foundations.pipeline.PipelineInput
getNestedPipelineInput, getPipelineInput, getRootPipelineInput, isSaveLineage, setSaveLineageMethods inherited from class com.northconcepts.datapipeline.foundations.core.FoundationObject
addExceptionProperties, assertValid, assertValid, clone, exception, exception, exception, getInternalId, getInternalName, resetInternalIdMethods inherited from class java.lang.Object
equals, finalize, getClass, hashCode, notify, notifyAll, wait, wait, waitMethods inherited from interface com.northconcepts.datapipeline.core.DataExceptionContributor
addExceptionProperties
-
Constructor Details
-
ParquetPipelineInput
public ParquetPipelineInput()
-
-
Method Details
-
getName
Description copied from class:PipelineInputReturns a display name for this input, such as the source's name followed by its format.- Overrides:
getNamein classFileSourcePipelineInput
-
getFileSource
- Specified by:
getFileSourcein classFileSourcePipelineInput
-
setFileSource
- Specified by:
setFileSourcein classFileSourcePipelineInput
-
createDataReader
-
generateJavaCode
Description copied from interface:JavaCodeGeneratorAppends Java source code representing this object to the given builder. -
isLineageSupported
public boolean isLineageSupported()Description copied from class:PipelineInputIndicates if this input can save data lineage on the records it reads (false unless a subclass overrides it).- Overrides:
isLineageSupportedin classPipelineInput
-
isDebug
public boolean isDebug()Indicates if debugging is enabled to print log statements. -
setDebug
Indicates if debugging is enabled to print log statements. -
getConfiguration
Returns the Parquet configuration parameters. -
setConfiguration
Sets the Parquet configuration parameters. -
setConfiguration
Sets a single Parquet configuration parameter, keeping the others. -
getFilter
public MetadataFilter getFilter()Returns the metadata filter the reader applies when it opens the file (default isNO_FILTER). -
setFilter
Sets the metadata filter the reader applies when it opens the file (default isNO_FILTER). -
getSchema
public MessageType getSchema()Indicates the schema used to read the file. -
setSchema
Indicates the schema used to read the file. -
isMakeRequiredFieldsOptional
public boolean isMakeRequiredFieldsOptional()Indicates if top-level required columns whose statistics report nulls are read as optional (default is false). -
setMakeRequiredFieldsOptional
Indicates if top-level required columns whose statistics report nulls are read as optional (default is false). -
isMakeOptionalFieldsRequired
public boolean isMakeOptionalFieldsRequired()Indicates if top-level optional columns with no nulls and equal min/max statistics may be read as required (default is false). -
setMakeOptionalFieldsRequired
Indicates if top-level optional columns with no nulls and equal min/max statistics may be read as required (default is false). -
isRemoveFieldsWithoutColumnMetadata
public boolean isRemoveFieldsWithoutColumnMetadata()Indicates if schema fields with no column metadata in any row group are dropped from the schema used for reading (default is false). -
setRemoveFieldsWithoutColumnMetadata
public ParquetPipelineInput setRemoveFieldsWithoutColumnMetadata(boolean removeFieldsWithoutColumnMetadata) Indicates if schema fields with no column metadata in any row group are dropped from the schema used for reading (default is false). -
isRemoveFieldsWithoutValues
public boolean isRemoveFieldsWithoutValues()Indicates if top-level columns that have no values in a row group are dropped from the schema used for reading (default is false). -
setRemoveFieldsWithoutValues
Indicates if top-level columns that have no values in a row group are dropped from the schema used for reading (default is false). -
toRecord
Description copied from interface:RecordSerializableConverts this object's state to a record thatRecordSerializable.fromRecord(Record)can load.- Specified by:
toRecordin interfaceRecordSerializable- Overrides:
toRecordin classPipelineInput
-
fromRecord
Description copied from interface:RecordSerializableLoads this instance's state from a record and returnsthis(for fluid API call chaining). For fluid API call chaining, the overridden method should change the declared return type to its class.- Specified by:
fromRecordin interfaceRecordSerializable- Overrides:
fromRecordin classPipelineInput- Parameters:
source-- Returns:
- this instance.
-
toXmlElement
Description copied from interface:XmlSerializableReturns an element, created withdocument, describing this object; the default implementation throws aDataException.- Specified by:
toXmlElementin interfaceXmlSerializable- Overrides:
toXmlElementin classPipelineInput
-
fromXmlElement
- Specified by:
fromXmlElementin interfaceXmlSerializable- Overrides:
fromXmlElementin classPipelineInput
-