Class ParquetDataReader


public class ParquetDataReader extends IntegrationReader
Read records from Apache Parquet columnar files. See Apache Parquet columnar storage.
  • Constructor Details

    • ParquetDataReader

      public ParquetDataReader(InputFile inputFile)
      Reads parquet data from an InputFile.
  • Method Details

    • isDebug

      public boolean isDebug()
      Indicates if debugging is enabled to print log statements.
    • setDebug

      public ParquetDataReader setDebug(boolean debug)
      Indicates if debugging is enabled to print log statements.
    • getConfiguration

      public Configuration getConfiguration()
      Returns the Parquet configuration parameters.
    • setConfiguration

      public ParquetDataReader setConfiguration(Configuration configuration)
      Sets the Parquet configuration parameters.
    • getFilter

      public MetadataFilter getFilter()
      Returns the filter settings.
    • setFilter

      public ParquetDataReader setFilter(MetadataFilter filter)
      Sets the filter settings.
    • getSchema

      public MessageType getSchema()
      Indicates the schema used to read the file.
    • getModifiedSchema

      public MessageType getModifiedSchema()
      Indicates the modified schema used to read the file.
    • setSchema

      public ParquetDataReader setSchema(MessageType schema)
      Indicates the schema used to read the file.
    • isMakeRequiredFieldsOptional

      public boolean isMakeRequiredFieldsOptional()
      When true, check for required columns that can be made optional using
      invalid reference
      ColumnChunkMetaData#getStatistics()
      . See getRequiredColumnsWithNullValues() for details.
    • setMakeRequiredFieldsOptional

      public ParquetDataReader setMakeRequiredFieldsOptional(boolean makeRequiredFieldsOptional)
      When true, check for required columns that can be made optional using
      invalid reference
      ColumnChunkMetaData#getStatistics()
      . See getRequiredColumnsWithNullValues() for details.
    • isMakeOptionalFieldsRequired

      public boolean isMakeOptionalFieldsRequired()
      Used whether to check for optional columns that can be made required using
      invalid reference
      ColumnChunkMetaData#getStatistics()
      during open(). See getOptionalColumnsWithoutNullValues() ()} for details.
    • setMakeOptionalFieldsRequired

      public ParquetDataReader setMakeOptionalFieldsRequired(boolean makeOptionalFieldsRequired)
      Used whether to check for optional columns that can be made required using
      invalid reference
      ColumnChunkMetaData#getStatistics()
      during open(). See getOptionalColumnsWithoutNullValues() ()} for details.
    • isRemoveFieldsWithoutColumnMetadata

      public boolean isRemoveFieldsWithoutColumnMetadata()
      When true, remove fields in the schema does not have a corresponding ColumnChunkMetaData in the file. See getFieldsWithoutColumnMetadata() for details.
    • setRemoveFieldsWithoutColumnMetadata

      public ParquetDataReader setRemoveFieldsWithoutColumnMetadata(boolean removeFieldsWithoutColumnMetadata)
      Used whether to remove fields in the schema does not have a corresponding ColumnChunkMetaData in the file during open(). See getFieldsWithoutColumnMetadata() for details.
    • isRemoveFieldsWithoutValues

      public boolean isRemoveFieldsWithoutValues()
      When true, remove fields in the schema whose value count is invalid input: '<'= 0 See getFieldsWithoutColumnMetadata() for details.
    • setRemoveFieldsWithoutValues

      public ParquetDataReader setRemoveFieldsWithoutValues(boolean removeFieldsWithoutValues)
      Used whether to remove fields in the schema whose value count is invalid input: '<'= 0 during open(). See getFieldsWithoutColumnMetadata() for details.
    • open

      public void open() throws DataException
      Description copied from class: DataEndpoint
      Makes this endpoint ready for reading or writing.
      Overrides:
      open in class IntegrationReader
      Throws:
      DataException
    • close

      public void close() throws DataException
      Description copied from class: DataEndpoint
      Indicates that this endpoint has finished reading or writing.
      Overrides:
      close in class DataEndpoint
      Throws:
      DataException
    • readImpl

      protected Record readImpl() throws Throwable
      Description copied from class: DataReader
      Overridden by subclasses to read the next record from this DataReader. The default implementation of DataReader.read() now insures that this method will not be called again after it returns a null.

      If no record is available, null will be returned.

      Contract for subclasses (see also docs/authoring/DataReader.md):

      Specified by:
      readImpl in class DataReader
      Throws:
      Throwable
    • readRecord

      protected Record readRecord(Group group, int depth)
      Converts a Parquet group into a record with one field per group field; depth is 0 for a row and grows with nesting.
    • readGroupValue

      protected void readGroupValue(Group group, int depth, int fieldIndex, ValueNodeContainer valueContainer)
      Adds the value of the group's field at fieldIndex to the container: primitives as values, lists as arrays, and maps and other groups as records.
    • readPrimitiveValue

      protected void readPrimitiveValue(Group group, int fieldIndex, ValueNodeContainer valueContainer, Type fieldType)
      Adds the values of the group's primitive field at fieldIndex to the container, converted using its logical type, or a typed null if it has none; times and timestamps are truncated to milliseconds.
    • int96ToTimestamp

      protected Date int96ToTimestamp(byte[] int96Bytes)
      Converts a 12-byte Parquet INT96 timestamp to a Date in the system default time zone.
    • isLineageSupported

      public boolean isLineageSupported()
      Description copied from class: DataReader
      Indicates if this reader can capture record and field lineage (false unless overridden by a reader that supports it).
      Overrides:
      isLineageSupported in class DataReader
    • addLineage

      protected Record addLineage(Record record)
      Description copied from class: DataReader
      Called by DataReader.read() for each record from DataReader.readImpl() while lineage is saved; the default copies recordLineage into every field along with its original index and name. Overrides set their source details on recordLineage first and end with super.addLineage(record).
      Overrides:
      addLineage in class DataReader
    • addExceptionProperties

      public DataException addExceptionProperties(DataException exception)
      Description copied from class: Endpoint
      Adds this endpoint's current state to a DataException. Since this method is called whenever an exception is thrown, subclasses should override it to add their specific information.
      Overrides:
      addExceptionProperties in class DataReader