Class ParquetPipelineOutput

All Implemented Interfaces:
DataExceptionContributor, JsonSerializable, RecordSerializable, XmlSerializable, DataWriterFactory, JavaCodeGenerator, Serializable

public class ParquetPipelineOutput extends FileSinkPipelineOutput
A pipeline output that writes an Apache Parquet file to a FileSink's path using a ParquetDataWriter.
See Also:
  • Constructor Details

    • ParquetPipelineOutput

      public ParquetPipelineOutput()
  • Method Details

    • getName

      public String getName()
      Description copied from class: PipelineOutput
      Returns a display name for this output, such as the destination's name followed by its format.
      Overrides:
      getName in class FileSinkPipelineOutput
    • createDataWriter

      public DataWriter createDataWriter()
    • getFileSink

      public FileSink getFileSink()
      Specified by:
      getFileSink in class FileSinkPipelineOutput
    • setFileSink

      public ParquetPipelineOutput setFileSink(FileSink fileSink)
      Specified by:
      setFileSink in class FileSinkPipelineOutput
    • getCompressionCodecName

      public CompressionCodecName getCompressionCodecName()
      Indicates the compression used for writing (default UNCOMPRESSED).
    • getConfiguration

      public Map<String,String> getConfiguration()
      Returns the Parquet configuration parameters.
    • setConfiguration

      public ParquetPipelineOutput setConfiguration(Map<String,String> configuration)
      Sets the Parquet configuration parameters.
    • setConfiguration

      public ParquetPipelineOutput setConfiguration(String key, String value)
      Sets a single Parquet configuration parameter, keeping the others.
    • setCompressionCodecName

      public ParquetPipelineOutput setCompressionCodecName(CompressionCodecName compressionCodecName)
      Indicates the compression used for writing (default UNCOMPRESSED).
    • isDefaultAdjustedToUTC

      public boolean isDefaultAdjustedToUTC()
      Indicates if all datetime fields should be marked as AdjustedToUTC.
    • setDefaultAdjustedToUTC

      public ParquetPipelineOutput setDefaultAdjustedToUTC(boolean defaultAdjustedToUTC)
      Indicates if all datetime fields should be marked as AdjustedToUTC.
    • getRoundingMode

      public RoundingMode getRoundingMode()
      Indicates the rounding algorithm used for all BigDecimal values (default is RoundingMode.HALF_UP).
    • setRoundingMode

      public ParquetPipelineOutput setRoundingMode(RoundingMode roundingMode)
      Indicates the rounding algorithm used for all BigDecimal values (default is RoundingMode.HALF_UP).
    • getSchema

      public MessageType getSchema()
      Returns the schema used to write the file.
    • setSchema

      public ParquetPipelineOutput setSchema(MessageType schema)
      Sets the schema used to write the file.
    • getDefaultBigDecimalScale

      public int getDefaultBigDecimalScale()
      Returns the default scale used when writing BigDecimal values (default 5).
    • setDefaultBigDecimalScale

      public ParquetPipelineOutput setDefaultBigDecimalScale(int defaultBigDecimalScale)
      Sets the default scale used when writing BigDecimal values (default 5).
    • getDefaultBigNumberPrecision

      public int getDefaultBigNumberPrecision()
      Returns the default precision used when writing BigDecimal invalid input: '&' BigInteger values (default 25).
    • setDefaultBigNumberPrecision

      public ParquetPipelineOutput setDefaultBigNumberPrecision(int defaultBigNumberPrecision)
      Sets the default precision used when writing BigDecimal invalid input: '&' BigInteger values (default 25).
    • setMaxRecordsAnalyzed

      public ParquetPipelineOutput setMaxRecordsAnalyzed(Long maxRecordsAnalyzed)
      Indicates how many records should be analyzed and cached to generate the Parquet schema if no schema was explicitly set on this writer (default is 1000). This value will not be used if a schema was set on this writer.

      Passing in null will cause all records to be read and cached to determine the schema.

      The value will be set to 1 if a value less than 1 is passed in.

      Note: Using null or a high record count can significantly slow down processing and cause an OutOfMemoryError.
    • getCacheFolder

      public String getCacheFolder()
    • setCacheFolder

      public ParquetPipelineOutput setCacheFolder(String cacheFolder)
    • getRecordsPerCacheFile

      public int getRecordsPerCacheFile()
    • setRecordsPerCacheFile

      public ParquetPipelineOutput setRecordsPerCacheFile(int recordsPerCacheFile)
    • isRemoveUnsupportedChars

      public boolean isRemoveUnsupportedChars()
      Indicates if characters other than letters, digits, '_' and '-' are removed from field names (default is true).
    • setRemoveUnsupportedChars

      public ParquetPipelineOutput setRemoveUnsupportedChars(boolean removeUnsupportedChars)
      Indicates if characters other than letters, digits, '_' and '-' are removed from field names (default is true).
    • generateJavaCode

      public void generateJavaCode(JavaCodeBuilder code)
      Description copied from interface: JavaCodeGenerator
      Appends Java source code representing this object to the given builder.
    • toRecord

      public Record toRecord()
      Description copied from interface: RecordSerializable
      Converts this object's state to a record that RecordSerializable.fromRecord(Record) can load.
      Specified by:
      toRecord in interface RecordSerializable
      Overrides:
      toRecord in class Bean
    • fromRecord

      public ParquetPipelineOutput fromRecord(Record source)
      Description copied from interface: RecordSerializable
      Loads this instance's state from a record and returns this (for fluid API call chaining). For fluid API call chaining, the overridden method should change the declared return type to its class.
      Specified by:
      fromRecord in interface RecordSerializable
      Overrides:
      fromRecord in class PipelineOutput
      Parameters:
      source -
      Returns:
      this instance.
    • toXmlElement

      public Element toXmlElement(Document document)
      Description copied from interface: XmlSerializable
      Returns an element, created with document, describing this object; the default implementation throws a DataException.
      Specified by:
      toXmlElement in interface XmlSerializable
      Overrides:
      toXmlElement in class PipelineOutput
    • fromXmlElement

      public ParquetPipelineOutput fromXmlElement(Element pipelineOutputElement)