Class OrcPipelineOutput

All Implemented Interfaces:
DataExceptionContributor, JsonSerializable, RecordSerializable, XmlSerializable, DataWriterFactory, JavaCodeGenerator, Serializable

public class OrcPipelineOutput extends FileSinkPipelineOutput
A pipeline output that writes an Apache ORC file to a FileSink's path using an OrcDataWriter.
See Also:
  • Constructor Details

    • OrcPipelineOutput

      public OrcPipelineOutput()
  • Method Details

    • getName

      public String getName()
      Description copied from class: PipelineOutput
      Returns a display name for this output, such as the destination's name followed by its format.
      Overrides:
      getName in class FileSinkPipelineOutput
    • getFileSink

      public FileSink getFileSink()
      Specified by:
      getFileSink in class FileSinkPipelineOutput
    • setFileSink

      public OrcPipelineOutput setFileSink(FileSink fileSink)
      Specified by:
      setFileSink in class FileSinkPipelineOutput
    • createDataWriter

      public DataWriter createDataWriter()
    • generateJavaCode

      public void generateJavaCode(JavaCodeBuilder javaCodeBuilder)
      Description copied from interface: JavaCodeGenerator
      Appends Java source code representing this object to the given builder.
    • getSchema

      public TypeDescription getSchema()
      Returns the schema used to write the file.
    • setSchema

      public OrcPipelineOutput setSchema(TypeDescription schema)
      Sets the schema used to write the file.
    • getConfiguration

      public Map<String,String> getConfiguration()
      Returns the Orc configuration parameters.
    • addConfiguration

      public OrcPipelineOutput addConfiguration(Map<String,String> configuration)
      Sets the Orc configuration parameters.
    • addConfiguration

      public OrcPipelineOutput addConfiguration(String key, String value)
      Sets the Orc configuration parameters.
    • getBatchSize

      public int getBatchSize()
      Indicates the maximum number of records to buffer when writing ORC data (default 1024).
    • setBatchSize

      public OrcPipelineOutput setBatchSize(int batchSize)
      Indicates the maximum number of records to buffer when writing ORC data (default 1024).
    • isUseTimestampInstant

      public boolean isUseTimestampInstant()
      Indicates if Instants should be used to store datetime values.
    • setUseTimestampInstant

      public OrcPipelineOutput setUseTimestampInstant(boolean useTimestampInstant)
      Indicates if Instants should be used to store datetime values.
    • isOverwrite

      public boolean isOverwrite()
      Overwrites the file if true, otherwise, throws an exception if the file already exists.
    • setOverwrite

      public OrcPipelineOutput setOverwrite(boolean overwrite)
      Overwrites the file if true, otherwise, throws an exception if the file already exists.
    • isUTC

      public boolean isUTC()
      Indicates if datetime values are adjusted to UTC.
    • setUTC

      public void setUTC(boolean UTC)
      Indicates if datetime values are adjusted to UTC.
    • isWriteUuidAsString

      public boolean isWriteUuidAsString()

      Indicates whether UUID values are written as strings or bytes. If set to false, UUID values will be written as bytes.


      Default value is true.

    • setWriteUuidAsString

      public OrcPipelineOutput setWriteUuidAsString(boolean writeUuidAsString)

      Indicates whether UUID values are written as strings or bytes. If set to false, UUID values will be written as bytes.


      Default value is true.

    • getRowBatchVersion

      public RowBatchVersion getRowBatchVersion()
      Returns RowBatchVersion to store ORC data. precision > TypeDescription.MAX_DECIMAL64_PRECISION (18) then DecimalColumnVector is used instead of Decimal64ColumnVector.
    • setRowBatchVersion

      public OrcPipelineOutput setRowBatchVersion(RowBatchVersion rowBatchVersion)
      Sets RowBatchVersion to store ORC data. If BigDecimal value has precision > TypeDescription.MAX_DECIMAL64_PRECISION (18) then DecimalColumnVector is used instead of Decimal64ColumnVector.
    • getBigDecimalPrecision

      public int getBigDecimalPrecision()
      Returns default precision value for BigDecimal type of values. Default value is HiveDecimal.MAX_PRECISION (38).
    • setBigDecimalPrecision

      public OrcPipelineOutput setBigDecimalPrecision(int bigDecimalPrecision)
      Sets default precision value for BigDecimal type of values. Default value is HiveDecimal.MAX_PRECISION (38).
    • getBigDecimalScale

      public int getBigDecimalScale()
      Indicates the scale used when writing BigDecimal values (default 10).
    • setBigDecimalScale

      public OrcPipelineOutput setBigDecimalScale(int bigDecimalScale)
      Indicates the scale used when writing BigDecimal values (default 10).
    • getRoundingMode

      public RoundingMode getRoundingMode()
      Indicates the rounding algorithm used for all BigDecimal values (default is RoundingMode.HALF_UP).
    • setRoundingMode

      public OrcPipelineOutput setRoundingMode(RoundingMode roundingMode)
      Indicates the rounding algorithm used for all BigDecimal values (default is RoundingMode.HALF_UP).
    • getCompressionKind

      public CompressionKind getCompressionKind()
      Indicates the compression used for writing (default NONE).
    • setCompressionKind

      public OrcPipelineOutput setCompressionKind(CompressionKind compressionKind)
      Indicates the compression used for writing (default NONE).
    • toRecord

      public Record toRecord()
      Description copied from interface: RecordSerializable
      Converts this object's state to a record that RecordSerializable.fromRecord(Record) can load.
      Specified by:
      toRecord in interface RecordSerializable
      Overrides:
      toRecord in class Bean
    • fromRecord

      public OrcPipelineOutput fromRecord(Record source)
      Description copied from interface: RecordSerializable
      Loads this instance's state from a record and returns this (for fluid API call chaining). For fluid API call chaining, the overridden method should change the declared return type to its class.
      Specified by:
      fromRecord in interface RecordSerializable
      Overrides:
      fromRecord in class PipelineOutput
      Parameters:
      source -
      Returns:
      this instance.
    • toXmlElement

      public Element toXmlElement(Document document)
      Description copied from interface: XmlSerializable
      Returns an element, created with document, describing this object; the default implementation throws a DataException.
      Specified by:
      toXmlElement in interface XmlSerializable
      Overrides:
      toXmlElement in class PipelineOutput
    • fromXmlElement

      public OrcPipelineOutput fromXmlElement(Element pipelineOutputElement)