Class XmlReader

Direct Known Subclasses:
JavaBeanReader, JsonReader

public class XmlReader extends DataReader
Obtains records from an XML stream. See the Read an XML file example.

XmlReader's addField(String, String), addField(String, String, boolean), and addRecordBreak(String) methods use a subset of the XPath 1.0 location paths notation to identify field values and demarcate records.

Axis Specifiers

Axis Abbreviated Syntax Supported Examples
ancestor      
ancestor-or-self      
attribute @ yes @lang or attribute::lang
child   yes title or child::title
descendant   yes  
descendant-or-self // yes //book or /descendant-or-self::book/
following      
following-sibling      
namespace      
parent ..    
preceding      
preceding-sibling      
self . yes  

Node Tests

  • comment(), text(), processing-instruction(), node() are all supported

Predicates

  • None supported

Functions and Operators

  • None supported
  • Field Details

    • recordBreaks

      protected final List<XmlRecordBreak> recordBreaks
    • fields

      protected final List<XmlField> fields
    • reader

      protected final XmlNodeReader reader
    • file

      protected final File file
    • currentRecord

      protected Record currentRecord
    • hasCascadingFields

      protected boolean hasCascadingFields
  • Constructor Details

    • XmlReader

      public XmlReader(File file)
    • XmlReader

      public XmlReader(Reader reader)
    • XmlReader

      public XmlReader(XMLStreamReader streamReader)
      Creates a reader for an existing StAX stream; closing this reader closes the stream reader but not its underlying input.
    • XmlReader

      public XmlReader(XmlNodeReader reader)
      Creates a reader over the given node source; used by subclasses that present other formats, such as JSON or Java beans, as XML nodes.
  • Method Details

    • isDebug

      public boolean isDebug()
      Indicates if each node read is logged at debug level (default is false).
    • setDebug

      public XmlReader setDebug(boolean debug)
      Indicates if each node read is logged at debug level (default is false).
    • getXmlNodeReader

      protected XmlNodeReader getXmlNodeReader()
    • isAddTextToParent

      public boolean isAddTextToParent()
      Return true if each child node's text should be concatenated to its parent during parsing (defaults to false).
    • setAddTextToParent

      public XmlReader setAddTextToParent(boolean addTextToParent)
      Indicates if each child node's text should be concatenated to its parent during parsing (defaults to false). Setting this to true will result in higher memory consumption.
    • addRecordBreak

      public XmlReader addRecordBreak(String locationPathAsString)
      Adds a location path (for example //book); a record is completed each time a matching element ends.
    • addField

      public XmlReader addField(XmlField field)
    • addField

      public XmlReader addField(String name, String locationPathAsString, boolean cascadeValues)
      Identifies a new field using XPath in the XML stream.
      Parameters:
      name - the field name to create
      locationPathAsString - the XPath to match to populate this field
      cascadeValues - indicates if this reader should return the last value seen for this field when no matches are available.
    • addField

      public XmlReader addField(String name, String locationPathAsString, String cascadeResetLocationPath)
      Identifies a new field using XPath in the XML stream. If cascadeResetLocationPath is not null and not empty, it will indicate when this field should be cleared, otherwise this field will return the last value seen when no new matches are available.
      Parameters:
      name - the field name to create
      locationPathAsString - the XPath to match to populate this field
      cascadeResetLocationPath - the XPath to identify when cascading values (i.e. the last value seen for this field) should be cleared.
    • addField

      public XmlReader addField(String name, String locationPathAsString)
      Identifies a new field using XPath in the XML stream.
      Parameters:
      name - the field name to create
      locationPathAsString - the XPath to match to populate this field
    • getDuplicateFieldPolicy

      public XmlReader.DuplicateFieldPolicy getDuplicateFieldPolicy()
      Returns how a field matched more than once in a record is handled (default is XmlReader.DuplicateFieldPolicy.USE_LAST_VALUE).
    • setDuplicateFieldPolicy

      public XmlReader setDuplicateFieldPolicy(XmlReader.DuplicateFieldPolicy duplicateFieldPolicy)
      Sets how a field matched more than once in a record is handled; null restores the default XmlReader.DuplicateFieldPolicy.USE_LAST_VALUE.
    • isIgnoreNamespaces

      public boolean isIgnoreNamespaces()
      Indicates if namespaces on elements and attributes are ignored when matching expressions (default to true).
    • setIgnoreNamespaces

      public XmlReader setIgnoreNamespaces(boolean ignoreNamespaces)
      Indicates if namespaces on elements and attributes are ignored when matching expressions (default to true).
    • createRecord

      protected void createRecord()
      Starts a new current record with one field per added XmlField, copying the values of cascading fields from the previous record.
    • setSaveLineage

      public XmlReader setSaveLineage(boolean saveLineage)
      Description copied from class: DataReader
      Indicates if record and field lineage is captured for each record read (default is false); enabling it throws if DataReader.isLineageSupported() is false or the product edition does not include lineage.
      Overrides:
      setSaveLineage in class DataReader
    • setDescription

      public XmlReader setDescription(String description)
      Description copied from class: Endpoint
      Sets the optional text shown for this endpoint in Endpoint.toString() and exception properties.
      Overrides:
      setDescription in class Endpoint
    • readImpl

      protected Record readImpl() throws Throwable
      Description copied from class: DataReader
      Overridden by subclasses to read the next record from this DataReader. The default implementation of DataReader.read() now insures that this method will not be called again after it returns a null.

      If no record is available, null will be returned.

      Contract for subclasses (see also docs/authoring/DataReader.md):

      Specified by:
      readImpl in class DataReader
      Throws:
      Throwable
    • expandAndPushRecords

      protected void expandAndPushRecords(Record record)
    • expandListFieldsAsFields

      protected void expandListFieldsAsFields(Record record)
    • setRecordField

      protected void setRecordField(XmlNode node, XmlField xmlField, LocationPath path)
      Stores the value path selects from the node in the current record's field, replacing the field's value under XmlReader.DuplicateFieldPolicy.USE_LAST_VALUE and adding to it otherwise.
    • saveFieldValues

      protected void saveFieldValues(XmlNode node)
      Sets the fields whose location paths match the node and clears cascading fields whose reset path matches it.
    • saveAncestorAttributeFieldValues

      protected void saveAncestorAttributeFieldValues(XmlNode node)
      Sets the attribute fields (paths ending in an attribute step) matching the record-break node or an ancestor.
    • saveAncestorNodeFieldValues

      protected void saveAncestorNodeFieldValues(XmlNode node)
      Fills the still-empty element fields from ancestors of the record-break node that match their location paths.
    • addFieldValue

      protected void addFieldValue(Field recordField, Object value)
    • getFieldValues

      protected void getFieldValues(long sequence)
    • isRecordBreak

      protected boolean isRecordBreak(XmlNode node)
      Returns true if the node matches any location path added with addRecordBreak(String).
    • open

      public void open() throws DataException
      Description copied from class: DataEndpoint
      Makes this endpoint ready for reading or writing.
      Overrides:
      open in class DataEndpoint
      Throws:
      DataException
    • close

      public void close() throws DataException
      Description copied from class: DataEndpoint
      Indicates that this endpoint has finished reading or writing.
      Overrides:
      close in class DataEndpoint
      Throws:
      DataException
    • isLineageSupported

      public boolean isLineageSupported()
      Description copied from class: DataReader
      Indicates if this reader can capture record and field lineage (false unless overridden by a reader that supports it).
      Overrides:
      isLineageSupported in class DataReader
    • addLineage

      protected Record addLineage(Record record)
      Description copied from class: DataReader
      Called by DataReader.read() for each record from DataReader.readImpl() while lineage is saved; the default copies recordLineage into every field along with its original index and name. Overrides set their source details on recordLineage first and end with super.addLineage(record).
      Overrides:
      addLineage in class DataReader
    • addExceptionProperties

      public DataException addExceptionProperties(DataException exception)
      Description copied from class: Endpoint
      Adds this endpoint's current state to a DataException. Since this method is called whenever an exception is thrown, subclasses should override it to add their specific information.
      Overrides:
      addExceptionProperties in class DataReader
    • isAutoCloseReader

      public boolean isAutoCloseReader()
      Indicates if the Reader passed to the constructor is closed when this reader closes (default is true); a file opened by this reader is always closed.
    • setAutoCloseReader

      public XmlReader setAutoCloseReader(boolean autoCloseReader)
      Indicates if the Reader passed to the constructor is closed when this reader closes (default is true); a file opened by this reader is always closed.