Class CSVReader
Obtains records from a Comma Separated Value (CSV) or delimited stream.
-
Nested Class Summary
Nested classes/interfaces inherited from class com.northconcepts.datapipeline.core.DataEndpoint
DataEndpoint.State -
Field Summary
Fields inherited from class com.northconcepts.datapipeline.core.TextReader
EOF, readerFields inherited from class com.northconcepts.datapipeline.core.AbstractReader
currentRecord, fieldNames, lastRow, startingRowFields inherited from class com.northconcepts.datapipeline.core.DataReader
fieldLineage, recordLineageFields inherited from class com.northconcepts.datapipeline.core.DataEndpoint
lastRecord, PRODUCT, PRODUCT_VERSION, VENDOR, XML_INPUT_FACTORY_KEYFields inherited from class com.northconcepts.datapipeline.core.Endpoint
BUFFER_SIZE, captureElapsedTime, DEFAULT_READ_BUFFER_SIZEFields inherited from class com.northconcepts.datapipeline.core.DataObject
id, log, name, TIMESTAMP_FORMAT -
Constructor Summary
Constructors -
Method Summary
Modifier and TypeMethodDescriptionaddExceptionProperties(DataException exception) Adds this endpoint's current state to aDataException.protected RecordaddLineage(Record record) Called byDataReader.read()for each record fromDataReader.readImpl()while lineage is saved; the default copiesrecordLineageinto every field along with its original index and name.protected booleanfillRecord(Record record) Populates the given record with the next row's values; returns false at the end of the input.longReturns the 0-based character offset the parser has reached in the current line (not a field index).Returns the string used to mark the ending of text values that may contain commas and other special characters (defaults to double quote - ").Returns the string used between each field value (defaults to comma).String[]Returns the strings used to separate records ornullif none were explicitly assigned.Returns the starting quote - same as calling getStartingQuote() (defaults to double quote - ").Returns the string used to mark the beginning of text values that may contain commas and other special characters (defaults to double quote - ").booleanIndicates if line separators are allowed inside quoted fields.booleanIndicates if unescaped quotes are allowed in fields (default to false).booleanIndicates if this reader can capture record and field lineage (false unless overridden by a reader that supports it).booleanIndicates if space and tab characters at the beginning and ending of non-quoted-values should be removed.static RecordListUtility method to quickly parse lines of CSV text.static RecordListparse(String text, String fieldSeparator, String quoteChar, Boolean allowMultiLineText, Boolean trimFields) Utility method to quickly parse lines of CSV text.static RecordUtility method to quickly parse a line of CSV text.static RecordparseLine(String line, String fieldSeparator, String quoteChar, Boolean allowMultiLineText, Boolean trimFields) Utility method to quickly parse a line of CSV text.setAllowMultiLineText(boolean allowMultiLineText) Indicates if line separators are allowed inside quoted fields.setAllowQuoteInField(boolean allowQuoteInField) Indicates if unescaped quotes are allowed in fields (default to false).setEndingQuote(String endingQuote) Assigns the string used to mark the ending of text values that may contain commas and other special characters (defaults to double quote - ").setFieldNames(String... fieldNames) Sets the names given to each record's fields, taking precedence over names in the first row; must be called beforeAbstractReader.open().setFieldNames(Collection<String> fieldNames) Sets the names given to each record's fields, taking precedence over names in the first row; must be called beforeAbstractReader.open().setFieldNamesInFirstRow(boolean fieldNamesInFirstRow) Indicates if the first row holds the field names instead of data (default is false); must be called beforeAbstractReader.open().setFieldSeparator(char columnSeparator) Assigns the string used between each field value (defaults to comma).setFieldSeparator(String columnSeparator) Assigns the string used between each field value (defaults to comma).setLastRow(int lastRow) Sets the 0-based index of the last row to read or -1 to read to the end (default is -1).setLineSeparator(String lineSeparator) Indicates the string used to separate records.setLineSeparators(String... lineSeparators) Indicates the string or strings used to separate records.setQuoteChar(char quoteChar) Assigns startingQuote and endingQuote to the same value (defaults to double quote - ").setQuoteChar(String quoteChar) Assigns startingQuote and endingQuote to the same value (defaults to double quote - ").setSaveLineage(boolean saveLineage) Indicates if record and field lineage is captured for each record read (default is false); enabling it throws ifDataReader.isLineageSupported()is false or the product edition does not include lineage.setSkipEmptyRows(boolean skipEmptyRows) Indicates that rows with only null values should not be returned by the reader (default is false).setStartingQuote(String startingQuote) Assigns the string used to mark the beginning of text values that may contain commas and other special characters (defaults to double quote - ").setStartingRow(int startingRow) Sets the 0-based index of the first row to read; earlier rows are skipped when the reader is opened (default is 0).setTrimFields(boolean trimFields) Indicates if space and tab characters at the beginning and ending of non-quoted-values should be removed.Methods inherited from class com.northconcepts.datapipeline.core.TextReader
available, close, getFile, getLineNumber, open, readImplMethods inherited from class com.northconcepts.datapipeline.core.AbstractReader
getFieldNames, getLastRow, getStartingRow, isFieldNamesInFirstRow, isSkipEmptyRows, readMethods inherited from class com.northconcepts.datapipeline.core.DataReader
getBufferSize, getNestedEndpoint, getNestedReader, getReader, getRootEndpoint, getRootReader, isExhausted, isSaveLineage, peek, pop, push, skipMethods inherited from class com.northconcepts.datapipeline.core.DataEndpoint
decrementRecordCount, enableJmx, getLastRecord, getRecordCount, getRecordCountAsBigInteger, getRecordCountAsString, incrementRecordCount, isRecordCountBigInteger, resetRecordCount, toStringMethods inherited from class com.northconcepts.datapipeline.core.Endpoint
addElapsedtime, assertClosed, assertNotOpened, assertOpened, finalize, getClosedOn, getDescription, getElapsedTime, getElapsedTimeAsString, getOpenedOn, getOpenElapsedTime, getOpenElapsedTimeAsString, getSelfTime, getSelfTimeAsString, getState, isCaptureElapsedTime, isClosed, isOpen, setCaptureElapsedTime, setDescription
-
Constructor Details
-
CSVReader
-
CSVReader
-
-
Method Details
-
parseLine
public static Record parseLine(String line, String fieldSeparator, String quoteChar, Boolean allowMultiLineText) Utility method to quickly parse a line of CSV text.- Parameters:
line- the delimited line of text to parsefieldSeparator-quoteChar-allowMultiLineText-- Returns:
-
parseLine
public static Record parseLine(String line, String fieldSeparator, String quoteChar, Boolean allowMultiLineText, Boolean trimFields) Utility method to quickly parse a line of CSV text.- Parameters:
line-fieldSeparator-quoteChar-allowMultiLineText-trimFields-- Returns:
-
parse
public static RecordList parse(String text, String fieldSeparator, String quoteChar, Boolean allowMultiLineText) Utility method to quickly parse lines of CSV text.- Parameters:
text- the delimited lines of text to parsefieldSeparator-quoteChar-allowMultiLineText-- Returns:
-
parse
public static RecordList parse(String text, String fieldSeparator, String quoteChar, Boolean allowMultiLineText, Boolean trimFields) Utility method to quickly parse lines of CSV text.- Parameters:
text-fieldSeparator-quoteChar-allowMultiLineText-trimFields-- Returns:
-
getFieldSeparator
Returns the string used between each field value (defaults to comma). -
setFieldSeparator
Assigns the string used between each field value (defaults to comma). -
setFieldSeparator
Assigns the string used between each field value (defaults to comma). -
getQuoteChar
Returns the starting quote - same as calling getStartingQuote() (defaults to double quote - ").getStartingQuote() -
setQuoteChar
Assigns startingQuote and endingQuote to the same value (defaults to double quote - "). -
setQuoteChar
Assigns startingQuote and endingQuote to the same value (defaults to double quote - ").setStartingQuote(String)setEndingQuote(String) -
getStartingQuote
Returns the string used to mark the beginning of text values that may contain commas and other special characters (defaults to double quote - "). -
setStartingQuote
Assigns the string used to mark the beginning of text values that may contain commas and other special characters (defaults to double quote - "). -
getEndingQuote
Returns the string used to mark the ending of text values that may contain commas and other special characters (defaults to double quote - "). -
setEndingQuote
Assigns the string used to mark the ending of text values that may contain commas and other special characters (defaults to double quote - "). -
getLineSeparators
Returns the strings used to separate records ornullif none were explicitly assigned. -
setLineSeparators
Indicates the string or strings used to separate records. -
setLineSeparator
Indicates the string used to separate records. -
isAllowMultiLineText
public boolean isAllowMultiLineText()Indicates if line separators are allowed inside quoted fields. -
setAllowMultiLineText
Indicates if line separators are allowed inside quoted fields. -
isAllowQuoteInField
public boolean isAllowQuoteInField()Indicates if unescaped quotes are allowed in fields (default to false). -
setAllowQuoteInField
Indicates if unescaped quotes are allowed in fields (default to false). -
isTrimFields
public boolean isTrimFields()Indicates if space and tab characters at the beginning and ending of non-quoted-values should be removed. If line separators are explicitly set on this reader then any non-line-separator char matchingCharacter.isWhitespace(int)will be removed (not only spaces and tabs). This is helpful when CSV files may contain spaces between the commas (field separators) and values.
Whitespace characters used as field or line separators will not be removed.
The default istrue. -
setTrimFields
Indicates if space and tab characters at the beginning and ending of non-quoted-values should be removed. If line separators are explicitly set on this reader then any non-line-separator char matchingCharacter.isWhitespace(int)will be removed (not only spaces and tabs). This is helpful when CSV files may contain spaces between the commas (field separators) and values.
Whitespace characters used as field or line separators will not be removed.
The default istrue. -
getParser
-
getColumnNumber
public long getColumnNumber()Returns the 0-based character offset the parser has reached in the current line (not a field index). -
isLineageSupported
public boolean isLineageSupported()Description copied from class:DataReaderIndicates if this reader can capture record and field lineage (false unless overridden by a reader that supports it).- Overrides:
isLineageSupportedin classDataReader
-
addLineage
Description copied from class:DataReaderCalled byDataReader.read()for each record fromDataReader.readImpl()while lineage is saved; the default copiesrecordLineageinto every field along with its original index and name. Overrides set their source details onrecordLineagefirst and end withsuper.addLineage(record).- Overrides:
addLineagein classTextReader
-
addExceptionProperties
Description copied from class:EndpointAdds this endpoint's current state to aDataException. Since this method is called whenever an exception is thrown, subclasses should override it to add their specific information.- Overrides:
addExceptionPropertiesin classTextReader
-
fillRecord
Description copied from class:AbstractReaderPopulates the given record with the next row's values; returns false at the end of the input.- Specified by:
fillRecordin classAbstractReader- Throws:
Throwable
-
setSkipEmptyRows
Description copied from class:AbstractReaderIndicates that rows with only null values should not be returned by the reader (default is false).- Overrides:
setSkipEmptyRowsin classAbstractReader
-
setFieldNames
Description copied from class:AbstractReaderSets the names given to each record's fields, taking precedence over names in the first row; must be called beforeAbstractReader.open().- Overrides:
setFieldNamesin classAbstractReader
-
setFieldNames
Description copied from class:AbstractReaderSets the names given to each record's fields, taking precedence over names in the first row; must be called beforeAbstractReader.open().- Overrides:
setFieldNamesin classAbstractReader
-
setFieldNamesInFirstRow
Description copied from class:AbstractReaderIndicates if the first row holds the field names instead of data (default is false); must be called beforeAbstractReader.open().- Overrides:
setFieldNamesInFirstRowin classAbstractReader
-
setStartingRow
Description copied from class:AbstractReaderSets the 0-based index of the first row to read; earlier rows are skipped when the reader is opened (default is 0).- Overrides:
setStartingRowin classAbstractReader
-
setLastRow
Description copied from class:AbstractReaderSets the 0-based index of the last row to read or -1 to read to the end (default is -1).- Overrides:
setLastRowin classAbstractReader
-
setSaveLineage
Description copied from class:DataReaderIndicates if record and field lineage is captured for each record read (default is false); enabling it throws ifDataReader.isLineageSupported()is false or the product edition does not include lineage.- Overrides:
setSaveLineagein classDataReader
-