Class LocalFileDataset
java.lang.Object
com.northconcepts.datapipeline.foundations.core.Bean
com.northconcepts.datapipeline.foundations.core.FoundationObject
com.northconcepts.datapipeline.foundations.pipeline.dataset.Dataset
com.northconcepts.datapipeline.foundations.pipeline.dataset.LocalFileDataset
- All Implemented Interfaces:
DataExceptionContributor,JsonSerializable,RecordSerializable,XmlSerializable,Closeable,Serializable,AutoCloseable,Iterable<Record>
Caches the dataset's records on disk as binary data.
- See Also:
-
Nested Class Summary
Nested classes/interfaces inherited from class com.northconcepts.datapipeline.foundations.pipeline.dataset.Dataset
Dataset.ColumnsDataReader -
Field Summary
Fields inherited from class com.northconcepts.datapipeline.foundations.core.FoundationObject
internalId, internalName, log, TIMESTAMP_FORMATFields inherited from interface com.northconcepts.datapipeline.core.RecordSerializable
SERIALIZED_CLASS_NAME, TYPEFields inherited from interface com.northconcepts.datapipeline.core.XmlSerializable
XML_SERIALIZED_CLASS_NAME -
Constructor Summary
ConstructorsModifierConstructorDescriptionprotectedLocalFileDataset(String rootFolder, String fileNamePrefix) Opens a dataset saved inrootFolderwithfileNamePrefix.protectedLocalFileDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline) Creates a dataset that, when loaded, saves the pipeline's records to files inrootFolderwhose names start withfileNamePrefix. -
Method Summary
Modifier and TypeMethodDescriptionaddExceptionProperties(DataException exception) Adds this object's current state to aDataException.protected voidCalled during the data loading process after all the column stats have been loaded.protected voidCalled at the end of the data loading process after all the records and column stats have been loaded.protected voidCalled at the start of the data loading process, but before any records or column stats have been loaded.voidclose()static LocalFileDatasetcreateDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline) Creates a persistent dataset on disk in the specifiedrootFolderdirectory.static LocalFileDatasetcreateDataset(String rootFolder, String fileNamePrefix, DataReaderFactory dataReaderFactory) Creates a persistent dataset on disk in the specifiedrootFolderdirectory.protected DataWriterWrites records to this dataset's cache after clearing it.static LocalFileDatasetcreateTempDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline) Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits.static LocalFileDatasetcreateTempDataset(String rootFolder, String fileNamePrefix, DataReaderFactory dataReaderFactory) Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits.getColumn(int index) Returns the column with the given name or null if there is none.longprotected FilegetDataFile(long fileIndex) Returns the data file holding the given 0-based chunk ofgetRecordsPerFile()records.protected FileReturns the JSON file holding the dataset's record, column and data file counts and its records per file.protected ColumngetOrCreateColumn(String name, int index) Returns the named column, creating it with the given 0-based field index if needed; called while collecting column stats.getRecord(long index) Returns the loaded record at the given 0-based index.longReturns the number of records loaded into this dataset so far.getRecordList(long offset, int count) Get a subset of the records cached in this dataset.intReturns the number of records stored in each data file (defaults to 1000).protected FileReturns the file holding the serialized column stats.booleanIndicates if this dataset's files are deleted when it is closed (defaults to false).load(Integer maxRecords, JobCallback<DataReader, DataWriter> callback) Starts the asynchronous loading of records from the pipeline into this dataset.protected InputStreamnewInputStream(File file) Opens one of this dataset's files for reading.protected OutputStreamnewOutputStream(File file) Creates one of this dataset's files for writing, replacing any existing content.static LocalFileDatasetopenDataset(String rootFolder, String fileNamePrefix) Loads an existing dataset from the specifiedrootFolderinvalid input: '&'fileNamePrefix.setDeleteFilesOnClose(boolean deleteFilesOnClose) Indicates if this dataset's files are deleted when it is closed (defaults to false).setPipeline(AbstractPipeline pipeline) setRecordsPerFile(int recordsPerFile) Sets the number of records stored in each data file (defaults to 1000); only allowed before the data is loaded.Methods inherited from class com.northconcepts.datapipeline.foundations.pipeline.dataset.Dataset
addField, afterRecordsLoaded, cancelLoad, createColumnsDataReader, createDataReader, createDataReader, finalize, forEach, fromRecord, getColumnStatsException, getColumnStatsReaderThreads, getDataLoadException, getJob, getJobExecutor, getMaxColumnStatsRecords, getMaxColumnsToAnalyze, getMaxRecordsToLoad, getPipeline, isCollectUniqueValues, isColumnStatsLoaded, isDataLoaded, isDataLoading, isDetectBigNumberValues, isDetectBooleanValues, isDetectNumericValues, isDetectTemporalValues, isDetectUuidValues, isInferStringTypes, isRecordsLoaded, iterator, load, load, setCollectUniqueValues, setColumnStatsLoaded, setColumnStatsReaderThreads, setDetectBigNumberValues, setDetectBooleanValues, setDetectNumericValues, setDetectTemporalValues, setDetectUuidValues, setInferStringTypes, setJobExecutor, setMaxColumnStatsRecords, setMaxColumnsToAnalyze, setRecordsLoaded, stream, toRecord, updateColumns, waitForColumnStatsToLoad, waitForColumnStatsToLoad, waitForRecordsToLoad, waitForRecordsToLoad, waitUntilJobFinishedMethods inherited from class com.northconcepts.datapipeline.foundations.core.FoundationObject
assertValid, assertValid, clone, exception, exception, exception, getInternalId, getInternalName, resetInternalIdMethods inherited from class java.lang.Object
equals, getClass, hashCode, notify, notifyAll, wait, wait, waitMethods inherited from interface com.northconcepts.datapipeline.core.DataExceptionContributor
addExceptionPropertiesMethods inherited from interface java.lang.Iterable
spliteratorMethods inherited from interface com.northconcepts.datapipeline.core.RecordSerializable
fromJson, fromJson, toJson, toJson, toJsonMethods inherited from interface com.northconcepts.datapipeline.core.XmlSerializable
fromXml, fromXml, fromXmlElement, toXml, toXml, toXml, toXml, toXml, toXmlElement
-
Constructor Details
-
LocalFileDataset
Creates a dataset that, when loaded, saves the pipeline's records to files inrootFolderwhose names start withfileNamePrefix. -
LocalFileDataset
Opens a dataset saved inrootFolderwithfileNamePrefix. Its metadata and stats files are read now and its records on demand.
-
-
Method Details
-
createTempDataset
public static LocalFileDataset createTempDataset(String rootFolder, String fileNamePrefix, DataReaderFactory dataReaderFactory) Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits. The temporary file will be created in the specified directory (or the default temporary-file directory ifnull), using the prefix "dataset-" and suffix ".mvstore" to generate its name.- Parameters:
rootFolder- the directory to store the temporary file ornullto use the default temporary-file directory.fileNamePrefix- the base file name of data invalid input: '&' stats.dataReaderFactory- the source of the dataset's data
-
createTempDataset
public static LocalFileDataset createTempDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline) Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits. The temporary files will be created in the default temporary-file or specified directory, using the specified prefix.- Parameters:
rootFolder- the directory to store the temporary file ornullto use the default temporary-file directory.fileNamePrefix- the base file name of data invalid input: '&' stats.pipeline- the source of the dataset's data
-
createDataset
public static LocalFileDataset createDataset(String rootFolder, String fileNamePrefix, DataReaderFactory dataReaderFactory) Creates a persistent dataset on disk in the specifiedrootFolderdirectory. The data invalid input: '&' stats files will remain on disk even after the dataset is closed and the JVM exits.- Parameters:
rootFolder- the root directory where data invalid input: '&' stats files will be saved.fileNamePrefix- the base file name of data invalid input: '&' stats.dataReaderFactory- the source of the dataset's data
-
createDataset
public static LocalFileDataset createDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline) Creates a persistent dataset on disk in the specifiedrootFolderdirectory. The data invalid input: '&' stats files will remain on disk even after the dataset is closed and the JVM exits.- Parameters:
rootFolder- the root directory where data invalid input: '&' stats files will be saved.fileNamePrefix- the base file name of data invalid input: '&' stats.pipeline- the source of the dataset's data
-
openDataset
Loads an existing dataset from the specifiedrootFolderinvalid input: '&'fileNamePrefix.- Parameters:
rootFolder- the root directory where data invalid input: '&' stats files are available.fileNamePrefix- the base file name of data invalid input: '&' stats.
-
getDataFile
Returns the data file holding the given 0-based chunk ofgetRecordsPerFile()records. -
getStatsFile
Returns the file holding the serialized column stats. -
getMetadataFile
Returns the JSON file holding the dataset's record, column and data file counts and its records per file. -
newInputStream
Opens one of this dataset's files for reading. Every read of a data, statistics, or metadata file goes through here, so overriding this (together withnewOutputStream(File)) is enough to change how the dataset is stored on disk — to hold it encrypted, for example.Implementations must be symmetric with
newOutputStream(File)and must not depend on subclass state:LocalFileDataset(String, String)reads the metadata and statistics files while the superclass constructor is still running, before subclass fields are assigned. -
newOutputStream
Creates one of this dataset's files for writing, replacing any existing content. Every write of a data, statistics, or metadata file goes through here. SeenewInputStream(File). -
setPipeline
- Overrides:
setPipelinein classDataset
-
isDeleteFilesOnClose
public boolean isDeleteFilesOnClose()Indicates if this dataset's files are deleted when it is closed (defaults to false). -
setDeleteFilesOnClose
Indicates if this dataset's files are deleted when it is closed (defaults to false). -
getRecordsPerFile
public int getRecordsPerFile()Returns the number of records stored in each data file (defaults to 1000). -
setRecordsPerFile
Sets the number of records stored in each data file (defaults to 1000); only allowed before the data is loaded. -
close
public void close() -
getRecordCount
public long getRecordCount()Description copied from class:DatasetReturns the number of records loaded into this dataset so far.- Specified by:
getRecordCountin classDataset
-
getRecord
Description copied from class:DatasetReturns the loaded record at the given 0-based index. -
getRecordList
Description copied from class:DatasetGet a subset of the records cached in this dataset.- Overrides:
getRecordListin classDataset
-
getColumnCount
public long getColumnCount()- Specified by:
getColumnCountin classDataset
-
getColumnNames
- Specified by:
getColumnNamesin classDataset
-
getColumn
-
getColumn
Description copied from class:DatasetReturns the column with the given name or null if there is none. -
getOrCreateColumn
Description copied from class:DatasetReturns the named column, creating it with the given 0-based field index if needed; called while collecting column stats.- Specified by:
getOrCreateColumnin classDataset
-
getColumns
- Specified by:
getColumnsin classDataset
-
beforeLoad
protected void beforeLoad()Description copied from class:DatasetCalled at the start of the data loading process, but before any records or column stats have been loaded.- Specified by:
beforeLoadin classDataset
-
afterLoad
protected void afterLoad()Description copied from class:DatasetCalled at the end of the data loading process after all the records and column stats have been loaded. -
afterColumnStatsLoaded
protected void afterColumnStatsLoaded()Description copied from class:DatasetCalled during the data loading process after all the column stats have been loaded. The records would have already been loaded when this method is called since column stats require additional processing.- Overrides:
afterColumnStatsLoadedin classDataset
-
createDataWriter
Description copied from class:DatasetWrites records to this dataset's cache after clearing it.- Specified by:
createDataWriterin classDataset
-
load
Description copied from class:DatasetStarts the asynchronous loading of records from the pipeline into this dataset. This method returns immediately and does not wait for loading to complete. SeeDataset.waitForRecordsToLoad()andDataset.waitForRecordsToLoad(long, long). -
addExceptionProperties
Description copied from class:FoundationObjectAdds this object's current state to aDataException. Since this method is called whenever an exception is thrown, subclasses should override it to add their specific information.- Specified by:
addExceptionPropertiesin interfaceDataExceptionContributor- Overrides:
addExceptionPropertiesin classFoundationObject
-