All Implemented Interfaces:
DataExceptionContributor, JsonSerializable, RecordSerializable, XmlSerializable, Closeable, Serializable, AutoCloseable, Iterable<Record>

public class LocalFileDataset extends Dataset
Caches the dataset's records on disk as binary data.
See Also:
  • Constructor Details

    • LocalFileDataset

      protected LocalFileDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline)
      Creates a dataset that, when loaded, saves the pipeline's records to files in rootFolder whose names start with fileNamePrefix.
    • LocalFileDataset

      protected LocalFileDataset(String rootFolder, String fileNamePrefix)
      Opens a dataset saved in rootFolder with fileNamePrefix. Its metadata and stats files are read now and its records on demand.
  • Method Details

    • createTempDataset

      public static LocalFileDataset createTempDataset(String rootFolder, String fileNamePrefix, DataReaderFactory dataReaderFactory)
      Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits. The temporary file will be created in the specified directory (or the default temporary-file directory if null), using the prefix "dataset-" and suffix ".mvstore" to generate its name.
      Parameters:
      rootFolder - the directory to store the temporary file or null to use the default temporary-file directory.
      fileNamePrefix - the base file name of data invalid input: '&' stats.
      dataReaderFactory - the source of the dataset's data
    • createTempDataset

      public static LocalFileDataset createTempDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline)
      Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits. The temporary files will be created in the default temporary-file or specified directory, using the specified prefix.
      Parameters:
      rootFolder - the directory to store the temporary file or null to use the default temporary-file directory.
      fileNamePrefix - the base file name of data invalid input: '&' stats.
      pipeline - the source of the dataset's data
    • createDataset

      public static LocalFileDataset createDataset(String rootFolder, String fileNamePrefix, DataReaderFactory dataReaderFactory)
      Creates a persistent dataset on disk in the specified rootFolder directory. The data invalid input: '&' stats files will remain on disk even after the dataset is closed and the JVM exits.
      Parameters:
      rootFolder - the root directory where data invalid input: '&' stats files will be saved.
      fileNamePrefix - the base file name of data invalid input: '&' stats.
      dataReaderFactory - the source of the dataset's data
    • createDataset

      public static LocalFileDataset createDataset(String rootFolder, String fileNamePrefix, AbstractPipeline pipeline)
      Creates a persistent dataset on disk in the specified rootFolder directory. The data invalid input: '&' stats files will remain on disk even after the dataset is closed and the JVM exits.
      Parameters:
      rootFolder - the root directory where data invalid input: '&' stats files will be saved.
      fileNamePrefix - the base file name of data invalid input: '&' stats.
      pipeline - the source of the dataset's data
    • openDataset

      public static LocalFileDataset openDataset(String rootFolder, String fileNamePrefix)
      Loads an existing dataset from the specified rootFolder invalid input: '&' fileNamePrefix.
      Parameters:
      rootFolder - the root directory where data invalid input: '&' stats files are available.
      fileNamePrefix - the base file name of data invalid input: '&' stats.
    • getDataFile

      protected File getDataFile(long fileIndex)
      Returns the data file holding the given 0-based chunk of getRecordsPerFile() records.
    • getStatsFile

      protected File getStatsFile()
      Returns the file holding the serialized column stats.
    • getMetadataFile

      protected File getMetadataFile()
      Returns the JSON file holding the dataset's record, column and data file counts and its records per file.
    • newInputStream

      protected InputStream newInputStream(File file)
      Opens one of this dataset's files for reading. Every read of a data, statistics, or metadata file goes through here, so overriding this (together with newOutputStream(File)) is enough to change how the dataset is stored on disk — to hold it encrypted, for example.

      Implementations must be symmetric with newOutputStream(File) and must not depend on subclass state: LocalFileDataset(String, String) reads the metadata and statistics files while the superclass constructor is still running, before subclass fields are assigned.

    • newOutputStream

      protected OutputStream newOutputStream(File file)
      Creates one of this dataset's files for writing, replacing any existing content. Every write of a data, statistics, or metadata file goes through here. See newInputStream(File).
    • setPipeline

      public LocalFileDataset setPipeline(AbstractPipeline pipeline)
      Overrides:
      setPipeline in class Dataset
    • isDeleteFilesOnClose

      public boolean isDeleteFilesOnClose()
      Indicates if this dataset's files are deleted when it is closed (defaults to false).
    • setDeleteFilesOnClose

      public LocalFileDataset setDeleteFilesOnClose(boolean deleteFilesOnClose)
      Indicates if this dataset's files are deleted when it is closed (defaults to false).
    • getRecordsPerFile

      public int getRecordsPerFile()
      Returns the number of records stored in each data file (defaults to 1000).
    • setRecordsPerFile

      public LocalFileDataset setRecordsPerFile(int recordsPerFile)
      Sets the number of records stored in each data file (defaults to 1000); only allowed before the data is loaded.
    • close

      public void close()
      Specified by:
      close in interface AutoCloseable
      Specified by:
      close in interface Closeable
      Overrides:
      close in class Dataset
    • getRecordCount

      public long getRecordCount()
      Description copied from class: Dataset
      Returns the number of records loaded into this dataset so far.
      Specified by:
      getRecordCount in class Dataset
    • getRecord

      public Record getRecord(long index)
      Description copied from class: Dataset
      Returns the loaded record at the given 0-based index.
      Specified by:
      getRecord in class Dataset
    • getRecordList

      public RecordList getRecordList(long offset, int count)
      Description copied from class: Dataset
      Get a subset of the records cached in this dataset.
      Overrides:
      getRecordList in class Dataset
    • getColumnCount

      public long getColumnCount()
      Specified by:
      getColumnCount in class Dataset
    • getColumnNames

      public List<String> getColumnNames()
      Specified by:
      getColumnNames in class Dataset
    • getColumn

      public Column getColumn(int index)
      Specified by:
      getColumn in class Dataset
    • getColumn

      public Column getColumn(String name)
      Description copied from class: Dataset
      Returns the column with the given name or null if there is none.
      Specified by:
      getColumn in class Dataset
    • getOrCreateColumn

      protected Column getOrCreateColumn(String name, int index)
      Description copied from class: Dataset
      Returns the named column, creating it with the given 0-based field index if needed; called while collecting column stats.
      Specified by:
      getOrCreateColumn in class Dataset
    • getColumns

      public List<Column> getColumns()
      Specified by:
      getColumns in class Dataset
    • beforeLoad

      protected void beforeLoad()
      Description copied from class: Dataset
      Called at the start of the data loading process, but before any records or column stats have been loaded.
      Specified by:
      beforeLoad in class Dataset
    • afterLoad

      protected void afterLoad()
      Description copied from class: Dataset
      Called at the end of the data loading process after all the records and column stats have been loaded.
      Specified by:
      afterLoad in class Dataset
    • afterColumnStatsLoaded

      protected void afterColumnStatsLoaded()
      Description copied from class: Dataset
      Called during the data loading process after all the column stats have been loaded. The records would have already been loaded when this method is called since column stats require additional processing.
      Overrides:
      afterColumnStatsLoaded in class Dataset
    • createDataWriter

      protected DataWriter createDataWriter()
      Description copied from class: Dataset
      Writes records to this dataset's cache after clearing it.
      Specified by:
      createDataWriter in class Dataset
    • load

      public LocalFileDataset load(Integer maxRecords, JobCallback<DataReader,DataWriter> callback)
      Description copied from class: Dataset
      Starts the asynchronous loading of records from the pipeline into this dataset. This method returns immediately and does not wait for loading to complete. See Dataset.waitForRecordsToLoad() and Dataset.waitForRecordsToLoad(long, long).
      Overrides:
      load in class Dataset
      Parameters:
      maxRecords - the maximum records to load or null to load all records.
      callback - the object to notify as data is being loaded.
    • addExceptionProperties

      public DataException addExceptionProperties(DataException exception)
      Description copied from class: FoundationObject
      Adds this object's current state to a DataException. Since this method is called whenever an exception is thrown, subclasses should override it to add their specific information.
      Specified by:
      addExceptionProperties in interface DataExceptionContributor
      Overrides:
      addExceptionProperties in class FoundationObject