All Implemented Interfaces:
DataExceptionContributor, JsonSerializable, RecordSerializable, XmlSerializable, Closeable, Serializable, AutoCloseable, Iterable<Record>

public class MvStoreDataset extends Dataset
Caches the dataset's records on disk using MVStore.
See Also:
  • Constructor Details

    • MvStoreDataset

      protected MvStoreDataset(File databaseFile, AbstractPipeline pipeline)
      Creates a dataset that, when loaded, stores the pipeline's records in the given MVStore file.
    • MvStoreDataset

      protected MvStoreDataset(File databaseFile)
      Opens the dataset saved in the given MVStore file, with its records and column stats already loaded.
  • Method Details

    • createTempDataset

      public static MvStoreDataset createTempDataset(AbstractPipeline pipeline)
      Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits. The temporary file will be created in the default temporary-file directory, using the prefix "dataset-" and suffix ".mvstore" to generate its name.
      Parameters:
      pipeline - the source of the dataset's data
    • createTempDataset

      public static MvStoreDataset createTempDataset(File databaseFolder, AbstractPipeline pipeline)
      Creates a temporary dataset on disk that will be deleted when the dataset is closed or when the JVM exits. The temporary file will be created in the specified directory (or the default temporary-file directory if null), using the prefix "dataset-" and suffix ".mvstore" to generate its name.
      Parameters:
      databaseFolder - the directory to store the temporary file or null to use the default temporary-file directory.
      pipeline - the source of the dataset's data
    • createDataset

      public static MvStoreDataset createDataset(File databaseFile, AbstractPipeline pipeline)
      Creates a persistent dataset on disk in the specified databaseFile. The file will remain on disk even after the dataset is closed and the JVM exits.
      Parameters:
      databaseFile - the file to store the dataset's contents.
      pipeline - the source of the dataset's data
    • openDataset

      public static MvStoreDataset openDataset(File databaseFile)
      Loads an existing dataset from the specified databaseFile.
      Parameters:
      databaseFile - the existing file to containing the dataset's contents to load.
    • setPipeline

      public MvStoreDataset setPipeline(AbstractPipeline pipeline)
      Overrides:
      setPipeline in class Dataset
    • getDatabaseFile

      public File getDatabaseFile()
    • isDeleteDatabaseFileOnClose

      public boolean isDeleteDatabaseFileOnClose()
      Indicates if the MVStore file is deleted when this dataset is closed (defaults to false).
    • setDeleteDatabaseFileOnClose

      public MvStoreDataset setDeleteDatabaseFileOnClose(boolean deleteDatabaseFileOnClose)
      Indicates if the MVStore file is deleted when this dataset is closed (defaults to false).
    • close

      public void close()
      Specified by:
      close in interface AutoCloseable
      Specified by:
      close in interface Closeable
      Overrides:
      close in class Dataset
    • getRecordCount

      public long getRecordCount()
      Description copied from class: Dataset
      Returns the number of records loaded into this dataset so far.
      Specified by:
      getRecordCount in class Dataset
    • getRecord

      public Record getRecord(long index)
      Description copied from class: Dataset
      Returns the loaded record at the given 0-based index.
      Specified by:
      getRecord in class Dataset
    • getColumnCount

      public long getColumnCount()
      Specified by:
      getColumnCount in class Dataset
    • getColumnNames

      public List<String> getColumnNames()
      Specified by:
      getColumnNames in class Dataset
    • getColumn

      public Column getColumn(int index)
      Specified by:
      getColumn in class Dataset
    • getColumn

      public Column getColumn(String name)
      Description copied from class: Dataset
      Returns the column with the given name or null if there is none.
      Specified by:
      getColumn in class Dataset
    • getOrCreateColumn

      protected Column getOrCreateColumn(String name, int index)
      Description copied from class: Dataset
      Returns the named column, creating it with the given 0-based field index if needed; called while collecting column stats.
      Specified by:
      getOrCreateColumn in class Dataset
    • getColumns

      public List<Column> getColumns()
      Specified by:
      getColumns in class Dataset
    • addField

      protected Column addField(Record record, Field field, int fieldIndex)
      Description copied from class: Dataset
      Adds the field to the stats of its column, creating the column if needed; runs on a column stats worker thread.
      Overrides:
      addField in class Dataset
    • beforeLoad

      protected void beforeLoad()
      Description copied from class: Dataset
      Called at the start of the data loading process, but before any records or column stats have been loaded.
      Specified by:
      beforeLoad in class Dataset
    • afterLoad

      protected void afterLoad()
      Description copied from class: Dataset
      Called at the end of the data loading process after all the records and column stats have been loaded.
      Specified by:
      afterLoad in class Dataset
    • createDataWriter

      protected DataWriter createDataWriter()
      Description copied from class: Dataset
      Writes records to this dataset's cache after clearing it.
      Specified by:
      createDataWriter in class Dataset
    • load

      public Dataset load(Integer maxRecords, JobCallback<DataReader,DataWriter> callback)
      Description copied from class: Dataset
      Starts the asynchronous loading of records from the pipeline into this dataset. This method returns immediately and does not wait for loading to complete. See Dataset.waitForRecordsToLoad() and Dataset.waitForRecordsToLoad(long, long).
      Overrides:
      load in class Dataset
      Parameters:
      maxRecords - the maximum records to load or null to load all records.
      callback - the object to notify as data is being loaded.