Use a Regex Rule to Detect Records

This example shows how to tell the record detector that keys matching a regular expression are really rows. Calling addRegexNodeNameRule("row\\d+") on the TreeDetectionStrategy adds a RegexNodeNameRule that is matched against the whole node name; row1, row2 and row3 are classified as dynamic names, collapse into one repeating node and are read as three records. The key is kept as the sheet_id field. The method can be called more than once with different expressions.

Input JSON file

{
    "sheet": {
        "row1": { "col1": "Alice", "col2": 30 },
        "row2": { "col1": "Bob", "col2": 25 },
        "row3": { "col1": "Carol", "col2": 41 }
    }
}

Java Code Listing

package com.northconcepts.datapipeline.foundations.examples.tree;

import java.io.File;

import com.northconcepts.datapipeline.core.DataWriter;
import com.northconcepts.datapipeline.core.StreamWriter;
import com.northconcepts.datapipeline.foundations.pipeline.tree.Tree;
import com.northconcepts.datapipeline.foundations.pipeline.tree.detect.TreeDetectionStrategy;
import com.northconcepts.datapipeline.job.Job;
import com.northconcepts.datapipeline.json.JsonReader;

public class UseARegexRuleToDetectRecords {

    public static void main(String[] args) {
        File inputFile = new File("example/data/input/tree/regex_named_rows.json");

        TreeDetectionStrategy strategy = TreeDetectionStrategy.standard(true)
                .addRegexNodeNameRule("row\\d+");

        Tree tree = Tree.loadJson(inputFile, strategy);

        JsonReader reader = new JsonReader(inputFile);
        tree.getAllFields().forEach(node -> reader.addField(node.getFieldName(), node.getXpathExpression()));
        tree.getAllRecordBreaks().forEach(node -> reader.addRecordBreak(node.getXpathExpression()));

        DataWriter writer = StreamWriter.newSystemOutWriter();

        Job.run(reader, writer);
    }

}

Code Walkthrough

  1. inputFile points at regex_named_rows.json.
  2. TreeDetectionStrategy.standard(true) returns the default detection strategy; the true argument turns on tree optimization, which merges dynamically named nodes into a single wildcard node while the tree is being built. addRegexNodeNameRule("row\\d+") adds the RegexNodeNameRule; the Java string escapes the backslash, so the expression is row\d+.
  3. Tree.loadJson(inputFile, strategy) loads the document into a Tree with the three rows merged into one node, which the detection passes mark as the record break with col1 and col2 as fields.
  4. A JsonReader is created on the same file and configured from the tree: tree.getAllFields() supplies the name and XPath expression of every detected field to addField(). sheet_id holds the original key.
  5. tree.getAllRecordBreaks() supplies the XPath of the record node to addRecordBreak().
  6. Job.run() streams the records to a StreamWriter on the console.

Console Output

-----------------------------------------------
0 - Record (MODIFIED) {
    0:[sheet_id]:STRING=[row1]:String
    1:[col1]:STRING=[Alice]:String
    2:[col2]:LONG=[30]:Long
}

-----------------------------------------------
1 - Record (MODIFIED) {
    0:[sheet_id]:STRING=[row2]:String
    1:[col1]:STRING=[Bob]:String
    2:[col2]:LONG=[25]:Long
}

-----------------------------------------------
2 - Record (MODIFIED) {
    0:[sheet_id]:STRING=[row3]:String
    1:[col1]:STRING=[Carol]:String
    2:[col2]:LONG=[41]:Long
}

-----------------------------------------------
3 records
Mobile Analytics