Skip to content

Latest commit

 

History

History
201 lines (133 loc) · 8.26 KB

File metadata and controls

201 lines (133 loc) · 8.26 KB

The Matching Part

The matching part is the third and last part of a csvpath. It comes behind the root and the scanning part. Like the scanning part, it is bracketed. The matching part is where validation and/or data upgrading happens.

Match Components

The matching part is built from space separated "match components" that are ANDed or ORed together. When a match component "votes" that a line is a match, its value is True. That means as far as that match component is concerned, the line matches.

Csvpaths are configured to logically AND match components together by default. When match components are ANDed, all the components must match for the line to match as a whole. When the components are ORed, if any one match component matches the line is considered to match.

Csvpaths may be configured so that matched lines are:

  • Considered valid or invalid
  • Collected as output from evaluation

The match components' order is meaningful. Generally, components are tested left to right, top to bottom. The exception to this order is if a match component is set to only match if all other components match. In that case, that match component must come last. If all match components are set to only match if all other match components match, the order of all the components is preserved.

Types of Match Components

A match component is one of these types:

These components can be combined in endless ways. The organization of a csvpath's match part is [x x x x] where each x is a match component. All of the match components are ANDed or ORed together. There can be any number of match components in a csvpath statement.

Since equalities are match components, [ "x" == "y" "z" == "z" ] is a legal csvpath matching part containing two top-level match components. Each of those two components is an Equality. Each Equality holds two Term component literals, "x" and "y", and "z" and "z".

In this case, if the csvpath ANDs match components, the default, this statement will never match because "x" never equals "y". If evaluation is switched to OR, the statement will always match because "z" always equals "z". The switch from AND to OR can be done programmatically or, more typically, using a mode declaration in a csvpath comment. Modes are covered in the page on comments.

Qualifiers

Some of these component types can be modified with qualifiers. A qualifier changes the behavior of a match component. It is set by adding a dot-name to the match component name.

For example, count.cars(#color=="blue") is a variation on count(#color=="blue"). The difference is that behind the scenes the count() function's variable is named cars, rather than a random string. Likewise count.cars.onmatch(#color=="blue") increments the count of the cars variable only if the rest of the line matches.

Read more about qualifiers here.

Testing Required

There is no limit to the functionality you can include in a single csvpath using match components. However, functions have different performance characteristics. You should test both the performance and functionality of your paths, just as you would when working with SQL or another language.

Match Component Types

Term

A string, number, or regular expression value.

Returns Matches Examples
A value Always matches
  • "a value"
  • 3

Read about terms here.

Function

A composable unit of functionality called once for every row scanned. CsvPath Validation Language has over 150 functions. As a language, it is very much functions-oriented.

Returns Matches Examples
Calculated Calculated count()

Read about functions here.

Variable

A value that is set or retrieved once per row scanned. Generally, variables last from the line they are created on through the end of the file. Variables are available at the end of evaluation.

Returns Matches Examples
A value True when set. (Unless the onchange qualifier is used). Alone a variable is an existence test. @firstname

Read about variables here.

Header

A named header or a header identified by its 0-based index. (CsvPath avoids the word "column" for reasons we'll go into later).

Returns Matches Examples
A value Calculated. Used alone it is an existence test. #area_code

Read about headers here.

Equality

Two of the other match component types joined with an "=" or "==" or the when-do operator ->.

Returns Matches Examples
Calculated Calculated #area_code == 617
Calculated When-do matches when left side matches #area_code -> print("area code is $.headers.area_code")
No value Assignment always matches @code = #area_code

Reference

References are a way of pointing to data generated by csvpaths. Referenced data comes from the currently running csvpath or from named-results.

Named-results are the results of past csvpath runs. The name is the same as the named-paths group that generated it. Named-path groups and results are higher-level CsvPath Framework concepts covered on csvpath.org.

As match components, references can point to:

  • Variables
  • Headers
  • Csvpath runtime indicators

A csvpath runtime indicator is a metric or fact about the running csvpath. These include things like the line count, the start time, if the file is considered valid, etc.

References are also used more broadly outside of csvpaths for data and run management. That topic is also covered on csvpath.org.

The form of a reference is:

    $my-name.variables.firstname

This reference looks for results named my-name. The keyword variables indicates the value is the firstname variable.

Returns Matches Examples
Calculated True, if used in an assignment, otherwise calculated. @q = $orders.variables.quarter

References in print statements

The most common use for references in csvpaths is within the print() function. In this use, the reference is used by a match component, but it is not itself a match component. Nevertheless, because it is common we cover it briefly here.

Print function references are local. This means that they refer to the csvpath that contains them. Because a print reference is local, the syntax is shortened by not specifying the name.

References in the print() function can point to user-defined metadata, as well as variables, headers, and runtime indicators.

A print() variable reference looks like

    print("$.variables.firstname")

A print() header reference looks like:

    print("$.headers.lastname")

A print() runtime indicator looks like:

    print("$.csvpath.line_number")

And a print() metadata field looks like:

    print("$.metadata.my_description")

Read more about references here.