diff --git a/docs/explanation/processing.md b/docs/explanation/processing.md index d1ef69a1a..7c862cb11 100644 --- a/docs/explanation/processing.md +++ b/docs/explanation/processing.md @@ -61,9 +61,10 @@ is still processing (or scheduled to be processed). In most scenarios, entry processing is not triggered individually, but as part of an upload processing. Many entries of one upload might be processed at the same time. Some order -can be enforced through *processing levels*. Levels are part of the parser metadata and -entries paired to parsers with a higher level are processed after entries with a -parser of lower level. See also [how to write parsers](../howto/plugins/types/parsers.md). +can be enforced through the *execution order* of parsers. The execution order is part of the +parser entry point and entries paired to parsers with a higher execution order are processed +after entries with a parser of lower execution order. See also +[how to control parser execution order](../howto/plugins/types/parsers.md#control-parser-execution-order). ## Customize processing diff --git a/docs/howto/plugins/types/normalizers.md b/docs/howto/plugins/types/normalizers.md index 8c58ae829..9d23e5c1c 100644 --- a/docs/howto/plugins/types/normalizers.md +++ b/docs/howto/plugins/types/normalizers.md @@ -1,6 +1,6 @@ # How to create a normalizer -A normalizer takes the archive of an entry as input and manipulates (usually expands) the given archive. This way, a normalizer can add additional sections and quantities based on the information already available in the archive. All normalizers are executed in the order [determined by their `level`](#control-normalizer-execution-order) after parsing, but the normalizer may decide to not do anything based on the entry contents. +A normalizer takes the archive of an entry as input and manipulates (usually expands) the given archive. This way, a normalizer can add additional sections and quantities based on the information already available in the archive. All normalizers are executed in the order [determined by their `execution_order`](#control-normalizer-execution-order) after parsing, but the normalizer may decide to not do anything based on the entry contents. This documentation shows you how to create a plugin entry point for a normalizer. You should read the [introduction to plugins](../plugins.md) to have a basic understanding of how plugins and plugin entry points work in the NOMAD ecosystem. @@ -113,18 +113,21 @@ Here, we used the schema definition for the `run` section defined in this [plugi ## Control normalizer execution order -`NormalizerEntryPoints` have an attribute `level`, which you can use to control their execution order. Normalizers are executed in order from lowest level to highest level. The default level for normalizers is `0`, but this can be changed per installation using `nomad.yaml`: +`NormalizerEntryPoints` have an attribute `execution_order`, which you can use to control their execution order. Normalizers are executed in order from lowest to highest execution order. The default execution order for normalizers is `0`, but this can be changed per installation using `nomad.yaml`: ```yaml plugins: entry_points: options: "nomad_example.normalizers:mynormalizer1": - level: 1 + execution_order: 1 "nomad_example.normalizers:mynormalizer2": - level: 2 + execution_order: 2 ``` +!!! note + `execution_order` replaces the deprecated `level` attribute. Existing configurations that use `level` keep working, but should be migrated to `execution_order`. + ## Running the normalizer If you have the plugin package and `nomad-lab` installed in your Python environment, you can run the normalization as a part of the parsing process using the NOMAD CLI: diff --git a/docs/howto/plugins/types/parsers.md b/docs/howto/plugins/types/parsers.md index cd6d4234f..44067c682 100644 --- a/docs/howto/plugins/types/parsers.md +++ b/docs/howto/plugins/types/parsers.md @@ -110,6 +110,59 @@ myparser = MyParserEntryPoint( You can find all of the available matching criteria in the [`ParserEntryPoint` reference](../../../reference/plugins.md#parserentrypoint) +### Control parser matching order + +Each file is assigned to at most one parser. NOMAD checks the parsers one by one and the first parser whose matching criteria fit the file is used, even if other parsers would also match it. When several parsers can match the same files, for example a generic parser and a more specialized one, you can use the `matching_order` attribute to decide which parser gets to claim them first. + +Parsers with a lower `matching_order` are checked first. Parsers with the same `matching_order` are checked in alphabetical order of their entry point id. The default value is `0`. If all parsers use the default value, the order in which the parsers were registered is used instead. You can set the matching order in the entry point: + +```python +myparser = MyParserEntryPoint( + name='MyParser', + description='My custom parser.', + mainfile_name_re='.*\.myparser', + matching_order=-1, +) +``` + +or change it per installation using `nomad.yaml`: + +```yaml +plugins: + entry_points: + options: + "nomad_example.parsers:myparser": + matching_order: -1 +``` + +Here a negative value makes sure that `myparser` is checked before all parsers that use the default value. + +## Control parser execution order + +`ParserEntryPoints` have an attribute `execution_order`, which you can use to control the order in which entries are processed within an upload. Entries matched by parsers with a higher execution order are processed after all entries matched by parsers with a lower execution order. This is useful when a parser needs to read the results of other entries in the same upload. The default execution order for parsers is `0`, but this can be changed per installation using `nomad.yaml`: + +```yaml +plugins: + entry_points: + options: + "nomad_example.parsers:myparser": + execution_order: 1 +``` + +!!! note + `execution_order` replaces the deprecated `level` attribute. Existing configurations that use `level` keep working, but should be migrated to `execution_order`. + +### Matching order vs. execution order + +The two attributes control different steps of processing and are independent of each other: + +| Attribute | Controls | Step | Effect | +| --- | --- | --- | --- | +| `matching_order` | Which parser is assigned to a file | Matching, when entries are created | Lower values are checked first. The first matching parser claims the file. | +| `execution_order` | When the entries of a parser are processed | Processing, after matching | Lower values are processed first. Entries with higher values wait until the entries with lower values are processed. | + +For example, `matching_order` decides whether a file is parsed by your parser or by another parser that also matches it. It has no effect on when that file is processed. `execution_order` only becomes relevant once the file has been matched to your parser, and it does not change which parser a file gets. + ## Running a parser Parsers automatically run for the matched files within a NOMAD distribution, but it is also possible to run the manually for specific files. This can be useful for testing and for connecting them into external software.