Building a trajectory dataset (CLI)

This page walks through building a trajectory dataset from a recipe using the anemoi-datasets command line tool. For the recipe syntax itself, see Building your own datasets.

A trajectory dataset stores forecast fields indexed by a base date (the model-run time) and a forecast step, rather than a single validity time. The on-disk array is 5-D (base_dates, variables, ensembles, steps, cells).

Recipe

Set output.layout to trajectories and replace the usual dates: block with two blocks, base_dates: and steps:. The set of samples written on disk is the Cartesian product of the base dates and the steps.

base_dates:
  start: 2021-01-01 00:00:00
  end:   2021-01-02 00:00:00
  frequency: 12h

steps:
  start: 6
  end: 30
  frequency: 6h

input:
  mars:
    type: fc
    # ... source definition ...

output:
  layout: trajectories

Note

base_dates: and steps: are required for layout: trajectories and dates: is rejected. Conversely, for any other layout dates: is required. See Trajectories for the full set of rules.

One-shot creation

The simplest way to build the dataset is the create command, which runs every step in a single process:

anemoi-datasets create dataset.yaml dataset.zarr --overwrite

Incremental / parallel

For large datasets, build the dataset step by step so the loading can be split across processes, terminals or SLURM jobs.

  1. Initialise the (empty) dataset:

    anemoi-datasets init dataset.yaml dataset.zarr --overwrite
    
  2. Load the data in parts. Parts are numbered 1/NN/N (1-based); each part loads a subset of the base dates and can be run in any order and in parallel:

    anemoi-datasets load dataset.zarr --parts 1/10
    anemoi-datasets load dataset.zarr --parts 2/10
    # ... up to ...
    anemoi-datasets load dataset.zarr --parts 10/10
    
  3. Finalise the dataset (merge statistics, write metadata and attributes, clean up temporary files):

    anemoi-datasets finalise dataset.zarr
    
  4. Patch the metadata:

    anemoi-datasets patch dataset.zarr
    

You can follow the progress at any time and clean up leftover temporary files with:

anemoi-datasets inspect dataset.zarr
anemoi-datasets cleanup dataset.zarr

See also