Skip to content
SciStack

Labeled multidimensional data

Arrays with named dimensions and coordinates, as in climate, imaging, and simulation output.

Daily global temperature on a 0.25° grid for 40 years is an array of 14,610 days by 721 latitudes by 1,440 longitudes: 15 billion numbers, 61 GB in single precision. By position, you have to remember which axis is which and whether latitude runs north to south, and a slice that picks the Arctic in one product picks the Antarctic in the next. A labeled array carries the names and coordinates of its dimensions with the values, so t.sel(time="2003-08").mean("time") is the mean of August 2003 however the file is laid out.

Arrays like this come from climate models, from time-lapse confocal microscopy stored as time by depth by y by x by channel, from Raman maps with a spectrum at every pixel, from simulations swept over temperature and pressure, and from seismic surveys with a trace for every source and receiver.

In Python, xarray provides DataArray and Dataset, selection by label with sel and by position with isel, arithmetic that aligns and broadcasts by dimension name, groupby("time.month") for climatologies, resample for changing the time step, and open_mfdataset to open a directory of NetCDF files as one dataset. With Dask installed, the data load lazily in chunks and nothing is read until you call compute. In Julia, DimensionalData.jl provides DimArray with selectors such as At and Near, and Rasters.jl builds on it for gridded earth data.

Reduce before you load. Group by month and average while the data are still lazy, and the 61 GB cube becomes a climatology of 12 by 721 by 1,440 values, 50 MB. Start with one DataArray and selection with sel, then broadcasting by name and groupby, then lazy computation over many files. Two-dimensional tables are data wrangling, and the NetCDF and Zarr files these arrays live in are scientific file formats.

What belongs here

Arrays whose dimensions have names and coordinates, such as time by latitude by longitude: selecting and aligning by label, broadcasting by name, groupby and resample along a dimension, and lazy computation on data larger than memory, with xarray in Python and DimensionalData.jl in Julia. Two-dimensional tables belong to data-wrangling; how the data are stored on disk to data-formats.

0 tutorials by type and language

PythonJulia
Concept – –
Tool – –
Recipe – –
Visualization – –
Project – –