Skip to content
SciStack

Scientific file formats

HDF5, NetCDF, Parquet, and instrument formats: storing and reading scientific data with its metadata.

A simulation that produces a 1000 × 1000 grid of temperatures in double precision holds 8 MB in memory, and np.save writes a file of the same size. np.savetxt at its default format writes 25.5 MB of text. The pandas route, to_csv and back with read_csv, writes 19.6 MB and returns 324,361 of the million numbers changed in their last digits. Text is the wrong home for floating-point results. The detector frames of a beamline, the output of a climate model, the records of a seismometer network, and the vendor files of a mass spectrometer all raise the same question: how to store numbers exactly, with their units and history, in a file that others can still open years later.

In Python, h5py reads and writes HDF5. A file holds a tree of groups and datasets, each dataset comes back as a NumPy array, and attributes carry units and provenance. NetCDF, the format of climate and ocean science, has been built on HDF5 since version 4 and is read with netCDF4. Tables go to Parquet through pyarrow, which pandas calls in to_parquet and read_parquet. In Julia, HDF5.jl offers h5open, h5read, and h5write, and Arrow.jl reads the Arrow files that pyarrow and the to_feather method of pandas write, so a table crosses between the languages without CSV.

Start with HDF5: write one array with its unit as an attribute, read it back, and compare it bit for bit. Then chunking and compression, which let you read a slice of a dataset too large for memory, NetCDF for gridded data, and Parquet for tables. Convert a vendor format to an open one once and keep the converter. A file that only the instrument's software opens lasts as long as the license. Loaded data go on to multidimensional data or data wrangling.

What belongs here

How scientific data are stored and read back: binary formats such as HDF5, NetCDF, and Parquet, their metadata and units, reading instrument and vendor formats, compression and chunking, and choosing a format that others can open in twenty years. Text tables and what to do with them once loaded belong to data-wrangling.

0 tutorials by type and language

PythonJulia
Concept – –
Tool – –
Recipe – –
Visualization – –
Project – –