Skip to content
SciStack
Tool Python Intermediate 30 min

Parallel runs with concurrent.futures: a parameter sweep on every core

Afterwards you can run independent computations in parallel with concurrent.futures, seed each one reproducibly, and predict when parallelism pays.

Field
Cross-disciplinary
Libraries
matplotlib 3.11.2numpy 2.4.3psutil 7.2.2threadpoolctl 3.7.0
Download notebook Save Mark as done

py-concurrent-futures.ipynb, executed with the versions above. The download needs a free account

Run it yourself. In a terminal, this installs exactly the versions above:

pip install numpy==2.4.3 psutil==7.2.2 matplotlib==3.11.2 threadpoolctl==3.7.0 jupyterlab

The problem: 64 simulations, one busy core

A parameter sweep of 64 simulations, each a fifth to a third of a second, takes 14 to 21 s as a plain loop on the machine that built this page, depending on the run. That machine has four cores, and for all that time three of them sit idle. The simulations do not depend on each other, which is the easy case of parallel computing: concurrent.futures.ProcessPoolExecutor, from Python's standard library, starts a few more Python processes, hands each a share of the runs, and collects the results. With four workers the same sweep takes under 6 s, and every result agrees with the loop's to the last bit.

Each simulation is 2,000 random walkers that start at a wall at 0 and take 8,000 steps drawn from a normal distribution with standard deviation s and mean −b, a drift toward the wall. A walker that would cross the wall is reflected. One run returns one number, the spread of the final positions, and the sweep covers eight step sizes and eight drifts. Sediment grains settling in water, Brownian particles under gravity, and molecules near a surface behave like this, but here the walkers stand in for any simulation that takes a parameter set and returns a result. Swap in your own: the parallel part does not change, and a task that draws no random numbers, such as an ODE solve, needs no seeds.

Speedup over a plain loop against the number of worker processes, both on log scales. Red dots: tasks of a fifth to a third of a second, rising with the number of workers up to the core count and flat beyond it. Blue squares: tasks with a thousand times less work, below 1 at every worker count. A gray line marks the ideal, k workers k times faster.

This is where we end up: each point is one timed sweep, and its height says how many times faster k workers finished it than one loop, on the machine that built this page.

Setup

Imports, seed, parameter grids, core counts, and the style block. The helpers remember and timed serve this page, not you. They time a call with time.perf_counter, the clock behind timeit from the profiling recipe, store the seconds in timings.json to two significant digits, and return the stored value on every later run, so the page shows the same numbers on each rebuild. Delete the file and the notebook measures your machine. The file also records the date, the versions, and the load average: about how many cores other programs kept busy over the last minute. The logical and the physical core count can differ, and Step 5 says why.

import datetime, json, os, platform, time
import multiprocessing as mp
from concurrent.futures import ProcessPoolExecutor, as_completed
from functools import partial
from itertools import product
import numpy as np
import matplotlib.pyplot as plt
import psutil

SEED = 20261009
STEP_SIZES = 0.25 * np.arange(1, 9)      # s, standard deviation of one step
BIASES = 0.01 * np.arange(8)             # b, drift toward the wall per step
CORES = os.cpu_count()                   # logical cores
PHYSICAL = psutil.cpu_count(logical=False) or CORES   # psutil returns None where it cannot tell

plt.rcParams.update({
    "figure.figsize": (7, 3.6), "figure.dpi": 110,
    "axes.spines.top": False, "axes.spines.right": False,
    "axes.grid": True, "grid.alpha": 0.25,
    "font.size": 11, "lines.linewidth": 1.8,
})
INK, ACCENT, SECOND, MUTED = "#1f2a44", "#c8553d", "#2a7f9e", "#8a8f98"

TIMINGS = "timings.json"
if os.path.exists(TIMINGS):
    with open(TIMINGS) as f:
        timings = json.load(f)
else:
    timings = {"date": str(datetime.date.today()),
               "machine": f"{platform.machine()}, {CORES} logical and {PHYSICAL} physical cores, "
                          f"Python {platform.python_version()}, NumPy {np.__version__}",
               "load": f"{os.getloadavg()[0]:.1f}" if hasattr(os, "getloadavg") else "unknown"}

def remember(label, seconds):
    # returns the STORED time if label is in timings.json: delete the file after changing code
    if label not in timings:
        timings[label] = float(f"{seconds:.2g}")
    with open(TIMINGS, "w") as f:
        json.dump(timings, f, indent=1)
    return timings[label]

def timed(label, func):
    """Call func once; return its result and its run time in seconds."""
    t0 = time.perf_counter()
    result = func()
    return result, remember(label, time.perf_counter() - t0)

print(f"{CORES} logical cores, {PHYSICAL} physical cores, {len(STEP_SIZES) * len(BIASES)} tasks")
4 logical cores, 4 physical cores, 64 tasks

Step 1: Put the simulation in a file and run the sweep in a loop

A process pool starts several separate Python programs, called processes, each with its own memory and free to run on a core of its own. The notebook hands them tasks, and each process finds the function it should run by importing it from a file. So the simulation goes into a file now, written by %%writefile, and Step 2 hands that file to the pool:

%%writefile sweep.py
import numpy as np

def final_spread(step, bias, seed, walkers=2_000, steps=8_000):
    """Standard deviation of the walkers' positions after the last step."""
    rng = np.random.default_rng(seed)
    x = np.zeros(walkers)                                   # every walker starts at the wall
    for _ in range(steps):
        x = np.abs(x + rng.normal(-bias, step, walkers))   # a walker that crosses 0 is reflected
    return x.std()
Writing sweep.py

The simulation has two checks built in. Without drift the walkers spread like the absolute value of a free random walk, whose spread after N steps is \(s\sqrt{N(1 - 2/\pi)}\). With strong drift they settle into a layer whose density falls off exponentially with height, like the air in the barometric formula. The layer is \(s^2/(2b)\) thick, and its spread equals its thickness. The last two output lines divide the simulation by these values: ratios near 1 say the code is right before any of it runs in parallel. Each task gets its own child of SeedSequence(SEED).spawn(64), bound to its index, as in Step 5 of the random-numbers tutorial, so a task's result does not depend on which process runs it, or when.

from sweep import final_spread

seeds = np.random.SeedSequence(SEED).spawn(64)           # one stream per task, by task index
task_steps, task_biases = zip(*product(STEP_SIZES, BIASES))

serial, t_serial = timed("serial", lambda: [final_spread(s, b, q)
                                            for s, b, q in zip(task_steps, task_biases, seeds)])
spread = np.array(serial).reshape(8, 8)                   # rows: step size, columns: bias
print(f"loop over 64 tasks: {t_serial:.2g} s, {t_serial / 64:.2f} s per task\n")

print("final spread    b = " + " ".join(f"{b:5.2f}" for b in BIASES))
for s, row in zip(STEP_SIZES, spread):
    print(f"         s = {s:4.2f}     " + " ".join(f"{v:5.1f}" for v in row))

half_normal = STEP_SIZES * np.sqrt(8_000 * (1 - 2 / np.pi))   # no drift
layer = STEP_SIZES**2 / (2 * BIASES[-1])                       # strong drift
print("\nsimulation / theory, b = 0:    " + " ".join(f"{r:.2f}" for r in spread[:, 0] / half_normal))
print("simulation / theory, b = 0.07: " + " ".join(f"{r:.2f}" for r in spread[:, -1] / layer))
loop over 64 tasks: 19 s, 0.30 s per task

final spread    b =  0.00  0.01  0.02  0.03  0.04  0.05  0.06  0.07
         s = 0.25      13.1   3.1   1.5   1.1   0.8   0.6   0.5   0.4
         s = 0.50      27.3  11.8   6.2   4.4   3.2   2.4   2.1   1.8
         s = 0.75      39.0  25.0  13.5   9.3   7.4   5.5   4.7   3.9
         s = 1.00      53.9  37.3  24.7  16.3  12.5  10.3   7.9   7.5
         s = 1.25      66.8  50.4  34.9  25.6  19.4  15.5  12.6  11.2
         s = 1.50      78.7  61.5  48.8  35.8  28.6  21.4  19.0  14.8
         s = 1.75      94.2  77.5  61.5  47.0  39.6  29.4  23.8  22.1
         s = 2.00     106.1  91.8  72.3  62.6  47.5  40.6  31.9  27.2

simulation / theory, b = 0:    0.97 1.01 0.96 1.00 0.99 0.97 1.00 0.98
simulation / theory, b = 0.07: 0.98 1.02 0.97 1.05 1.00 0.92 1.01 0.95

The loop takes 14 to 21 s, a fifth to a third of a second per task. Along each row the spread falls as the drift pulls the walkers toward the wall, and down each column it grows with the step size. Both checks hold: without drift the ratios lie within 4 % of 1, at b = 0.07 within 9 %. An exponential layer has a long tail, and its spread estimated from 2,000 walkers scatters by about 3 % from seed to seed.

Step 2: Hand the sweep to ProcessPoolExecutor.map

ProcessPoolExecutor is the pool. The with block starts the workers and, at its end, waits for every task and shuts the workers down. pool.map(f, xs, ys, zs) takes one iterable per argument of f, like the built-in map, and returns the results in task order, whatever order the workers finish them in. Arguments and results travel between processes pickled, turned into bytes and back.

How a worker starts matters. spawn starts each worker as a fresh, empty Python that knows only what it imports, here sweep. fork, the Linux default before Python 3.14, instead copies the running Python with everything defined in it, notebook functions included, so code that works on Linux can fail on macOS and Windows, where spawn is the default. Asking for spawn explicitly makes this page behave the same on every machine. In a script, put the code that creates the pool under if __name__ == "__main__":, because every spawned worker imports the main script again. A notebook needs no guard.

SPAWN = mp.get_context("spawn")

def map_sweep(workers, task=final_spread):
    with ProcessPoolExecutor(max_workers=workers, mp_context=SPAWN) as pool:
        return list(pool.map(task, task_steps, task_biases, seeds))

pooled, t_pool = timed(f"map {CORES}", lambda: map_sweep(CORES))
print(f"{CORES} workers: {t_pool:.2g} s, {t_serial / t_pool:.1f} times faster than the loop")
print("identical to the loop:", np.array_equal(pooled, serial))
4 workers: 5.3 s, 3.6 times faster than the loop
identical to the loop: True

With four workers the sweep runs at least 2.5 times faster, and the result is identical to the loop's to the last bit, because each task carried its own seed. Your machine prints its own speedup.

Step 3: Collect results as they finish with submit and as_completed

map hands back results in task order only. pool.submit(f, *args) sends one task and at once returns a Future, a placeholder for a result that does not exist yet. A dictionary maps each future to its row and column in the grid. as_completed(futures) yields the futures in the order they finish, and future.result() returns the value, or raises in the notebook the exception the task raised in its worker.

def as_completed_sweep(workers):
    grid = np.empty((8, 8))
    t0 = time.perf_counter()
    with ProcessPoolExecutor(max_workers=workers, mp_context=SPAWN) as pool:
        futures = {pool.submit(final_spread, s, b, q): divmod(n, 8)     # future -> (row, column)
                   for n, (s, b, q) in enumerate(zip(task_steps, task_biases, seeds))}
        for done, future in enumerate(as_completed(futures), start=1):
            grid[futures[future]] = future.result()     # a task's exception is raised here
            if done == 1:
                t_first = time.perf_counter() - t0
            if done % 16 == 0:
                print(f"{done} of 64 done")
    return grid, t_first

(grid, t_first), t_all = timed(f"as_completed {CORES}", lambda: as_completed_sweep(CORES))
t_first = remember(f"first result {CORES}", t_first)
print(f"first result after {t_first:.2g} s, all 64 after {t_all:.2g} s")
print("identical to the loop:", np.array_equal(grid, spread))
16 of 64 done
32 of 64 done
48 of 64 done
64 of 64 done
first result after 0.52 s, all 64 after 5.1 s
identical to the loop: True

The first result arrives in under a second, the whole sweep in under 6 s, as with map. A long sweep can report progress or save partial results long before it ends, and the grid is again identical to the loop's. The output prints only counts, because the order in which the tasks finish changes from run to run.

Step 4: Check whether NumPy already uses every core

A thread is a second line of work inside one process that shares the process's memory, so one process can keep several cores busy. BLAS, the Basic Linear Algebra Subprograms, is the compiled library NumPy hands matrix work to, and it runs on several threads. That covers matrix products (@, np.dot) and np.linalg (solve, eig, inv, svd). Element-wise arithmetic, np.exp and the other ufuncs, sums, and random draws run on one core, which is why final_spread gets no help from BLAS. A second file holds a task that does: three products of a 1,200 × 1,200 random matrix.

%%writefile blas_work.py
import numpy as np
from threadpoolctl import threadpool_limits

def products(seed, n=1_200):
    """Three products of an n x n random matrix: work that NumPy hands to BLAS."""
    a = np.random.default_rng(seed).random((n, n)) / n
    return np.trace(a @ a @ a @ a)

def one_thread():
    threadpool_limits(1)    # no with block: the limit stays until the worker exits
Writing blas_work.py

Called as a plain function, not in a with block, threadpool_limits(1) holds for the rest of the process's life. Passed to the pool as initializer, a function each worker runs once when it starts, it gives every worker a single BLAS thread. threadpool_info() lists the thread pools NumPy has loaded. Sixteen tasks, timed four ways, test whether a pool can lose to a plain loop when NumPy already keeps every core busy: a loop and a pool of four workers, each with BLAS threads free and with one thread.

from threadpoolctl import threadpool_info, threadpool_limits
from blas_work import products, one_thread

blas = [lib for lib in threadpool_info() if lib["user_api"] == "blas"][0]
print(f"BLAS library: {blas['internal_api']}, {blas['num_threads']} threads\n")

blas_seeds = np.random.SeedSequence(SEED).spawn(16)

def blas_loop(threads):
    with threadpool_limits(threads):
        return [products(q) for q in blas_seeds]

def blas_pool(initializer=None):
    with ProcessPoolExecutor(max_workers=CORES, mp_context=SPAWN, initializer=initializer) as pool:
        return list(pool.map(products, blas_seeds))

t_blas = {}
_, t_blas["loop, BLAS threads free"] = timed("blas loop", lambda: blas_loop(None))
_, t_blas["loop, one BLAS thread"] = timed("blas loop 1", lambda: blas_loop(1))
_, t_blas[f"{CORES} workers, BLAS threads free"] = timed("blas pool", blas_pool)
_, t_blas[f"{CORES} workers, one BLAS thread each"] = timed("blas pool 1", lambda: blas_pool(one_thread))
for label, t in t_blas.items():
    print(f"16 tasks, {label:34s} {t:4.1f} s")
BLAS library: openblas, 4 threads

16 tasks, loop, BLAS threads free             1.4 s
16 tasks, loop, one BLAS thread               3.4 s
16 tasks, 4 workers, BLAS threads free        2.4 s
16 tasks, 4 workers, one BLAS thread each     1.2 s

It can. With its BLAS threads free the loop runs more than twice as fast as with one, and four workers with free threads run 4 × 4 = 16 threads on four cores, called oversubscription, taking at least 30 % longer than the loop. One thread per worker cuts the pool's time by a third or more, to within 35 % of the loop's. So if your code spends its time in matrix products or np.linalg, time the plain loop first, and give each worker of a pool one thread; OMP_NUM_THREADS=1, set before NumPy is imported, does the same for every process. On Apple silicon, recent NumPy wheels hand matrix work to Apple's Accelerate, whose threads neither the variable nor threadpoolctl reaches.

Step 5: Measure the speedup against the number of workers

The scan runs the whole sweep in a fresh pool for each number of workers k, the start of the pool included, up to twice the core count to show what surplus workers cost. It does the same for short tasks, the same function with 8 steps instead of 8,000. functools.partial fixes steps=8 and pickles like the function itself. Each short sweep is the best of five runs. The last timing reuses one running pool, to separate starting a pool from passing tasks to it.

WORKERS = sorted({k for k in (1, 2, 3, 4, 6, 8, 12, 16, 24, 32) if k <= CORES} | {CORES, min(2 * CORES, 32)})
short_task = partial(final_spread, steps=8)     # the same function, a thousand times less work

def best_of_five(label, func):
    times = []
    for _ in range(5):
        t0 = time.perf_counter()
        func()
        times.append(time.perf_counter() - t0)
    return remember(label, min(times))

t_short_serial = best_of_five("short serial", lambda: [short_task(s, b, q)
                                                       for s, b, q in zip(task_steps, task_biases, seeds)])
t_long = np.array([timed(f"long {k}", lambda: map_sweep(k))[1] for k in WORKERS])
t_short = np.array([best_of_five(f"short {k}", lambda: map_sweep(k, short_task)) for k in WORKERS])
long_speedup, short_speedup = t_serial / t_long, t_short_serial / t_short
with ProcessPoolExecutor(max_workers=1, mp_context=SPAWN) as pool:   # the first of five runs starts the worker
    t_warm = best_of_five("short warm", lambda: list(pool.map(short_task, task_steps, task_biases, seeds)))
t_start, t_pass = t_short[0] - t_warm, (t_warm - t_short_serial) / 64

print(f"measured {timings['date']} on {timings['machine']}, load {timings['load']}")
print(f"loop: long tasks {t_serial:.2g} s, short tasks {1000 * t_short_serial:.2g} ms\n")
print(" workers    long tasks  speedup    short tasks  speedup")
for k, tl, sl, ts, ss in zip(WORKERS, t_long, long_speedup, t_short, short_speedup):
    print(f"{k:8d} {tl:11.2g} s {sl:8.2f} {1000 * ts:11.0f} ms {ss:8.2f}")
print(f"\nshort tasks, one running worker: {1000 * t_warm:.0f} ms, so starting the pool took "
      f"{t_start:.2f} s and passing a task {1000 * t_pass:.2f} ms")
measured 2026-10-09 on x86_64, 4 logical and 4 physical cores, Python 3.12.3, NumPy 2.4.3, load 0.6
loop: long tasks 19 s, short tasks 18 ms

 workers    long tasks  speedup    short tasks  speedup
       1          19 s     1.00         200 ms     0.09
       2         9.8 s     1.94         230 ms     0.08
       3         6.3 s     3.02         260 ms     0.07
       4         5.3 s     3.58         300 ms     0.06
       8         5.2 s     3.65         500 ms     0.04

short tasks, one running worker: 28 ms, so starting the pool took 0.17 s and passing a task 0.16 ms

The same numbers on log scales, against the ideal of k workers being k times faster:

k = np.array(WORKERS)
fig, ax = plt.subplots()
ax.plot(k, k, color=MUTED, lw=1.1)
ax.axhline(1, color=MUTED, lw=1.1)
ax.axvline(CORES, color=MUTED, lw=1, ls="--")
if PHYSICAL != CORES:
    ax.axvline(PHYSICAL, color=MUTED, lw=1, ls="--")
ax.plot(k, long_speedup, "o-", color=ACCENT, ms=6)
ax.plot(k, short_speedup, "s-", color=SECOND, ms=6)

best = long_speedup.argmax()
ax.annotate(f"{long_speedup[best]:.1f} times faster", (k[best], long_speedup[best]),
            xytext=(6, -16), textcoords="offset points", color=ACCENT)
ax.text(2.1, 1.3, f"tasks of {t_serial / 64:.2f} s", color=ACCENT)
ax.text(1.05, short_speedup[0] * 0.55, f"tasks of {1000 * t_short_serial / 64:.2g} ms", color=SECOND)
ax.text(CORES * 0.92, 6.0, "ideal: k workers, k times faster", color=MUTED, ha="right")
ax.text(k[-1], 1.1, "no gain", color=MUTED, ha="right")
ax.text(CORES * 1.04, 0.45, f"{CORES} cores", color=MUTED)
if PHYSICAL != CORES:
    ax.text(PHYSICAL * 1.04, 0.3, "physical cores", color=MUTED)

ax.set_xscale("log", base=2)
ax.set_yscale("log", base=2)
y_ticks = 2.0 ** np.arange(np.floor(np.log2(short_speedup.min())), np.ceil(np.log2(k[-1])) + 1)
ax.set_xticks(k, [str(v) for v in k])
ax.set_yticks(y_ticks, [f"{v:g}" if v >= 1 else f"1/{1 / v:g}" for v in y_ticks])
ax.minorticks_off()
ax.set(xlabel="number of workers k", ylabel="speedup = loop time / pool time")
plt.show()
Speedup over a plain loop against the number of worker processes, both on log scales. Red dots: tasks of a fifth to a third of a second, rising with the number of workers up to the core count and flat beyond it. Blue squares: tasks with a thousand times less work, below 1 at every worker count. A gray line marks the ideal, k workers k times faster.

Long tasks come within 30 % of the ideal up to the four cores, and past them the speedup stays flat to within 20 %: eight workers take turns on four cores. Short tasks lose at every k, and the last line says why. The loop over them takes 15 to 30 ms, while starting a one-worker pool, a fresh Python that imports NumPy and sweep, takes 0.1 to 0.25 s, and more workers take longer to start. Once the pool runs, each task costs under 0.3 ms more than in the loop, for pickling the task and its result and passing them between processes.

That is enough to predict your own scan. Time one task and multiply by the number of tasks: a pool pays when this loop time is many times the start, and when each task is long against the cost of passing it, a millisecond at the very least. The speedup is then at most the number of physical cores.

os.cpu_count() counts logical cores, and ProcessPoolExecutor() without max_workers starts that many workers. With hyperthreading each physical core appears as two logical ones that share its arithmetic units, so number-crunching tasks gain little past the physical count, which psutil.cpu_count(logical=False) reports. Here the two agree. On a shared cluster node, len(os.sched_getaffinity(0)) gives the cores your job may use (Linux only).

Pitfalls

The same generator in every task. Pass one generator as the seed of four identical tasks. The loop and the pool disagree, and the pool returns one number four times:

rng = np.random.default_rng(SEED)
quick = partial(final_spread, steps=100)
in_loop = [quick(1.0, 0.0, rng) for _ in range(4)]
with ProcessPoolExecutor(max_workers=4, mp_context=SPAWN) as pool:
    in_pool = list(pool.map(quick, [1.0] * 4, [0.0] * 4, [rng] * 4))
print("loop:", " ".join(f"{v:.4f}" for v in in_loop))
print("pool:", " ".join(f"{v:.4f}" for v in in_pool))
loop: 5.9497 5.7333 6.0775 5.9985
pool: 6.1245 6.1245 6.1245 6.1245

default_rng hands a generator back unchanged. In the loop the four tasks draw one after another from one stream, so their numbers differ, but each depends on the tasks that ran before it, and no task can be rerun on its own. The pool pickles the generator in its current state once per task, so every worker draws the same numbers. The fix is Step 1's: one spawned child SeedSequence per task.

A worker function defined in the notebook. Write final_spread in a notebook cell under the name local_spread, hand it to a spawn pool, and the sweep dies with BrokenProcessPool: A process in the process pool was terminated abruptly while the future was running or pending. Above it, every worker's traceback ends in AttributeError: Can't get attribute 'local_spread' on <module '__main__' (<class '_frozen_importlib.BuiltinImporter'>)>. A spawned worker looks the function up by module and name, and the notebook's __main__ is not a module it can import. The fix is the file from Step 1. Under fork the copied process already holds the function, which hides the problem until the same code runs on macOS or Windows.

Returning huge arrays. Keep every walker's full path and one run returns 2,000 × 8,000 float64 values, 128 MB. Sixty-four tasks then pickle 8 GB in the workers and send it to the notebook, which has to hold all of it at once. Reduce in the worker and return the summary, here one number. If you need the arrays, have each task write its own with np.save under a name built from its index, and return the file name.

Variations

  • Thousands of tiny tasks. pool.map(f, xs, chunksize=100) sends the tasks in batches of 100, one pickled message per batch instead of one per task.
  • Threads instead of processes. ThreadPoolExecutor has the same submit and map and starts no new Python. Python lets only one thread at a time run Python code, a rule called the global interpreter lock, so threads help with tasks that wait on files or the network, or that spend their time in compiled code that releases the lock.
  • joblib. Parallel(n_jobs=-1)(delayed(final_spread)(s, b, q) for s, b, q in zip(task_steps, task_biases, seeds)), where n_jobs=-1 asks for one worker per core and delayed packs a function and its arguments into one task. scikit-learn runs its parallel loops this way, and joblib picks the batch size on its own.
  • Many machines. mpi4py.futures.MPIPoolExecutor keeps the submit and map interface and places the workers on the nodes of a cluster.

Cheat sheet

from worker import f, one_thread                        # both live in a file, not in the notebook
seeds = np.random.SeedSequence(SEED).spawn(n_tasks)     # one seed per task; a deterministic f needs none
with ProcessPoolExecutor(max_workers=psutil.cpu_count(logical=False),  # in a script, under if __name__ == "__main__":
                         mp_context=mp.get_context("spawn"),           # the same behavior on every OS
                         initializer=one_thread) as pool:              # or OMP_NUM_THREADS=1 before importing NumPy
    results = list(pool.map(f, xs, ys, seeds))          # either: all results, in task order
    out = [None] * n_tasks                              # or: one future per task, collected as they finish
    futures = {pool.submit(f, x, y, q): n for n, (x, y, q) in enumerate(zip(xs, ys, seeds))}
    for fut in as_completed(futures):
        out[futures[fut]] = fut.result()                # raises the task's exception, if it raised one

Further reading

Was this tutorial helpful? Sign in to tell the author with one click.

Found a mistake, or something unclear? Report a problem (with a free account).

Cite this tutorial

SciStack (2026). Parallel runs with concurrent.futures: a parameter sweep on every core. https://scistack.dev/t/py-concurrent-futures/ (accessed 2026-10-09).

@online{scistack-py-concurrent-futures,
  author  = {{SciStack}},
  title   = {Parallel runs with concurrent.futures: a parameter sweep on every core},
  date    = {2026-10-09},
  url     = {https://scistack.dev/t/py-concurrent-futures/},
  urldate = {2026-10-09},
  note    = {numpy 2.4.3, psutil 7.2.2, matplotlib 3.11.2, threadpoolctl 3.7.0}
}

Tags

as_completedconcurrent.futuresmatplotlibmultiprocessingnumpyparallelprocesspoolexecutorpsutilseedsequencethreadpoolctl

Comments

No comments yet.

Sign in to comment, with a free account.