Skip to content

Parallel Execution & Isolation

tiderace runs your tests itself — there is no pytest at runtime. Execution is built on three ideas: a warm wellspring that imports your project once, a parallel pool of wellsprings (one per core) fed by a locality-aware scheduler, and an isolation ladder that isolates each test the cheapest sound way.

The warm wellspring & fork model

A Wellspring (ADR-E003) is a CPython process that imports your project once — numpy, conftest, your modules — and then runs tests inside that already-warm interpreter. On the fork path, each test runs as a fork()ed copy-on-write child, so it gets a pristine view of imported state for roughly the cost of a fork, not a fresh python -c.

flowchart TB
    L["launch wellspring"] --> I["import project ONCE<br/>(numpy · conftest · user modules)"]
    I --> RDY["ready (warm)"]
    RDY --> REQ{"ExecRequest<br/>per test"}
    REQ -->|"fork path"| FK["fork() → COW child<br/>run body → report → _exit"]
    REQ -->|"no-fork path"| NF["run in-process<br/>(snapshot / restore)"]
    FK --> REQ
    NF --> REQ
  • Import once, fork many — the warm import is the expensive part; COW children share it.
  • Per-test deadline — a child exceeding its deadline is killed and reported Error. One deadline, DEFAULT_DEADLINE_MS (60 s, --timeout to change it), for tiderace run and every daemon mode alike: the daemon used to hand its pool 5 s, and a class whose set-up ran two worker interpreters timed out under the daemon only (TID-89). On the in-process tiers the same deadline is a SIGALRM armed around the test: it ends any wait CPython lets a signal interrupt (a lock, a sleep, a socket read), the node is reported as a timeout error and forks from the next run on. A wait the signal cannot reach leaves the worker silent, and the engine's read on it gives up ten seconds past the deadline: the worker is killed, the in-flight node is reported as the fault, the rest of its unit as not run, and the other workers finish the run (TID-93). Before that a test blocking on the in-process tier blocked the whole run, for good.
  • Fixture scopes — the shim keeps wider-scope fixtures live in the warm image and tears them down in reverse as a worker moves between modules and classes (tiderace_shim/engine.py).

What the engine does not promise

Two things a pytest suite may lean on without noticing, stated here so a divergence is read for what it is.

Execution order across modules. A module's tests run in one process, in file order, as pytest runs them (TID-80): a file whose tests hand each other state — through a fixture, a library, a mock's backend or the module's own globals — works as it does under pytest. Nothing is put back between a file's tests; what is put back is the module: when a worker leaves a file for the next, that file's globals, os.environ, sys.modules, the working directory, sys.path, the logging and warnings state, and the entries its tests added to library registries are restored to what the worker found on entering it (TID-81). So the next module starts clean, and the file itself behaves as its author saw it under pytest. What is not promised is the order between modules, or that two modules share a process: modules are units handed to workers from a queue, heaviest first. A test that asserts on what an earlier module left behind — the canonical case is assert "chromadb" not in sys.modules, true only if nothing before it imported the package — is order-dependent under pytest too (pytest -p randomly breaks it the same way), and the fix belongs in the test. When such a failure mentions sys.modules, the engine appends a line saying so rather than leaving a bare AssertionError to read as the runner's bug. This was the single divergence on a 4,652-test suite in the benchmark; with the file now run in its own order that test precedes the import that broke it and the suite agrees with pytest exactly, but the class of difference remains real and is not promised away. --shard-modules gives the old behaviour back — a file heavier than one worker's share split across workers, every core busy on a single-file suite, and such files may break — the trade pytest-xdist makes between --dist loadfile and --dist load.

Isolation inside a module. On the fork tier — a module holding a global that cannot be snapshotted, such as a client or a lock — the module's tests run sequentially in one forked child rather than one child each (TID-80). The child is the boundary between that module and the rest of the run, which is what opacity requires; inside it the file behaves as under pytest, including a generator one test advanced being advanced for the next. A child that dies mid-file is reported on the test that killed it and the file's remaining tests run in a fresh one.

A run shorter than its longest test. Units are drained from a shared queue heaviest-first, and on the second run (durations recorded) the schedule reaches its ideal: on pirn-core seven of eight workers finish within 0.1s of each other while the eighth runs one 22.9s test from t=0, and the wall is that test plus ~3s of start-up. tiderace run --report records each node's worker, unit and unit start/end, and benchmarks/harness/timeline.py draws them, so a slow run can be read as what it is — a schedule, a start-up, or a test — rather than guessed at (TID-78).

Being a plugin host. Parametrisation a pytest plugin injects — anyio's backends — is expanded so the node ids match, and the fixtures a plugin defines (mocker, anyio_backend_name) are registered at the lowest precedence, as ordinary fixture functions in an importable module, which is all they are (TID-87). The plugin's hooks never run. A suite whose purpose is to test a pytest plugin through pytester is testing pytest, and running it means becoming pytest. See 12-plugin-host for the boundary.

The parallel pool

The runner (engine-core/runner/run.rs) runs N workers, one per core; on Unix they are forked off one warm image (WellspringPool, engine-core/exec/tiers/pool.rs), which the daemon keeps between runs (engine-daemon/warm_image.rs). The LocalityScheduler (ADR-E010) packs work into per-worker batches with two goals at once:

flowchart LR
    T["tests + weights<br/>(timing history or equal)"] --> G["group by locality key<br/>(module = file part of node id)"]
    G --> LPT["LPT bin-packing<br/>(longest-processing-time first)<br/>across N workers"]
    LPT --> B["WorkerBatches<br/>(balanced, module-coherent)"]
  • Scope locality — a module's tests land on the same worker, so its module/session fixtures are set up once.
  • Load balance — LPT (longest-processing-time-first) greedy packing keeps workers evenly busy.

A RoundRobinScheduler exists as a simpler baseline.

Each batch runs on the platform's isolation tier, chosen once per run (WorkerStrategy::factory, engine-core/exec/tier.rs):

  • Unix — a ForkWorker: one warm wellspring, fork-per-test (the model above).
  • Windows — no fork(), so a SubprocessWorker runs the batch no-fork (in-process, with snapshot/restore between tests; opaque modules are refused rather than run without isolation). Parallelism still comes from N batches on N threads — one process per batch. This is what lets the parallel pool, and run --all, work on Windows at all.

The warm image (TID-84)

The pool's parent — the one process that imported the suite, from which the N workers are forked — used to exit with the run, so every tiderace run paid the import again. With tiderace daemon start, the daemon launches the parent persistent (WellspringPool::launch_persistent): it imports once, reports ready, and then answers spawn requests over its stdin/stdout for as long as the daemon holds it. Each RunFull request asks it for N fresh workers (spawn_workers), which connect back over a Unix socket and serve that run only; the parent is untouched by any of them, so the next run forks from the same clean image — no start-up, no imports, no state carried over.

What decides whether the image still describes the tree is a stamp over every .py file and pytest config file (pytest.ini, pyproject.toml, tox.ini, setup.cfg — the shim reads addopts at start-up) under the root: path, size, mtime, taken per request. A changed stamp drops the parent. A full run then launches a new one, which is the full import a full run pays anyway; an impacted run on a changed tree uses the one-shot pool and its selective import instead (TID-75), which is cheaper than re-importing everything into an image it may not need. So the warm image pays off for runs that change nothing — re-runs, gates on an unchanged tree, and -k runs of one test by name, whose selection travels with the request and is applied by the workers after the fork (TID-90) — and the source-edit inner loop stays where TID-75 put it.

The isolation ladder

We isolate tests from each other so one can't corrupt another's view of process-global state. The classic mechanism is fork() per test — but the fork (~4.5 ms) was the dominant cost, and most tests don't mutate shared state at all, so the fork buys them nothing. tiderace classifies each test and runs it the cheapest sound way. This is automatic (ADR-E014); there is no user flag.

flowchart TD
    START["test to run"] --> STATIC{"static pre-filter (AST):<br/>obviously mutates<br/>shared state?"}
    STATIC -->|"yes (global / env /<br/>process-global call)"| FORK
    STATIC -->|"no obvious impurity"| RESTORABLE{"module<br/>snapshot-restorable?<br/>(no opaque globals)"}
    RESTORABLE -->|"no (opaque globals)"| FORK["FORK<br/>COW child<br/>~4.5 ms · bulletproof"]
    RESTORABLE -->|"yes"| KNOWN{"known pure?<br/>(recorded verdict)"}
    KNOWN -->|"yes"| BARE["BARE NO-FORK<br/>run in-process, no snapshot<br/>~0.05 ms per trivial test"]
    KNOWN -->|"unknown / impure"| RESTORE["NO-FORK + RESTORE<br/>snapshot → run → undo<br/>~0.4–0.9 ms (5–14×)"]
    RESTORE --> VERIFY["purity guard verifies<br/>(records verdict for next time)"]
    BARE --> DONE["outcome + coverage + purity"]
    VERIFY --> DONE
    FORK --> DONE
Tier When Isolation mechanism Per-test cost ¹
bare no-fork test is known pure (recorded verdict) nothing to isolate ~0.05 ms (90×)
no-fork + restore restorable footprint, purity unknown/impure deep-copy snapshot of module globals + os.environ on entering the module, per-test verdict, restore on leaving it ² ~0.4–0.9 ms (5–14×)
fork module has opaque (un-deep-copyable) globals copy-on-write child ~4.5 ms (1×)

² A global whose type compares by identity (no __eq__) is snapshotted as itself, not deep-copied: a copy of it could never compare equal, so every test in its module read as impure. from __future__ import annotations binds one such global (annotations) in almost every module — on pirn-core it accounted for 4,491 of 4,499 impure verdicts and left 16 tests recorded pure; with the identity rule 4,488 are (TID-77). The verdict still catches a rebinding; mutation inside such an object is left to the fingerprint, as it always was.

The per-test snapshot is a verdict, not a restore, since TID-81: it says what each test touched (the bare tier and --report read it), and the restore happens once, at the module boundary. The one thing no boundary can put back is a thread a test left running; that test is re-run in the clean room and forks from then on, as before.

¹ Microbenchmark figures — one trivial test, against a fork from a light parent. They show the shape of each tier's overhead, not what a suite will see, and both halves of the ratio move:

  • The fork baseline scales with the parent. fork() copies page tables, so on a large-import project it is far more than 4.5 ms — ~29 ms per test against a 66 MB parent. Every multiplier above is relative to the light-parent figure.
  • The bare tier's reach depends on the suite, and is usually small. It engages only for tests that mutate nothing. Measured on a snapshot-heavy corpus of genuinely pure tests it delivers ~3.4× over no-fork + restore; on a real 4,545-test suite built with module-level test doubles and registries, no test at all measures pure, so the tier never engages. Suites of self-contained assertions benefit; suites that record into shared state — most suites worth optimising — do not. No-fork + restore is the tier real suites actually spend their time in.

Key properties:

  • Sound by construction. No-fork + restore contains mutation rather than predicting it; a non-restorable module always falls back to fork (isolation._restorable()). Correctness never depends on the purity verdict — the verdict is only an optimization that lets a known-pure test skip the snapshot.
  • No learning pass. Restore works on the very first run; the purity guard records verdicts as a free side effect of running, so subsequent runs can promote pure tests to the bare tier.

The daemon enables this by default: it sets TIDERACE_RESTORE=1 and requests no-fork on every test; the shim downgrades to fork only where unsound. TIDERACE_FORCE_FORK=1 reverts to fork-per-test as a debug / benchmark baseline only — it is not a user-facing tuning flag.

The sub-interpreter tier (Windows parallelism)

The ladder above removes the fork tax, but on Windows there is no fork() at all — so the pool falls back to the no-fork SubprocessWorker, which is sequential within each process. A pure-Python suite that flies in parallel on Linux runs one-test-at-a-time on Windows. The sub-interpreter tier (ADR-E015) is how tiderace gets parallel no-fork execution there.

Since CPython 3.14, concurrent.interpreters (PEP 734) exposes multiple interpreters in one process, each with its own GIL (PEP 684) — so they run Python genuinely in parallel across cores, no fork. The catch is that not every module can be imported into an isolated sub-interpreter: numpy's C core, for one, refuses to load in a sub-interpreter (which rules out pandas/scipy/torch with it). So the tier is conditional — detect first, route accordingly:

flowchart TD
    START["run --all<br/>TIDERACE_SUBINTERP=1"] --> PROBE["probe each module<br/>(import in an isolated<br/>sub-interpreter — safe?)"]
    PROBE --> PART{"module<br/>sub-interpreter-safe?"}
    PART -->|"yes (pure-Python /<br/>stdlib / sub-interp-friendly)"| SI["SubInterpWorker<br/>parallel pool, per-interpreter GIL<br/>no fork"]
    PART -->|"no (numpy &c.)"| REST["fork pool (Unix) /<br/>no-fork SubprocessWorker (Windows)"]
    SI --> OUT["outcomes"]
    REST --> OUT
  • Detect — tiderace-daemon probe imports each module in a throwaway isolated sub-interpreter and records safe / unsafe. The verdict is content-addressed and cached (.tiderace-state.json), so a module is re-probed only when its content changes — the same pattern as purity verdicts.
  • Route — safe modules go to the SubInterpWorker pool (parallel, no fork); everything else takes the ordinary fork/no-fork pool. A mixed suite gets partial parallelism: its pure-Python modules run in parallel, its numpy modules run the ordinary way.
  • Sound by the same rule as the ladder — a sub-interpreter has its own module dict and its own os.environ, so tests in the pool can't leak state into one another; anything undeterminable is treated as unsafe and routed away. The tier is verified result-identical to the fork pool on the safe subset.

It is opt-in (TIDERACE_SUBINTERP=1 on run --all; TIDERACE_SUBINTERP_WORKERS sizes the pool) because its payoff is Windows-specific: on Linux the fork pool already parallelizes, so the tier measures at parity there and buys nothing. Requires CPython 3.14+ (concurrent.interpreters); on older interpreters probe reports unknown and callers fall back to fork. See the CLI reference for the probe mode and the TIDERACE_SUBINTERP* env vars.

The transport seam

Execution reaches Python through one trait, ShimTransport (ADR-E011): send an ExecRequest, block for an ExecResponse. In production this is PipeTransport — length-prefixed JSON frames over the wellspring's pipes. An experimental InProcessTransport (②, ADR-E013) drives an embedded CPython over PyO3 FFI with no subprocess. The engine never knows which backend it's talking to. See ARCHITECTURE.md for the seam diagram.