An Order-Dependent Test Failure - Traced to a Monkeypatch Nobody Restored
A test passed in isolation and only failed inside the full suite. The culprit was a monkeypatch in another file
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
One particular test kept failing whenever I ran the full regression suite. Run that same test alone, and it always passed.
It reproduced a second time in the same week. That ruled out coincidence — something was systematically wrong.
The symptom
The failing test checked one of the safety gates around fund allocation (I won't go into the specific decision logic here).
Run the whole test suite together, and this test failed. Run it alone, and it passed. Even checking out the pre-change baseline with git stash and running the suite, it passed.
It looked like a classic flaky test. My first guess was timing, a random seed, or some network dependency.
First approach — cast a wide net and try to reproduce by combination
I suspected something touching global state or module-level objects. I shortlisted six files that swapped entries in sys.modules, touched a shared account object, or dealt with in-kind transfers and money-path logic.
I ran those six files in every combination I could think of, trying to reproduce the failure. 166 combinations later, still nothing.
But that process did surface one more clue. The value that showed up on failure was a flat round number — and it had nothing to do with the actual computation this test was supposed to verify (a difference between two legitimate input values). It wasn't a value a logic bug could produce. It looked exactly like a literal constant some monkeypatch had planted.
The actual cause
I reran the full regression suite in its real execution order and narrowed it down from there.
A test in a completely different file was monkeypatching a module-level function that returns an account snapshot, hardcoding it to that same round number. The problem: nothing ever restored it afterward.
In unittest, once a test replaces a module-level function, every other test file that runs afterward in the same process sees that same replaced function — there's no isolation between test modules unless something explicitly tears it down. Only when this test happened to run earlier in the suite did the later safety-gate test end up reading the leftover planted value instead of a real computed one, and fail.
That also explained why running it alone always passed: the test that caused the pollution simply never ran first.
The fix
I added addCleanup calls in both files to restore the original function.
The test doing the patching now saves the original function first, and registers a cleanup to restore it once the test finishes — pass or fail. I added the same kind of guard on the consuming side too, as a second line of defense in case a similar restoration gets missed somewhere else in the future.
After the fix, I reran the full suite in its actual execution order and confirmed the failure no longer reproduced.
Generalizing
Every monkeypatch needs a matching teardown. The moment a test directly overwrites a module attribute or global object, if you don't register a restore via addCleanup (or your framework's equivalent teardown hook), that change outlives the test. It survives in the process and gets inherited by whatever runs next.
"Passes alone, fails in the full suite" is almost always shared mutable state leaking between tests. Before chasing timing or randomness, check whether some test is mutating global or module-level state, and whether anything restores it afterward — that's the faster path.
The failure value itself is a clue. If what comes back on failure isn't a plausible wrong computation but a suspiciously clean literal — the kind of number someone would type into a mock — treat that as a sign of contamination before assuming a logic bug. Recognizing that the number here couldn't possibly come out of the real formula was the turning point in this investigation.
Reproducing the real execution order beat a wide bisect. I first tried bisecting across a shortlist of "files likely to touch global state," 166 combinations deep, and got nowhere. What actually worked was running the full regression in the order it actually runs. A guessed-at subset of candidates was less reliable than reproducing the exact sequence where the failure actually occurs.
More detail on this fund-allocation safety-gate layer is in the live-account deployment writeup(new tab).