Quant Trading Bot Devlog

한국어로 보기

"[Sep 28-Oct 4] Backup GPU Self-Healing Collapses, Paper-Trading Fill Mismatch Traced"

Backup GPU self-healing collapsed completely on its first night in production and got redesigned, and a paper-trading ledger residual turned out to trace back to a fill mismatch from nine days earlier.

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

This week kept circling back to the gap between "shipped" and "actually works in production."

Backup-GPU self-healing passed 100 reproduction runs clean, then collapsed completely on its first production night. A paper-trading ledger residual that had been written off as "a transient false positive that resolved itself" turned out not to have resolved at all. Both were cases of mistaking quiet surfaces for actual resolution.

Monday: wiring self-healing, dropping a wrong hypothesis

Recovery logic finally got attached to the dead-slot detector built last week(new tab). An overnight reproduction run passing all 100 attempts cleanly brought enough confidence to call it settled.

The same day, a hypothesis that "raising concurrency would reduce blast radius during an incident" got tested against the logs and the premise turned out wrong. Ticker groups and processing slots aren't fixed to each other, so "four groups collapsed at once" was really one slot dying at a moment when four groups happened to be assigned to it. Raising concurrency got shelved, and two real defects surfaced by this investigation — self-healing never wired into the production startup path, and the detector's internal state never resetting after a slot was repaired — got fixed and shipped instead.

Tuesday: the self-healing wired in Monday night collapsed

On its first production night, one slot died and within an hour all four processing slots were dead. Self-healing fired over 200 times but barely recovered anything, and the overnight run stopped having filled less than half its target.

Getting a second opinion from another AI on root cause, Monday's call to hold off on raising concurrency turned out unrelated to this incident. The actual problem was deciding that clearing and reviving a dead slot alone would be enough, which pushed a full server-restart fallback out of scope. Taking that feedback, a call pattern that let one ticker hang for a long stretch got removed, and an escalation path got added — restart the whole server if a healed slot dies again quickly — among five fixes that day.

The morning resume surfaced a separate problem. Restarting manually without disabling a regular-hours safety switch stopped four processing units partway through — the same mistake as a past incident. Missing tickers got backfilled from a logged last-known value, which turned out to belong to a run from days earlier rather than that day, so those tickers ended up re-run in isolation instead.

Wednesday: tracing a paper-trading residual back nine days

A reconciliation residual on one paper-trading(new tab) account, previously written off as "a transient false positive that resolved itself," turned out not to have resolved at all on re-examination. The residual had never shrunk, and the account had been frozen out of trading by a safety guard for days.

The root cause was a defect in the external paper-trading server itself. Order numbers reset from scratch each day on that server, and the fill-lookup code searched back across recent days by number alone without cross-checking ticker or direction, pulling in a fill from a completely different ticker nine days earlier. Price-range validation had passed and there was simply no check comparing quantity or ticker, so a dedicated ticker/direction guard got designed alongside the correction itself.

The same day, an alert registration that had fired pointlessly every day for over a month got retired for good — it had been set up for a condition the current account structure could never satisfy.

Sunday: concurrency increase shelved, a missing output-token cap found

Two higher concurrency tiers went into smoke testing to push model-serving throughput further, and both failed. Only a moderate single-tier increase combined with other serving-config tweaks passed, so that's the setting staying in place for now.

Code review turned up a side effect from the recent move of the backup serving path to a different GPU combination. The watchdog that checks whether the main process is alive judges by process name alone, and once the backup path started spawning similarly-named processes, the watchdog could mistake a healthy backup for the main process and kill it. Several related defects surfaced in the same review and got fixed together.

A ticker analysis that afternoon finished only after producing output several times longer than normal. The first instinct was that the output-token cap needed loosening, but a second opinion from another AI landed on the opposite conclusion.

The current serving config had no output-token cap at all, so the model kept generating until it filled the context window instead of stopping — the reverse of what the instinct suggested. The fix lands a new, modest cap rather than a higher one.


The thread running through this week was telling apart "gone quiet" from "actually resolved." Self-healing passing reproduction tests and then collapsing live, a ledger residual going unflagged without being fixed, and both first instincts about concurrency and token caps pointing the wrong direction are all the same pattern.

Next week watches whether the five fixes on the backup GPU path actually hold up before revisiting hardware priority, alongside the paper-trading ledger correction, the ticker/direction guard, and landing the output-token cap.