"[260927] R9700 Backup GPU - Bottleneck Root Cause and a Dead-Slot Detector"
I found why the backup path doesn't get faster with more concurrent requests, and built a detector that catches a slot dying and spewing runaway output.
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
Continuing the last few days' story
A few days ago, in my record of tuning a 35B MoE model on the R9700(new tab), I wrote that raising concurrent requests helps only up to 4, and beyond that it actually hurts.
After that, when I re-measured at a 24-stock scale, something strange showed up. Runs with 6 or 8 concurrent requests occasionally dropped a handful of stocks entirely.
Today was spent chasing both questions to their root cause: why throughput doesn't scale with more concurrent requests, and why stocks occasionally vanish. I sent an external advisor (fable5.1) everything measured so far and got back a prioritized roadmap, then worked through it in order.
The bottleneck isn't slot count, it's how fast expert weights get read
Re-measuring throughput at different concurrency levels, the added latency per extra request stayed roughly constant. That means the marginal cost of producing one more token is about the same no matter how many requests are running.
This model only activates a subset of its expert networks per token. As concurrency goes up, more distinct experts get activated at once, so each step ends up reading a wider slice of expert weights fresh from memory.
So adding more requests just adds proportionally more weight-reading time, and aggregate throughput plateaus past a certain point. That's the opposite of what happens with a dense model, where more concurrency reliably helps.
Given this diagnosis, I dropped the idea of pushing concurrency any higher today. Instead I moved the axes that can actually raise throughput further — lower-bit quantization to cut bytes read per token, and kernel-level efficiency — to the front of the next round of experiments.
Why stocks vanish — a slot dies quietly
Digging through the server logs from the runs that dropped stocks, the same one slot (a concurrent processing lane) was the culprit every time. Once a request on that slot ran abnormally long just once, every subsequent request landed on that same slot also ran the same abnormally long way.
The slot had effectively died and stayed dead afterward. Any call requiring structured output that landed on the dead slot failed outright, and that's what looked like a "stock going missing."
This was the same kind of failure I'd observed once before on the live-trading path, where I never pinned down the cause. This time, by narrowing the reproduction conditions, I caught it.
Once I knew the cause, detection turned out to be simple: if the same slot shows the same abnormal pattern twice in a row, that slot is dead. Today I extracted just that detection logic into its own piece of code.
I committed it as detection only, wired to nothing yet. To avoid disturbing another experiment chain currently running, I'm leaving the code dormant for a few days and plan to connect it to an actual remediation step (restarting the affected slot) after this weekend. In testing, it caught both damaged runs exactly and raised zero false alarms on two healthy runs.
The replay harness's degeneration watchdog was misfiring
As a side effect of this diagnosis, I found one more thing. Even when running replay experiments on historical data rather than live trading, the live-trading anomaly watchdog was still active alongside it.
That watchdog is meant to wait a fixed period and then force a cleanup if something in live trading stalls. In replay experiments, though, no live-trading signal ever arrives, so everything always looks "stalled" — and every single time, it waited out the full timeout and then force-killed the experiment process.
Replay experiments have nothing to do with live trading, so this watchdog should be off there. I fixed that today.
What's next
Next up is the kernel profiling run I scheduled today: measuring, without any stocks involved, exactly which operation the decode step spends its time on, across several configurations in 6-to-8-minute increments. It's set to run automatically Sunday evening.
Once this weekend passes, I'll wire an actual restart action into today's dead-slot detector, then move on to lowering the quantization level to cut the amount of data read per token.