Quant Trading Bot Devlog

한국어로 보기

"[261004] A Concurrency-Scaling Smoke Test, and Discovering There Was No Output Token Cap"

I tried to raise the serving concurrency for more throughput and the smoke tests caught it, and along the way I found out the output side had no token cap at all.

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

Trying to push throughput higher

I decided to try raising the concurrency level (often called "np," the number of requests served in parallel) for the 35B model currently in use, to push throughput higher.

I put two higher settings into the smoke test alongside the current one — one step up, and two steps up.

Neither passed. One failed on total run time, the other on individual request latency.

Instead, I retried with a configuration that raised the setting by just one modest step while also adjusting a couple of other serving parameters, and that one passed. For now I'm sticking with that setting, and the bigger jump is on hold.

The backup-path switch left a gap in the monitor

I'd recently moved the backup serving path to a different GPU pairing, and that move created an unexpected side issue.

The watchdog logic that checks whether the main process has died was judging purely by process name. Once the backup path started spawning processes with a similar name, the watchdog could mistake the backup for the main process.

In the worst case, the watchdog could have force-killed a perfectly healthy backup process, thinking it was the main one. Code review caught this along with several similar issues, and all of them got fixed.

I also patched a rerun script for the same underlying reason — now that the backup path shares a GPU with the main one, I added a guard against the old pattern of launching both at once.

A runaway output that never stopped generating

Going through the smoke logs in the afternoon, I noticed one stock analysis call that finished only after producing a vastly longer output than usual.

A normal response tops out at a few thousand tokens; this one blew past that by several multiples. My first instinct was that the output token cap needed to be loosened.

I asked an AI to dig into the root cause, and the conclusion was the opposite. The current serving configuration had no output token cap set at all.

With nothing to stop it, the model kept generating until it filled the remaining context window, and that bloated output then triggered a context-overflow error in the next stage. My instinct about the problem was right, but I had the lever pointed the wrong way.

The AI's recommendation was to set a new, lower cap instead of raising anything. The value it suggested wouldn't touch any normal response, only the runaway ones.

In the same pass, I found a separate bug in the tool-calling path used mid-analysis. Under a specific fallback condition, the server would return an error — something that had been too rare to notice until now.

Why this matters

Both incidents today share the same shape: my first instinct about the cause ran in the opposite direction from the real one.

Concurrency looked like it should go up, but the smoke test hit a wall first. The output looked like it needed a looser cap, but the real problem was that there was no cap at all.

In both cases, measuring first and getting a second opinion before changing a number kept me from running in the wrong direction.

What's next

Concurrency stays at the setting that passed for now; the bigger jump gets revisited later.

The lower output cap and the tool-call fix will get a fresh smoke test tomorrow during the market holiday — if it passes, both ship with the same restart; if not, they wait.