"[261005] Shipping the Output Token Cap and Expanding Order-Ledger Reconciliation Guards"
I rolled out yesterday's output-token-cap fix to both serving paths, and added two more layers to the order-ledger reconciliation safeguards.
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
Fixing yesterday's problem today
Last night I found out the output token cap wasn't set at all.
Today was the day to actually ship the fix. I set a sensible output cap on both the main and backup serving paths, and fixed the tool-calling bug alongside it.
I ran two rounds of smoke tests before shipping. Most checks passed, but one metric — how often a concurrent request gets preempted — failed the threshold both times.
Digging in, this metric turned out to be tied to a structural property of memory occupancy, unrelated to either the new cap or the tool-call fix. I asked another AI for a second opinion, and got the same read.
It also came out that this threshold had been set to "must be zero" without any real basis to begin with. Existing production logs showed the same rate happening every night regardless.
In the end I shipped on the grounds that this was the only failing metric and everything else passed. I didn't lower the threshold on the spot just because of this result — the actual threshold re-measurement got pushed to the next market holiday.
Two more layers for order-ledger reconciliation
I've been steadily expanding the order-fill ledger reconciliation safeguards(new tab) since going live, and added two more today.
One catches cases where the filled quantity exceeds the original order quantity. The other audits for an abnormal state where the ledger quantity goes negative at end of day.
In the same pass, I added a mechanism so that rejected lookup results get queued instead of discarded, so the same problem can be tracked if it recurs.
I built this in a separate workspace and had an AI run a code review on it. It raised three notable issues, all of which I addressed before merging into the main workspace.
One of them was that a test depended on the actual calendar date instead of a fixed date, so it would silently get skipped if run on a market holiday. I fixed it to pin the date inside the test.
Also: filled an observation gap
Serving has never had a record of how often preemption happens mid-request.
Today I stood up a lightweight external collector that polls the server state periodically from the outside, without touching the serving process itself. The very metric that caused today's threshold dispute will now be tracked as real numbers going forward.
Why this matters
Both changes today ran into the same question: what do you do with one failing metric?
For the output cap, I only moved forward after confirming with evidence that the threshold itself was wrong from the start, and I pinned down a separate re-measurement step rather than letting it slide.
For the order guards, I went the other way — I only merged after accepting every issue the AI review raised. Neither case treated "it works, so ship it" as good enough.
What's next
I'll watch both serving paths at tomorrow's first pre-dawn production run for cut-off responses or tool-call errors.
I'll also check that the new order-ledger guards stay at zero through tomorrow's first trading round and end-of-day close.
The threshold re-measurement is scheduled for the next market holiday.