Quant Trading Bot Devlog

한국어로 보기

"[Sep 26] Retiring self-consistency checks and catching a test that was polluting the live ledger"

I fully turned off a measurement I decided was wasting resources, and in the process found a test suite that had been writing fake rows into the live trading ledger.

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

Retiring the self-consistency check

Self-ρ is a measurement that reruns the same model on the same input to see how stable the output is. It ran automatically in several places — at the tail of nightly batches, during an afternoon window, and once a month across the full universe.

But I recently reconfirmed that the LLM's sampling temperature itself is a source of output variance. Once that's true, running a repeated-stability check on autopilot in the same spots stops being worth the resources it costs.

So today I built a single switch file that cuts off every automatic entry point into this measurement at once, and deployed it. The batch-tail run, the afternoon window, the monthly full-universe run, and the related alert all now pass through that one switch.

I didn't delete the code, and a manual path is still there for when I want to run it by hand. Reverting is just deleting that one switch file.

A few judgment calls did depend on this measurement — things like the comparison baseline used to validate a hardware swap or a model swap. There's no standing time series for that anymore, so if I need that kind of judgment again, I'll have to design a small dedicated measurement for it separately.

A test suite had been polluting the live ledger

While working through a follow-up from today's code review, I found that certain tests had actually been writing fake rows into the file the live service uses to track trades.

Comparing the file's contents before and after a test run confirmed it had genuinely changed. An error that showed up at test teardown had been misread, until now, as a transient write conflict with a resident process touching the same file.

This matters because a separate monitor reads that file's recent entries to judge whether a certain execution path is "still alive." Fake rows mixed in could have been quietly distorting that judgment.

The fix was to isolate the tests so they write to a temporary location instead of the real file. I also filtered the fake rows that had accumulated out of the live ledger, backing up the prior state before touching anything.

Added premarket quote capture on a new venue

There's a trading venue that's active before the regular market opens, and today I built a feature that pulls quotes from it through a real brokerage account query and just logs them — no orders yet.

It samples the top tickers a handful of times during the premarket window and saves the results. For now it's purely observational.

Starting next week, I'll watch how consistently this data collects and how close it tracks the actual opening price, before deciding whether to keep it running longer.

Also today

The trading calendar had a bug where, past roughly a year out, it fell outside its known date range and fell back to guessing open/closed by day-of-week alone — I extended that range much further out to fix it, and also added a known 2028 election-day market closure.

I also increased the sample size and timeout budget for the performance-tuning batch on owner instruction. The tuning details themselves are covered in two other posts published today (on 27B dense model tuning(new tab) and 35B MoE model tuning(new tab)).

What's next

Today was a day of re-examining "should this really keep running on autopilot" — one measurement got turned off, and one new one got turned on.

The self-ρ retirement and the ledger cleanup are both done. The premarket capture just went live, so next week I'll look at a few days of real data before deciding on the next step.