Quant Trading Bot Devlog

한국어로 보기

"[Sep 24] Shifting a Nightly Run Time and Reviewing a Standby-GPU Design"

Verification showed the time shift alone didn't achieve its original goal, and a design review for running the standby GPU full-time turned up a misread performance figure from the past.

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

Why I shifted the nightly run time

This bot runs every night, gathering the day's news and market conditions to prepare stock analysis for the next trading day.

The problem was that news published over a weekend or holiday wasn't making it into that preparation cleanly. Because the run time was pinned to "the evening of the last trading day," news published during a holiday stretch afterward sometimes didn't get picked up properly by the next run.

So I moved the run time to "the evening before the next trading day." Along with the code change, I found and fixed a regression where the monitoring logic had been missing delayed runs. All the tests passed, and I committed the change.

Verification showed it didn't achieve the goal

After committing, I got outside advisory input and traced through the code again — and found something I hadn't expected.

No matter how late the run happens, the news-gathering logic had always capped its collection window at "up through the last trading day." That cutoff was deliberate — it exists to keep timestamps aligned when comparing against other metrics later — but I'd missed that it conflicts with what I was trying to do here.

So even with the later run time, holiday news still falls outside the collection window. What the commit actually delivers is just a later GPU occupancy time and a bit more reliability in catching up on delayed runs.

To actually get "holiday news included," the news collection window needs to be decoupled from the trading-date cutoff as a separate piece of work. That's out of scope for this week, so I set it aside as a decision to make later.

This week also happens to fall in a multi-day holiday stretch, so I decided not to re-run the same date's data during it. The output from the last trading day is already feeding other validation work, and re-running it without any new information would just create confusion.

Also revisited: running the standby GPU full-time

Right now, one main card handles the core work that runs every night, and a standby card I picked up separately only gets used for occasional experiments.

There's a goal to eventually run that standby card full-time alongside the main one, so I brought in outside advisory input on two fronts today.

The first was about the locking structure. Right now, everything is locked as a single unit regardless of which card is involved, so even a job that only needs the standby card gets blocked whenever the main card is busy.

Splitting that into a two-tier lock, scoped per card, would fix that — but the more important finding was something else. Part of the existing code handles low-memory situations by broadly clearing related processes, and that pattern could, in the wrong moment, end up killing the standby card's processes while cleaning up after the main job.

That's something that needs narrowing down before full-time operation starts.

The second was about serving settings for the standby card. I'd previously referenced a figure suggesting a particular acceleration-path combination was faster, but tracing the source again showed that number actually came from an unrelated situation — a different model collapsing in performance under multiple concurrent requests.

It wasn't evidence that the combination I wanted to use was actually faster. So instead of rushing into that, I prioritized fixing something more basic first: the build I'm running is over two months stale. Rebuilding from the latest source moved to the top of the list.

Both advisories are now settled on direction and design; actual implementation hasn't started yet.

Also today

I also kicked off a background job re-grading a local model whose quality had been questioned, under the same conditions as the production model, for a blind comparison. I extended the safe-replay tool to support that model too, and after it passed a small-scale check, the full comparison is now running. I'll write up the results once they're in.

I also caught a bug in the performance-measurement tool: a run name containing a period would silently kill the re-registration step. It only showed up in a real run today, and I fixed it right away.

I also confirmed a minor issue where a live-trading timer fires a noisy alert on market holidays. The safeguard itself works correctly and no orders are affected — it's just a loud alert — so I plan to make it fail quietly instead.

What's next

Both threads today followed the same pattern: even after the code change and tests were done, tracing through whether it actually achieved what I wanted turned up something different from what I expected.

Neither result was bad enough to roll back the commit, but it was a good reminder that "fixed" and "achieved what I wanted" aren't always the same thing.