"[Sep 22] The Server Was Quietly Degrading - Building a Detect-Isolate-Recover Safeguard"
Right after closing out yesterday's news-input incident, a new defect showed up, and I built a safeguard against it in a single day.
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
Closing out yesterday's incident
Last night's discovery was that the input to the news-based stock scoring pipeline had quietly broken.
A change in how a data source served stock pages and news listings meant the bot couldn't even resolve ticker names anymore, and had been scoring stocks with essentially no news for over a week.
This morning, the fix deployed after yesterday's advisory review was confirmed to have run to completion.
It wasn't fully closed out, though. The days with missing data still linger inside the moving-average window for a few more sessions, so I noted the residual gap in the verdict record and pushed the deeper structural fix to after the next market holiday.
The server was quietly degrading
The next problem showed up almost immediately.
This morning, a handful of stocks that should have scored normally — including one I actually hold — came back marked as the vendor's default "hold" fallback instead.
Tracing it, this turned out to be unrelated to the news fix. The local inference server had been running for a long stretch and had slipped into a "degenerate" state on certain requests, repeating the same character over and over instead of producing a real answer.
Fortunately, it had no real effect on the day's rankings — I confirmed that separately.
Since this could recur, I built a detect-isolate-recover safeguard the same day. When a response shows a degenerate pattern, it's dropped instead of being scored, the server is automatically restarted, and just that stock is retried.
I also ran the detection logic against thousands of lines of historical incident logs first, and it caught every real incident with no false positives.
A separate bug in the smoke-test script — it crashed when run standalone — got fixed that evening too, letting me verify the full detect-restart-retry loop against the live server.
Keeping the 60-day window
A different kind of decision came up the same day.
I'd already pulled an experimental "shadow overlay" strategy out of the live account and left it running purely for observation, against a pre-committed 60-trading-day window. At day 49, I considered calling it early.
Statistically, the remaining days were unlikely to change the conclusion. But I decided to hold to the pre-committed window anyway — peeking at interim results and deciding whether to stop is a well-known statistical trap.
The fact that waiting longer cost essentially nothing (the strategy was already out of the live account) made that call easier.
Along the way, a separate issue turned up: this strategy churns through a large share of its holdings in a single day, which is now a separate item to look into.
What's next
Looking back at today, there's a pattern across all three threads: rather than jumping straight to a fix, I got an independent read on the root cause and options first, then made the final call on top of that.
Both the point where the news-input gap fully ages out of the moving average and the end of the shadow strategy's 60-day window are now scheduled to resolve automatically within the next few days.