Quant Trading Bot Devlog

한국어로 보기

"[Sep 20] Temperature Experiment — Direction Was Clear, Size Wasn't, Decision Deferred"

The two-day comparison test came back with an ambiguous verdict, so I had another AI independently re-check it, and cleared a backlog of fixes on the benchmark harness while I waited.

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

The main comparison run finished

The test comparing two settings for the model's response-randomness (temperature) parameter, which I'd started yesterday, finished its main comparison run early this morning.

The idea was to run the old setting and the new setting on the same inputs multiple times, and check statistically whether the outputs actually differ.

The first readout landed two of the decision thresholds in an ambiguous "gray" zone. One of the comparison runs also failed a quality gate outright and had to be thrown out as invalid.

So I set an automatic replacement run going for just that one, and spent the day waiting on it.

Direction was clear, size wasn't

Late in the afternoon the replacement run finished too, and I ran the final readout.

The tendency for the new setting to push the model further toward "hold" was consistent across several different statistical tests — hard to write off as chance.

The open question was whether that tendency was large enough to actually matter. On this data alone, the estimate spanned too wide a range to say for sure either way.

The original reason for adopting the new setting had been "more consistent results across repeated runs" — and that benefit didn't show up at all in this test. One of the justifications for the change was gone.

Asked another AI to check it twice, independently

Since the verdict was ambiguous, I handed the raw data to a different AI with no prior involvement in this experiment and asked it to recompute everything independently.

It reached the same conclusion: the direction was solid, but the magnitude couldn't be pinned down from this data alone.

I then asked a follow-up: given that, what's the better call to make by tomorrow evening? The recommendation that came back was to revert to the old setting and keep it frozen for a while.

The reasoning was that continuing to run the new setting needs some positive justification beyond "it's already running," and this test hadn't produced one.

I left the actual decision to be made by hand, deferred until tomorrow evening. Tonight I only kicked off one more run to sharpen the precision of the verdict a bit further.

Also cleared a backlog on the benchmark harness

Separately, a handful of small fixes had been piling up unaddressed on the LLM benchmark harness I built recently.

Today I knocked them all out in one sitting.

While rebuilding one of the helper programs used for measurement, I discovered that the rule for parsing its log output had never actually matched the real output format — it had effectively never worked correctly since it was written. Fixed the parsing to match the actual format.

Also fixed a bug where weekends and non-trading days were being blocked by a time-window restriction that shouldn't have applied to them, which had been forcing a manual workaround every time.

A report file that kept failing to install to its proper location because of a filename policy also got a real fix instead of the workaround I'd been using.

Yesterday's call to keep auto-resume interactive-only for now was reconfirmed today and left as-is.

After rolling out these fixes, I restarted the relevant resident services in the evening so the new code would actually take effect.

What's next

Deciding which way to finalize the temperature setting by tomorrow evening comes first.

I'll also check the results of tonight's extra run alongside that decision.

(How that decision actually played out is in the Sep 21 post(new tab).)