Quant Trading Bot Devlog

한국어로 보기

I Finally Changed the Temperature for Real - and Overrode a Pre-Registered Rule That Said "Reject"

The sampling-temperature experiment I believed I'd already run for weeks had never actually been executed

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

I ran an A/B experiment on the sampling temperature of the local LLM in production. The short version: the decision rule I'd fixed in advance said "reject the current setting," and I chose not to enforce it.

This post isn't arguing that the call was right. It's a record of what I based the override on, and how I documented it.

I had never changed the temperature

The project has a pipeline where a local LLM analyzes stocks(new tab). Last month I logged that I'd already run an A/B experiment comparing sampling settings for this model.

When I dug that record back up, that experiment turned out to be A versus A. The client code was dropping the temperature parameter instead of passing it to the server, so the B arm never received a temperature either. (I traced the cause in a separate postmortem(new tab).)

It wasn't "we discussed it but the validation was invalid." The reality was that I had never actually changed the temperature even once.

Production was running on the wrong mode's values

Looking at temperature properly for the first time, I found something bigger. Production ran with the model's reasoning (thinking) feature turned off, but no sampling values were specified, so whatever defaults were baked into the model file were being used.

Those defaults happened to be the recommended values for reasoning-on mode. The mode and the values didn't match, and it had been like that for weeks.

That same night I switched to the values the vendor recommends for reasoning-off mode. (That day's devlog(new tab)) At that point I had a rule of my own — "no production changes until the evaluation is done" — and I made an exception to it, knowingly.

What the "improved" numbers really were

For the first three days after the change, the numbers looked good. The rate at which two runs on the same input agreed on the grade went up. It looked like reproducibility had improved.

But the share of the Hold grade had gone up over the same period. When results pile onto one grade, the agreement rate rises on its own, because independent draws are more likely to coincide by chance.

So the increase in agreement wasn't evidence that reproducibility improved — it could be the same fact as "results are collapsing toward Hold" stated differently. That's when I upgraded my first "all clear" reading to an alert.

A pre-registered confirmation experiment

Now that I had a suspicion, I designed a confirmation experiment. Since the hypothesis came from looking at data, I wrote the decision rules down in a document before seeing any results.

The first A-arm run did trip a gate. Format parse failures and truncated responses exceeded the limit, so I threw that run out and ran a replacement. The discarded run's numbers were kept for reference only.

The rule's answer was "reject"

The results:

The rule detected exactly the hypothesis I'd set up. The rule wasn't wrong.

I kept B anyway

I asked a different AI for advice, and at first it recommended reverting to the old setting. The pre-registered rule said reject, and the one thing that would have justified keeping B, its reproducibility gain, had no evidence.

But I pointed out two things.

First, the old setting can't be called a "validated previous regime." I'd only run this model in production for a bit over ten days, and the old setting was a reasoning-mode default that had been applied by accident. The values the vendor recommends have surely been validated by far more usage than a few nights on my machine.

Second, in my investing, a lot of Hold isn't automatically bad. Treating it as a loss was an assumption I'd put into the endpoints when designing the experiment, without verifying it.

When I passed these two premises along, the advice flipped within a day. The side that had recommended reverting the night before now recommended keeping B. The data hadn't changed; the premises I gave it had. So this decision isn't "the experiment supported B" — it's more accurate to call it "I overrode a pre-registered rule with after-the-fact premises."

I wrote the override down

Overriding a rule is allowed, but not quietly. I recorded four things in the experiment document:

I also attached conditions. A non-blocking alert fires if the Hold share climbs too high, and I committed to not changing the sampling setting, whatever the results, until the next scheduled review. That review happens after the fix described below.

The input itself was degraded

After the experiment I learned something more awkward. Two external data sources had quietly died a few days earlier, so the news going into the prompts was nearly empty and even stock names were being replaced by codes.

The frozen input for this A/B happened to be a day from exactly that degraded period. So the results carry a caveat: they're results on "input with no news." The day the Hold collapse jumped also coincided with the first day the news vanished, so I can't pin the collapse on the new setting alone.

I fixed the inputs and added the degraded window as a boundary note in the evaluation document. The next review will use normal inputs.

Takeaways

Believing you changed a parameter and confirming it reached the server are two different things. A run now counts as valid only if the server launch command shows the flags. No evidence, no valid run.

A metric that looks better can just be distribution arithmetic. If an agreement rate went up, first check whether the results piled onto one side.

When you override a pre-registered rule, record whether the rule was wrong or the premise was wrong. Here the rule was right and the premise built into it had never been verified.

Advice depends on the premises you hand over. If a conclusion changes, you have to record what changed, the data or the premises, so you can revisit the decision later.

Some things remain unresolved. Whether the Hold collapse is actually costly, and whether normal inputs produce the same collapse, are things I'll only find out at the next review.