Quant Trading Bot Devlog

한국어로 보기

Making a 27B Dense Model 2x Faster on One RTX 5090 - The Bottleneck Was the KV Cache, Not the Concurrency Setting

How I cut a 100-stock nightly batch from six-plus hours to under three, and why you shouldn't fully trust that number

This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).

I decided to use a 27B dense model for the nightly 100-stock analysis batch. I ran experiments on how fast it can go on a single RTX 5090, and this post summarizes where I've got to.

The short version: I cut the production baseline of about 226 seconds per stock to about 102 seconds in a replay experiment. But that 102 is a single repetition on 24 stocks, so it isn't a production measurement yet. I'll write out those limits too.

What this batch actually does

For anyone who wants to reproduce the numbers, here's the workload first.

Analyzing one stock takes 17 LLM calls (a multi-agent pipeline). The per-node averages look like this.

Node Avg input Avg output Avg latency (3 slots)
Market analyst 13.3k tokens 3.0k tokens 68.4 s
Sentiment analyst 2.95k 0.87k 19.7 s
Portfolio manager 9.2k 0.52k 12.9 s
Trader 1.2k 0.35k 7.9 s

The total per stock is about 188k input tokens and 25.8k output tokens. Breaking down the time, prefill is about 38 seconds and decode about 540 seconds, so decode is 93%. That's why I judged there was little room in prefill optimization.

The node with the longest input is the Conservative Analyst. The largest input observed over nine production nights was 33,487 tokens (Bear 33,451, Bull 32,699), which is why I set the server's maximum length to 34,816. That's already tight, so I can't shrink it further. Anything over it fails the call with a 400 error.

Sampling settings

Sampling is temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5 — the model card's recommended values for non-thinking (reasoning-off) mode, used as-is.

The nine production nights before 09-24 were different. I had overridden only the temperature to 0.8 on top of the checkpoint defaults (1.0/0.95/20), so the effective sampling was 0.8 / 0.95 / 20 / presence 0. In effect I changed three axes at once.

The effect shows up most clearly in the output tail.

9 production nights (29,630 calls) New-sampling blind run (1,700 calls)
Output p50 / p90 / p99 1,530 / 2,837 / 5,148 tokens 1,337 / 2,670 / 4,402 tokens
Calls over 6,000 tokens 104 (0.35%) 2 (0.12%)
Max output 26,518 tokens 9,291 tokens

The presence penalty appears to have clipped the runaway tail. But since three axes changed at once, I couldn't separate which one deserves the credit.

The starting point was 21.7 hours on an XTX

I first ran this model on a single 7900 XTX(new tab) two months ago. On llama.cpp with 4-bit quantization (Q4_K_M), 100 stocks took around 30 hours.

Attaching a separate MTP (speculative decoding) head brought that down to about 21.7 hours, thanks to decode speed going from the mid-30s to the high-50s of tokens per second. That was still more than double my target of staying within roughly 10 hours, the window that fits overnight.

I learned one thing then: comparing by wall-clock time alone gives wrong answers. One run was over-measured by more than 3 hours simply because it happened to produce longer outputs. After that I compared arms by prefill/decode throughput instead of wall-clock time.

Moving to the 5090 and changing the stack

Moving to the 5090, I also switched the serving stack from llama.cpp to vLLM. I use a 4-bit floating-point (NVFP4) checkpoint and compress the KV cache to fp8.

The first production setting had 3 concurrent slots. Nine nights of production running on it measured about 226 seconds per stock, roughly 6.3 hours for 100 stocks. It cleared the target, but without much margin.

Before that, my first attempt with 14 concurrent slots had failed on memory. Back then I only looked at the number and thought, "more slots will make it faster."

Raising concurrency made it 1.77x faster

This time I froze the input (a replay experiment) and changed only the settings. The 3-slot baseline ran all 100 stocks at 204.5 seconds per stock (5.68 hours).

The 14-slot run took 115.8 seconds per stock (3.22 hours) on 28 stocks — 1.77x faster than baseline. Aggregate throughput went from 126 to 233 tokens per second.

But the server logs looked odd. The number of requests actually running at once wasn't 14 — it was 6 to 9. The queue always held 5 to 9 waiting requests, and per-call latency went from a median of 32 seconds to 70.

The bottleneck was KV cache capacity, not slot count

The cause was KV cache capacity. Even with 14 slots configured, once the cache fills up, new requests can't get in. Cache utilization hit 100% at peak.

This model isn't purely dense; it's a hybrid. Of its 64 layers, 48 use linear attention and only 16 use full attention, so KV accounting differs from a plain dense model. That's where the intuition "14 slots = 14x concurrency" breaks.

So the concurrency number I configure is meaningless; what matters is the concurrency actually realized. If I hadn't looked at the number of running requests in the log, I would have wrongly recorded this as "14-way, 1.77x."

Adding MTP doubled it

Next was speculative decoding. This time I turned on the MTP head built into the model, in vLLM, with 2 speculative tokens. At the same time I reduced the slots to 8 and raised GPU memory utilization to 0.95.

Running 24 stocks came to 102.0 seconds per stock (2.83 hours). That's 2.0x the baseline, and 12% faster than the 14-slot run without MTP. Aggregate throughput peaked at 437 tokens per second in the logs, with speculative-token acceptance of 65 to 79% depending on the window.

In this run too, the configuration said 8 slots but realized concurrency was 4 to 6, and KV utilization hit 100% at peak. The bottleneck is still the KV cache. Breaking down the time, decode is about 93% of it, so there's little room in prefill-side optimizations.

Why you shouldn't trust this number

The setting I've adopted is this 102 seconds. But the number has clear limits.

I also made one measurement mistake. Seconds per stock already reflects concurrency since it's wall-clock divided by stock count, yet I was about to divide by the slot count again. Dividing twice overstates speed by several times. Luckily review caught it.

Quality numbers

A speedup is meaningless if the output got worse, so here's what I know about quality too. The short version: I haven't been able to confirm quality yet.

So after speeding up the batch, my plan is to draw each stock several times and use the average score. Measuring IC from a single run's grades would let noise swallow the statistical power.

What I tried and dropped

What's next

On the first production night next week I'll measure the actual 100-stock time and see how much it differs from the replay experiment's 2.83 hours. That difference is the real conclusion of this post.

I plan to spend the leftover time on drawing each stock multiple times and averaging. I'll write a follow-up once the production result is in.

My record of tuning a 35B MoE model on the R9700 the same way is in the next post(new tab). The conclusions, from backend to concurrency to MTP, came out quite different.