Quant Trading Bot Devlog

한국어로 보기

Local LLM Speed by GPU Generation, and Our Production's Daily Token Volume

Model names, serving configs, and full prefill/decode measurements included

In my last post(new tab) I talked about switching to the RTX 5090. This time I'm going much deeper — exactly which model I run and how it's configured, precise prefill/decode speeds by hardware, and even the interim results of the model-adoption decision I'm currently running.

The hardware I've gone through

This bot's local LLM card has gone through one RTX 4070 Ti Super → two → one RX 7900 XTX → one RTX 5090.

Setup VRAM Result
4070 Ti Super ×1 16GB The model back then didn't fit in memory at all
4070 Ti Super ×2 (tensor-split) 32GB Nightly 100-stock batch: 5.81 hours
RX 7900 XTX ×1 24GB Same batch: 6.15 hours (one card replaced two)
RTX 5090 ×1 32GB See below

The actual model and serving config

The model running in production every night is Qwen3.6-35B-A3B (MoE, 3B active), GGUF Q4_K_M quantization. Weights are 20.75GiB + 1.16GiB of KV/compute buffer = 21.91GiB in actual use (91% of the XTX's 24GB). The vLLM server runs on port 8200 with max-num-seqs 14, fp8 KV cache, prefix caching off, temperature 0.8, thinking mode off — this is the frozen production command.

At the same time, I'm evaluating whether to adopt Qwen3.8-27B (dense). Same conditions (fp8 KV cache, prefix caching off, temp 0.8, thinking off), but max-num-seqs is dropped to 3 — dense has different memory/compute characteristics than MoE, so concurrency has to be lower for stability.

MTP (speculative decoding) is always on. Acceptance rate is 50.1% for our MoE model, 66.95% for the 27B dense model in an external measurement, and anywhere from 81-97% in other external measurements — it varies a lot by model and settings.

Prefill/decode measurements — 27B dense

Same 27B dense model, roughly 12.7k prefill tokens per call, limited to cards I've actually run myself (I'm leaving out AMD R9700 — I've never tested one myself, so it only appears as a cited reference in the public-benchmark section below):

Setup decode 100-stock equivalent Note
RX 7900 XTX, Q4_K_M, MTP on 62.5 t/s (single stream) 21.7h prefill 768 t/s, our measurement
RTX 5090, 27B dense 118 t/s (aggregate throughput) ~6.3h one complete 100-stock pass used for today's ranking — see "Today's real ranking run" below

These two numbers are measured differently (single-stream vs. aggregate), so a straight "N times faster" division isn't quite fair. Still, looking at the 100-stock runtime alone, it dropped from 21.7 hours on the XTX to about 6.3 hours on the 5090 dense run. That said, pitting MoE against dense directly on the same 5090 tells a different story, covered below.

Raw decode speed by card — public cross-GPU benchmark

Decode speed by card for the same 27B-dense-class model with MTP on (source: arca.live cross-GPU thread(new tab); context depth uncontrolled, reference only):

Card decode (t/s)
RTX 3090 40-45
AMD R9700 56
RTX PRO 6000 Max-Q 70
RX 7900 XTX 71
RTX 5090 80

The 5090 is about 1.13x faster than the XTX. I originally expected a much bigger gap based on the memory-bandwidth ratio (1.81x), but it turned out that conversion doesn't hold once MTP is on — speculative decoding frees decode from the bandwidth ceiling.

Decode speed from our own production logs, by hardware generation

ta_usage_tracker logs input/output tokens and elapsed time for every LLM call, per stock. I first averaged output_tokens ÷ elapsed_s per call, but vLLM batches multiple stock requests concurrently (continuous batching), so a single call's apparent speed swings with however many other requests happened to be in flight at that moment — that gave numbers that couldn't actually be used to compare hardware. So instead of averaging per-call rates, I recomputed it as total output tokens produced during that run ÷ the run's actual wall-clock duration (first call to last call).

Era Aggregate throughput (t/s) Note
4070 Ti Super ×2, MoE (33-day avg) 62.4 Normal full nightly batch
RX 7900 XTX, MoE (10-day avg) 71.6 Normal full nightly batch

For the 5090, instead of a "normal nightly batch" figure, here's exactly what actually produced today's ranking — this is the most accurate one.

Today's real ranking run — MoE vs dense on the 5090

I ran MoE and dense side by side for comparison, and the ranking that actually went out today used the dense results. Here's what it actually took each model to complete one full pass over 100 stocks (MoE ran yesterday evening, dense ran early this morning — each model's first complete pass):

Model Stocks Run time Input tokens Output tokens Wall-clock duration Aggregate throughput
MoE (existing) 100 yesterday 20:41-21:27 18,315,222 2,497,437 46 min 911 t/s
27B dense (used for today's ranking) 100 today 00:13-06:31 18,982,363 2,680,713 6h 18m 118 t/s

Same 5090, same 100 stocks — MoE is about 7.7x faster than dense. Dense being slower isn't a hardware issue — it's inherent to the model architecture (MoE only computes its active parameters; dense computes all of them every time).

The model-adoption decision, in progress

Which model gets formally adopted isn't decided yet. Running today's ranking with dense for comparison doesn't mean it's formally adopted — one of the ongoing verification methods is self-QWK: running the same stock set multiple times and measuring how consistent the resulting grades are (quadratic weighted kappa). For MoE, I already have three repeated runs on the same 99 stocks (the MoE run in "today's real ranking run" above is one of these three):

Pair QWK Raw agreement
rep1 vs rep2 0.0853 59/99
rep1 vs rep3 0.1260 56/99
rep2 vs rep3 0.1481 61/99
Average 0.1198

The self-QWK band observed historically on the XTX was 0.08-0.21, so this value (0.1198) falls within it. The grade distribution skews toward "Hold" (71-80%), and QWK mechanically comes out lower when one grade dominates, so the raw agreement rate (56-61%) needs to be read alongside it.

Self-QWK for dense hasn't been measured yet (planned for this weekend). The metric that actually decides adoption isn't self-QWK — it's forward RankIC (the correlation between predicted and actual return ranking), and there are zero measurements of that so far. In other words, right now there's only a qualitative expectation that dense will perform better; whether it actually does is still unverified.

Daily token volume

This project has an AI pipeline that analyzes stocks with a local LLM(new tab).

Normal weekday average (MoE only, no side-by-side comparison running): about 5,200 calls, roughly 55 million input tokens, roughly 7.5 million output tokens.

Same single 100-stock pass as "Today's real ranking run" above:

Model Stocks covered Calls/stock Input tokens Output tokens
MoE (existing) 100 17 18,315,222 2,497,437
27B dense (used for today's ranking) 100 17 18,982,363 2,680,713

Both models used the same call depth here (17 per stock), so the token volume itself is nearly identical — the difference shows up not in tokens, but in the wall-clock time above (46 min vs 6h 18m). For reference, on a normal day of MoE running solo, calls per stock can run much higher (55-71), since call depth itself depends on the purpose of that particular run (a formal production batch vs. a comparison run like this one).

What's next

Once forward RankIC is measured, I'll write up which model actually got adopted as a separate post.