Benchmark program · rewritten 11 Sep 2026

Local Lane

What two RTX 3090s, an RTX 2080 and a Ryzen AI MAX 390 measured between 17 Aug and 10 Sep 2026, and what the instrument got wrong along the way.

Measure quality once per model and speed once per machine. Neither one is decode tok/s.

What this is

I benchmark the three chips I have:

Every figure below comes from a committed run in my evaluation tree, and each one carries its run date, sample size and configuration. Where a number was later found to be wrong, the correction is stated rather than the number quietly replaced.

The measurement standard, as of 8 Sep. Models run as shipped. Only context length, temperature, thinking (off), the prompt and the probe length are pinned. Speed means measured task wall clock under a realistic prompt, not decode rate on an empty cache. So a ranking here is a ranking of each model at its shipped defaults, not of the model at its best.

The seat is qwen3.6:35b. I pick the seat from interactive use. The ranking informs that choice, it does not make it.

An earlier version of this page ranked the two chips by decode rate. It is archived in the site repository, and the sections below explain which of its numbers were wrong and why.

The ranking

One table, all three chips. Speed is measured task wall clock on two E9 fixtures, the mean of the warm trials with the first of each run discarded, normalised to gemma4:26b on the RTX 3090 = 1.00. Quality is the clean L2 pass count. Configuration: num_ctx 16384, temperature 0.8, thinking off, ollama 0.32.12 except the MAX 390 at 0.32.13.

#ChipModelArchshort task slong task sSpeedL2 clean
1RTX 3090qwen3.6:35b (seat)MoE8.815.61.9951/51
2=RTX 3090gemma4:26bMoE13.038.91.00149/178
2=RTX 3090qwen3.8:27bdense18.132.40.9663/63
4MAX 390qwen3.6:35bMoE27.340.60.7218/18
5RTX 3090gpt-oss:120bMoE31.153.40.5760/60
6=MAX 390gemma4:26bMoE26.3132.40.3916/18
6=MAX 390qwen3.8:27bdense35.396.80.3918/18
8RTX 3090gemma4:31bdense36.4128.40.3328/33
9RTX 2080gemma4:26bMoE41.3171.60.27 (0.23 to 0.31)15/15
10RTX 3090qwen3-nextMoE92.0325.70.1319/22

Ten is the whole list, not a cut from a longer one. Six models have wall clock in the corrected configuration on the RTX 3090, three on the MAX 390 and one on the RTX 2080, which comes to ten. Several further models have quality data but no wall clock under the corrected protocol, so they cannot be placed here at all. The ranking grows when those are re-run, not when the cutoff is moved.

Rows 2 and 3 are joint second. Bootstrapping the RTX 3090 ranking over warm trials and pass counts puts qwen3.8:27b ahead of gemma4:26b in only 79.1% of draws, which is not separable. Every other adjacent pair on that chip is ordered at 98.7% or better, and qwen3.6:35b takes first in every draw.

The RTX 2080 row is the weakest evidence in the table. That fixture ran three trials rather than four, so its warm mean rests on n=2 where every other row has n=3, and its long task was still falling across those two trials, 218.6 s then 124.5 s. Its two per-task ratios spread from 0.23 to 0.31, wider than the gap between several adjacent rows above it. Read the row as placing the card in this neighbourhood, not as fixing it at 0.27, and note the direction: with the long task still coming down, 0.27 more likely understates the card than flatters it. It is also the only roster member that runs on the 8 GB card at a usable rate, so there is no second 2080 row to check it against. Note it does not fit that card: its experts run from host RAM, as the correction further down sets out.

The quality column is not poolable across chips. The RTX 3090 counts accumulate over the seat corpus, the MAX 390 figures are an 18-cell floor battery, and the RTX 2080 is 15 cells at n=3 over five tasks. The denominators are printed because they are not interchangeable, and a rate computed across them would be arithmetic rather than a measurement.

Quality does not carry this ordering. Five of the ten rows sit at exactly 100%, so the order comes from speed. Finding a repeatable quality differentiator is open, and more trials cannot fix a ceiling.

Decode speed does not predict finishing

Five models, two E9 tasks, n=2 each, dual RTX 3090, 5 Sep.

Modeldecode tok/stask wall clock soutput tokenstool calls
qwen3.8:27b73.828.87884.8
qwen3.6:35b126.629.01,3807.0
gemma4:26b214.134.73,3665.8
gpt-oss:120b35.191.49216.2
gemma4:31b34.8117.12,7154.8

The fastest decoder is not the fastest finisher. gemma4:26b decodes 2.9x faster than qwen3.8:27b and finishes later, because it writes four times the output tokens. That is why the standard moved from decode rate to task wall clock. One correction to this run: its wall clock included about 4.5 s per trial of ledger bookkeeping, since removed. The ranking above is re-measured without it.

Decode rate also depends on the prompt. The same probe under a ~12.8k-token prompt, the loaded condition that matches my actual use, shows every model losing throughput, but by very different amounts:

Modeldecode, 121-token promptdecode, ~12.8k-token promptchange
gemma4:26b226170.1-25%
qwen3.8:27b78.162.3-21%
gemma4:31b37.733.8-11%
qwen3.6:35b130.0127.7-2%

Because the gap is model-dependent there is no conversion factor between the two. A decode figure means nothing without the prompt length it was measured under.

Across machines

The rows above interleave three chips on purpose. What they show when read by machine rather than by rank:

Quality rank transfers between chips. Speed rank does not. Both the desktop and the MAX 390 put qwen3.6:35b and qwen3.8:27b at ceiling with gemma4:26b below. There is no single chip multiplier either. The MAX 390 penalty runs from 1.95x to 3.40x against the same model on the RTX 3090 and reverses between tasks, so any single "N times slower" figure for this chip is wrong. That is also why the MAX 390's fastest model on the short task is not its fastest overall: gemma4:26b reads the prompt quicker than the seat and then loses three times the ground on the long one.

Stated confound: the MAX 390 ran ollama 0.32.13 against the desktop's 0.32.12. That affects cross-chip magnitudes, not ordering within a chip. MAX 390 quality is 18 cells per model, shown as a rate. Its numbers also depend on power state, so write the power state down or throw the number out.

A headless test bench with an 8 GB RTX 2080 reproduces the desktop's correctness (15/15 against the desktop's 28/30 on the same tasks) at 3.8x the wall clock, 141.1 s against 37.1 s. And capping VRAM on the 3090 is a good stand-in for a smaller card: a simulated 8 GB envelope ran gemma4:26b at 31.55 tok/s against a real 2080's 31.02, 1.7% apart.

What the 8 GB card can hold decides everything else about it. Same probe, num_ctx 16384, ollama 0.32.12:

ModelClassDense layers on the cardMoE expertsVRAM useddecode tok/ssame probe, RTX 3090
gemma4:26bMoE31 of 31host RAM7,416 MiB31.02226
qwen3.8:27bdense16 of 66n/a6,438 MiB3.8074 to 77

The dense model is the smaller file and runs 8.2x slower, because only 24% of its layers fit.

Correction, 12 Sep 2026. The gemma4:26b row was published as fully GPU-resident and it never was. That row previously read "31 of 31" with no other placement column, which reads as the whole model on the card. It is not. gemma4:26b is a 17.33 GB MoE blob whose full residency needs about 16.7 GB of VRAM; the card has 8,192 MiB with roughly 7.4 GB free. Ollama's own fitting pass resolves that shortfall by moving every expert tensor to system RAM and keeping only dense weights on the GPU, then logs offloaded 31/31 layers to GPU. Both statements are true and they describe different things: the layer counter counts layers, not experts.

This is checked, not inferred. Ten separate loads of this model on this card, from 20:19 on 11 Sep to 02:52 on 12 Sep, every one recording all MoE tensors moved to system memory with a shortfall of about 11.5 GB. Free VRAM ranged 6,478 to 7,413 MiB across those loads and never changed the outcome. This model has never been GPU-resident on this card and cannot be.

What that does and does not change. The 31.02 tok/s figure stands as measured, but it is a mixed-placement rate, not a resident one, and it is fragile in a way a resident rate would not be: the experts are memory-mapped from the model file, so decode depends on whether host RAM can cache them. Measured on the same card at the same placement, with the host under pressure, the same model returns 4.0 tok/s with about 580 major page faults per token and 38% iowait, an 8x swing with nothing on the GPU changing. Quality is untouched and is reported separately: 15/15 on this battery, 17/18 across the 11 Sep runs. Mixed placement changes how fast a row was produced, not whether it was right.

The instrument lesson. offloaded 31/31 layers to GPU is not sufficient evidence of MoE residency, and neither is ollama ps, which reported this model at 1.22 GB with a "25%/75% CPU/GPU" split while nvidia-smi showed the process holding 7,202 MiB. Only the load-time fitting decision, or a per-device VRAM measurement checked against the model's known full-residency requirement, establishes where an MoE model actually sits. The measurement standard has been changed accordingly, and no MoE residency claim on this site now rests on layer counts alone.

Still open: the Ryzen AI MAX 390 MoE rows. No fitting-pass evidence was captured in any run on that machine, so the same defect can be neither confirmed nor ruled out there. It is also an APU with unified memory, where the device-versus-host distinction is not the same question. Those rows are flagged pending a fresh load with the fit log captured.

The bench remains the functional-testing host, and anything about memory stays on the desktop.

Memory, context and placement

Pin num_ctx. gemma4:26b ships with a default context of 262,144 tokens. Unpinned, its cache takes it from 19.7 GB to 39.3 GB and splits it across both cards, and decode drops from 214.4 to 139.9 tok/s (n=6 interleaved, ranges do not overlap). The 35% loss was the oversized cache, not the split.

Splitting a model across two cards costs almost nothing when it fits on one. With context pinned, moving four models between one card and two changed decode by -4.7% to +9.3%, and under a realistic prompt by -8.0% to +6.7%. Decode walks the layers in order, so the two cards never work at once, and what crosses between them per token is tiny. What splitting does cost is about 1.4 GB of VRAM.

Modelsolo tok/sdual tok/sdecodesolo prefilldual prefill
gemma4:26b (control, one card either way)171.3170.1+0.7%3.5 s3.5 s
qwen3.6:35b132.7124.4+6.7%4.8 s3.9 s
qwen3.8:27b56.461.3-8.0%11.1 s11.0 s
gemma4:31b33.733.8-0.3%14.2 s8.5 s

The exception is prefill. Two of four models read their prompt much faster on two cards, gemma4:31b by 40% and qwen3.6:35b by 19%. I do not know why. My first explanation, that they were short of headroom on one card, was tested the same evening and did not hold.

Where the second card earns its place. At a configured context of 131,072 with a ~12.8k-token prompt, qwen3.8:27b ran at 61.4 tok/s on two cards and 11.2 on one, about 5.5x apart. This is a configured-capacity sweep, not a prompt-length sweep, and placement on the one-card side was not captured, so how much of that comes from spilling to CPU is provisional. Two models also fit side by side on the pair up to a context of 32,768, which makes switching between them take 0.6 s instead of a full reload.

Spilling to the CPU is expensive. Forcing layers of gemma4:26b onto the CPU (empty-cache probe, n=3) took it from 224.8 tok/s at 31 of 31 layers on the GPU to 172.7 at 30 of 31. One layer cost 23%. All four measured points fit a fixed 1.34 ms per CPU-resident layer. Whether that cost carries over to other models is not established. A rough estimate for qwen3.8:27b comes out near 5.5 ms per layer, but it divides a VRAM shortfall by an assumed layer count rather than counting layers, so it is arithmetic, not a measurement. And a dense model near its memory limit does not slow down smoothly: gemma4:31b at a reduced GPU layer count passed two tasks 3/3 and timed out on others. The limit is a stall with task-dependent odds, not a wall.

The interactive version of this is Where the Layers Go.

The quality ceiling and the noise floor

Small samples lie. One cell, gemma4:26b on csv-summarize-repair, had produced 5/6, 6/6 and 6/6 in separate runs, two of them read as clean results. At n=100 its true pass rate is 75/100. Six trials cannot distinguish that from 100%. Every result is now labelled with the evidence tier its sample size supports.

The battery is saturated at the hard rung and still works at the easier one. At L2 the current roster tops out the E9 tasks. At L1 the same tasks still separate models: gemma4:31b and gpt-oss:120b 30/30, qwen3.8:27b 26/30, gemma4:26b 24/30. Three fixtures that had been written and never run came back 18/18 for all seven models tested. The top of the ladder is exhausted.

Decode class does not predict correctness either. nemotron-3-nano decodes at 121.2 tok/s and passed 0 of 18 cells, missing the implementation rather than timing out.

A tool-grounded consultant bench

The newest bench asks consultant questions about a real medical-device registry, 2008 to 2025, three companies. The model gets a catalog of registered tools and one question at a time, has to actually call the tools against the live database within three turns, then commit an answer as JSON.

Each answer is scored three separate ways: protocol (well-formed turns), execution (it really called the tools when the task needs data) and truth (the answer matches ground truth computed independently by SQL over a pinned copy of the database, never taken from the tool path being judged). One task is planted: it asks about 2030, outside the data, and the only correct answer is that there is none.

Overnight on 10 Sep, one RTX 3090 per model:

Modellapstask pass ratetasks passed on every lapnever passed
qwen3.8:27b30283.3%01, 02, 03, 04, 0605
qwen3.6:35b61233.3%04, 06 (06 on 611 of 612)01, 02, 03, 05

It ran on off-peak electricity for under a dollar, estimated from the utility rate rather than metered.

Hosted models on the same bench, as controls:

Controllaps010203040506
DeepSeek chat (OpenRouter)133/1313/130/1313/130/1311/12
Claude Opus1passpasspasspassfailpass
GPT-5.6 Sol, reasoning off1failpasspasspassfailpass
GPT-5.6 Terra, reasoning off1passpassfailpassfailpass
GPT-5.6 Luna, reasoning medium1passpasspasspassfailpass
GPT-5.6 Luna, reasoning off21 of 21 of 20 of 22 of 20 of 20 of 2

Task 05 fails on every model tested, local and hosted, including a frontier model with reasoning on. When every model agrees, suspect the bench. I read it as a defect in how the task handles year-over-year figures, not as a ranking.

Everything except DeepSeek is a single lap and cannot be ranked against hundreds of local laps. Within this harness, though, the local 27b passes tasks 01 and 03 on every one of 302 laps, where DeepSeek managed 3 of 13 and 0 of 13.

Two hosted controls are left out on purpose. Claude Sonnet and Claude Haiku were driven through a command-line path that still exposed my account's own connected tools. Haiku replied that the bench's tools were not in its toolkit and named its email tools instead, and Sonnet answered "no data" without calling anything. That is the harness, not the models. It also shows a property of the bench: a model that refuses the tools still passes the 2030 task, so 1 of 6 is the score for doing nothing.

What reasoning effort costs

The same six tasks on GPT-5.6 Luna through a subscription, one lap with reasoning off and one at medium, 10 Sep. Token counts come from the API's own usage report and are priced at list, $1.25 per million input tokens and $10 per million output.

Luna, same six taskspassturnsinput tokensoutput tokensreasoning tokenscost at list
reasoning off2/61711,7227820$0.0225
reasoning medium5/62323,0547,2485,229$0.1013

Reasoning tokens are 72% of the output and about half the cost of the lap. Turning it on cost 4.5x as much and passed 2.5x as many tasks, which works out to 1.8x the cost per task solved. At list price, one lap is about two cents with reasoning off and ten with it on. A single lap did not visibly move the subscription's usage meter.

This is also what makes subscription value measurable: the bill is flat, and the list value of the tokens used is now counted per lap.

What this does not say

Related