9 minutes, 42 seconds

In this post I want to explore the pitfalls and working settings for running LLMs locally on an RDNA2 GPU. It mainly focuses on cache quantization but also offers more settings to improve performance and some rambling as well.

2026-03-28 -> The post was updated with a short description, nicer tables as well as further research into how different model architectures react to mixed quantization, based on a comment on Reddit by user a_beautiful_rhind

Rambling

I've been struggling with running LLMs locally on my AMD Radeon 6950XT. Running "frontier" open-weight models on 16GB VRAM and 32GB RAM is already not the easiest, and a lot of people seem to assume you have more than that. Just recently Mistral released their new Mistral 4 "Small" model at a whopping 120B parameters. At least it's MoE I guess.

It's not made easier by this generation, which is one of the latest ones that lack both FP8 and matrix calculation acceleration. Just one newer generation and you'd get both! At least the card has a funny name.

The best "bang for the buck" model I've seen so far are either the Qwen 3.5 or similar ones. GLM4.7-Flash-Reap is genuinely one of the fastest models I've run so far. If you trust the posted benchmark numbers though, even Qwen3.5-9B completely smokes it.

Using a finetune for coding called Omnicoder based on Qwen3.5-9B seemed like a nice choice. The only really valid alternatives seem to be Qwen3.5-35B-A3B, Qwen3-Coder-Next or the new Nemotron Cascade 2. The latter two are unreasonably large for my (V)RAM though and Qwen3.5-35B-A3B is awfully sensitive in my testings. I've had many instances where it would just completely run amok.

The 9B Omnicoder seemed like a good fit, as it would comfortably sit in my VRAM along with all the context I'd need and then some. Doing that I've been getting some pretty good numbers like 200 tok/s prompt processing and 20 tok/s token generation. Or so I thought.

Looking on Reddit and other sources, lots of people make it seem like this is a good performance for "low VRAM" setups, which 16GB is in the local LLMs space. It doesn't help that virtually everyone seems to be running Nvidia, specifically Blackwell, or Apple hardware. There's rarely any frame of reference to AMD. One comment I found only had that throughput on the much better AMD Radeon AI Pro 9700, albeit with a bigger model as well. I felt content with the performance I was getting, so to speak.

Cache Quantization

One recurring theme in performance guides on Reddit and otherwhere is KV cache quantization. Quantizing the KV cache can result in significant memory savings and usually insigificant quality loss. But, some models react more sensitively to specific quantizations. Some people thus advise to try different quantizations for key and value in the KV cache. You can easily try this out yourself with -ctv q8_0 -ctk f16. If you know any programming or data engineering this may make sense. It's a lot more important for keys to be unique and specific than it is for the value to be 100% accurate. So it's fine then, shouldn't everyone run with this?

Nope

I was going through the list of options llama.cpp supports. I kept reading "loading ggml-haswell backend" when launching llama.cpp and kept wondering if there's not an option to use a backend for newer CPU archs (like my Zen-3 5950X).

By chance the two options for cache quantization jump out to me again and I begin wondering, isn't it bad to have a differing layout for GPUs?

Benchmarks!

model size params backend ngl n_batch type_k type_v fa test t/s
qwen35 9B Q6_K 6.84 GiB 8.95 B Vulkan 99 1024 f16 q8_0 1 pp5000 334.27 ± 1.42
qwen35 9B Q6_K 6.84 GiB 8.95 B Vulkan 99 1024 f16 q8_0 1 tg128 53.53 ± 0.23
qwen35 9B Q6_K 6.84 GiB 8.95 B Vulkan 99 1024 q8_0 q8_0 1 pp5000 952.79 ± 0.46
qwen35 9B Q6_K 6.84 GiB 8.95 B Vulkan 99 1024 q8_0 q8_0 1 tg128 63.37 ± 0.06

My standard benchmark run tests a 5000 token prompt as I'm usually using the models for coding, so the default 512 token prompt isn't really representative.

As you can tell, same quantization can double throughput / differing quantization can halve throughput.

But L3tum, isn't f16 just more bandwidth and thus the lower performance can be explained by that?

Good point, but no! Here's the same test with KV cache at f16. As you can see, same performance as non-quantized.

model size params backend ngl n_batch fa test t/s
qwen35 9B Q6_K 6.84 GiB 8.95 B Vulkan 99 1024 1 pp5000 951.76 ± 0.30
qwen35 9B Q6_K 6.84 GiB 8.95 B Vulkan 99 1024 1 tg128 63.49 ± 0.02

What about other models?

After a Reddit user pointed me towards benchmarking those as well, here's the results for a few different models I had laying around.

GLM4.7-Flash Reap

This model is a little weird, in that it seems fine with mixed quantization for values, but destroys itself on any quantization that touches keys. It was running right at the maximum of VRAM though, so may have been an issue.

model size params backend ngl n_batch type_k type_v fa test t/s
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 q8_0 q8_0 1 pp5000 137.42 ± 0.21
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 q8_0 q8_0 1 tg128 107.44 ± 0.06
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 q8_0 f16 1 pp5000 137.73 ± 0.17
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 q8_0 f16 1 tg128 107.36 ± 0.03
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 f16 q8_0 1 pp5000 289.24 ± 0.18
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 f16 q8_0 1 tg128 118.25 ± 0.09
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 f16 f16 1 pp5000 289.08 ± 0.37
deepseek2 30B.A3B Q4_K - Medium 13.14 GiB 23.00 B Vulkan 99 1024 f16 f16 1 tg128 118.30 ± 0.08

Phi 4 Mini Instruct

This one seems to support my Qwen3.5 theory in that it gets destroyed in mixed quantization and is fine in either full quantization or no quantization. Notably quantization here does offer a little bit of performance advantage over no quantization.

model size params backend ngl n_batch type_k type_v fa test t/s
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 q8_0 q8_0 1 pp5000 692.43 ± 0.17
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 q8_0 q8_0 1 tg128 115.99 ± 0.03
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 q8_0 f16 1 pp5000 169.72 ± 0.81
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 q8_0 f16 1 tg128 51.15 ± 0.21
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 f16 q8_0 1 pp5000 113.79 ± 0.51
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 f16 q8_0 1 tg128 48.94 ± 3.93
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 f16 f16 1 pp5000 628.43 ± 0.81
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 1024 f16 f16 1 tg128 115.77 ± 0.45

IQuestCoder-14B

This one's a fairly old model I accidentally downloaded a while back cause it showed up with many downloads on HuggingFace. Same behaviour as Phi4 (or Phi3) though.

model size params backend ngl n_batch type_k type_v fa test t/s
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 q8_0 q8_0 1 pp5000 323.06 ± 0.09
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 q8_0 q8_0 1 tg128 45.02 ± 0.01
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 q8_0 f16 1 pp5000 101.46 ± 0.37
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 q8_0 f16 1 tg128 30.99 ± 0.21
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 f16 q8_0 1 pp5000 69.65 ± 0.21
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 f16 q8_0 1 tg128 30.81 ± 0.36
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 f16 f16 1 pp5000 295.63 ± 0.15
llama 3B Q6_K 11.03 GiB 14.44 B Vulkan 99 1024 f16 f16 1 tg128 45.20 ± 0.04

Devstral Small 2

One of the most used models for coding assistance. Also shows the same effect.

model size params backend ngl n_batch type_k type_v fa test t/s
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 q8_0 q8_0 1 pp5000 199.11 ± 0.08
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 q8_0 q8_0 1 tg128 31.42 ± 0.03
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 q8_0 f16 1 pp5000 81.42 ± 0.47
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 q8_0 f16 1 tg128 21.30 ± 0.12
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 f16 q8_0 1 pp5000 58.67 ± 0.11
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 f16 q8_0 1 tg128 21.18 ± 0.08
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 f16 f16 1 pp5000 188.07 ± 0.08
mistral3 14B IQ4_XS - 4.25 bpw 11.89 GiB 23.57 B Vulkan 99 1024 f16 f16 1 tg128 31.37 ± 0.04

Explanations?

Running more models made me check out my system utilization as well and made me notice that my CPU usage spikes to 66% and my GPU falls flat when mixed quantization is used and the performance craters. I think it's kind of impressive that my 5950X can still run prompt processing at that speed.

Though this does point towards a missing optimization in the Vulkan backend for llama.cpp when it comes to mixed quantization. It would be cool if other people using other backends could re-run the tests to see what is and is not affected.

66% of my 32-thread 5950X is roughly 24 threads that are being used. Not sure if that's a default in llama.cpp or some bottleneck elsewhere in the system.

Other settings that matter

My full "performance optimization" for my card:

-fa on --prio -1 --parallel 1 --batch-size 1024 --ubatch-size 1024 -ctv q8_0 -ctk q8_0

Here's a breakdown for the different options:

  • -fa on -> Flash Attention. Used to be the case that some models didn't like it, but nowadays can basically be always on for a nice speed boost
  • --prio -1 -> Sets process priority for llama-server to -1 (less than normal). This make everything else like windows or CLIs still responsive rather than freezing your desktop
  • --parallel 1 -> Theoretically llama-server should already only launch one client-slot and not allocate more context than for that one client, but some report setting it explicitly increases performance
  • --batch-size 1024 --ubatch-size 1024 -> Sets the batch size for processing and memory allocation. In my testing there's no change in throughput from 512 to 4096, but 1024 is more stable than 512. Lower values are slower
  • -ctv q8_0 -ctk q8_0 -> KV Cache quantization. You can theoretically leave it at f16 (default) for the 9B since it should fit into VRAM

Bonus llama.cpp version recommendation

The best version so far seems to be b8416. I've had some issues with versions after this as of today (2026-03-23) but YMMV.

Next Post