# Known-good llama.cpp serving shape for big Qwen3.5 MoE models on one GPU # (122B-class shown; the 397B-class uses the same shape with MORE expert # blocks on CPU). Graphometer, 2026-08-09; prior-art credit added # 2026-08-10; 0.75 provenance corrected 2026-08-11. Measured details and # the full story: graphometer.ai/qwen35/ # The four ideas: # 1. -ngl 99 asks for every layer on the GPU (any number >= the layer count # means "all"), then --n-cpu-moe N walks expert blocks back to the CPU # until the model fits your VRAM. # 2. TUNE N FOR YOUR CARD BEFORE TRUSTING ANY NUMBER HERE: start with N high # (most expert blocks on CPU), confirm it loads, then lower N step by # step watching nvidia-smi until you are near your VRAM ceiling. The 38 # below is OUR starting point for the 122B on a 32 GB card; it is not # universal, and the 397B wants a higher N on the same card. # 3. The MTP guessing head ships inside the GGUF. Turned on WITH its # confidence gate it was a large speed win on our machine. Without the # gate, creative-text decode measured SLOWER than no speculation at all # (measured on the 122B; we keep the same gate on the 397B on that # evidence). The gate flag is --spec-draft-p-min and the CLI flag # defaults to 0.00 (off) in the builds we used. # THE 0.75 IS NOT OUR NUMBER. It circulates because it was # llama.cpp's own shipped default for about fifteen months: set on # 2025-02-19 (commit abd4d0bc4, PR #11954) with no published # benchmark behind it, and dropped to 0.0 on 2026-05-19 (commit # d14ce3dab). The write-ups that recommend it as an acceptance # "sweet spot" (llama.cpp discussion #25198 uses it as a throughput # lever, a Chinese-language tuning writeup dated 2026-05-18 calls # >= 0.75 the sweet spot and gives guidance across 0.6 to 0.9, and # it appears in the reproduction command of an existing llama.cpp # issue) are, in effect, re-supplying by hand a default that was # silently dropped. What we measured is not the value but the # REVERSAL on this size class: the gate turning a measured negative # into a result above the no-speculation baseline on creative text. # NEITHER EFFECT IS NEW: # the creative slowdown was published on 2026-05-12 on a dense # Qwen3.6-27B, and the gated reversal on 2026-05-22 on a 27B (29.0 t/s # with no speculation, 48.9 with this gate). Ours is the independent # confirmation on a much larger mixture-of-experts model with most of # its expert blocks on the CPU; the field card carries both citations. # Tune it yourself; the knee for your model and quantization is not # something we located. # 4. --no-mmap copies the weights into RAM instead of memory-mapping them. # On our 397B runs, leaving mmap on cost 82 percent of prompt-reading # speed. If the weights fit in your RAM, copy them in. This also means # the weights MUST fit in RAM: for multi-file quants below, that is the # sum of all shards. # Note the file name: big quants ship as multiple shards. Point -m at the # FIRST shard and llama.cpp finds the rest. llama-server \ -m Qwen3.5-122B-A10B-UD-Q4_K_S-00001-of-00003.gguf \ -ngl 99 \ --n-cpu-moe 38 \ -c 65536 \ --no-mmap \ --spec-type draft-mtp \ --spec-draft-n-max 6 \ --spec-draft-p-min 0.75 \ --host 127.0.0.1 --port 8080 # Request-side settings that matter as much as the server flags: # - Thinking bills inside the output budget. Give requests thousands of # tokens of headroom (we use 4096 as the floor). # - Sampling for thinking-family Qwens: temperature 0.6, top-k 20, # top-p 0.95. # - Warm speeds only: never judge the first answer after a cold start. # All figures behind this file are one machine, one configuration, dated # recorded runs; the field card names the exact quants and numbers.