llama-server ^
--model A:\models\Qwen3.6-27B-NEO-CODE-2T-OT-Q5_K_S.gguf ^
--mmproj "S:\LLMs\Qwen3.6-27B-mmproj-BF16.gguf" ^
--image-min-tokens 1024 ^
--ctx-size 128000 ^
--batch-size 2048 ^
--gpu-layers all ^
--cache-type-k q8_0 ^
--cache-type-v q8_0 ^
--spec-draft-type-k q8_0 ^
--spec-draft-type-v q8_0 ^
--fit off ^
--no-prefill-assistant ^
--api-key sk-ik-llama ^
--chat-template-file "S:\LLMs\Qwen3.6-27B.jinja" ^
--no-mmap ^
--main-gpu 1 ^
--split-mode layer ^
--tensor-split 51,49 ^
--parallel 1 --kv-unified ^
--ctx-checkpoints 32 ^
--checkpoint-min-step 1024 ^
--cache-ram 12288 ^
--reasoning-preserve ^
--temp 0.6 --top-k 20 --top-p 0.95 --min-p 0.0 ^
--repeat-penalty 1.0 --presence-penalty 0.0
.\run_qwen3.6_27B_BeeLlama_4_0_no_spec.bat
0.00.134.194 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.377.319 I srv load_model: loading model 'A:\models\Qwen3.6-27B-NEO-CODE-2T-OT-Q5_K_S.gguf'
0.09.356.747 I srv load_model: loaded multimodal model, 'S:\LLMs\Qwen3.6-27B-mmproj-BF16.gguf'
0.09.409.186 I srv load_model: initializing, n_slots = 1, n_ctx_slot = 128000, kv_unified = 'true'
0.09.434.268 I srv llama_server: model loaded
0.09.434.278 I srv llama_server: listening on http://127.0.0.1:8080
0.26.717.254 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
0.26.717.321 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
0.31.626.380 I slot print_timing: id 0 | task 0 | n_decoded = 114, tg = 37.86 t/s, tg_3s = 37.86 t/s
0.34.648.361 I slot print_timing: id 0 | task 0 | n_decoded = 228, tg = 37.79 t/s, tg_3s = 37.72 t/s
0.37.655.918 I slot print_timing: id 0 | task 0 | n_decoded = 342, tg = 37.83 t/s, tg_3s = 37.90 t/s
0.40.664.332 I slot print_timing: id 0 | task 0 | n_decoded = 455, tg = 37.76 t/s, tg_3s = 37.56 t/s
0.43.680.994 I slot print_timing: id 0 | task 0 | n_decoded = 569, tg = 37.77 t/s, tg_3s = 37.79 t/s
0.46.712.479 I slot print_timing: id 0 | task 0 | n_decoded = 683, tg = 37.74 t/s, tg_3s = 37.61 t/s
0.49.717.824 I slot print_timing: id 0 | task 0 | n_decoded = 796, tg = 37.72 t/s, tg_3s = 37.60 t/s
0.52.740.749 I slot print_timing: id 0 | task 0 | n_decoded = 910, tg = 37.72 t/s, tg_3s = 37.71 t/s
0.55.756.300 I slot print_timing: id 0 | task 0 | n_decoded = 1023, tg = 37.69 t/s, tg_3s = 37.47 t/s
0.58.764.205 I slot print_timing: id 0 | task 0 | n_decoded = 1136, tg = 37.68 t/s, tg_3s = 37.57 t/s
1.01.776.239 I slot print_timing: id 0 | task 0 | n_decoded = 1249, tg = 37.66 t/s, tg_3s = 37.52 t/s
1.04.779.745 I slot print_timing: id 0 | task 0 | n_decoded = 1362, tg = 37.66 t/s, tg_3s = 37.62 t/s
1.07.795.484 I slot print_timing: id 0 | task 0 | n_decoded = 1475, tg = 37.65 t/s, tg_3s = 37.47 t/s
1.10.804.798 I slot print_timing: id 0 | task 0 | n_decoded = 1588, tg = 37.64 t/s, tg_3s = 37.55 t/s
1.13.812.032 I slot print_timing: id 0 | task 0 | n_decoded = 1701, tg = 37.64 t/s, tg_3s = 37.58 t/s
1.16.812.209 I slot print_timing: id 0 | task 0 | n_decoded = 1813, tg = 37.62 t/s, tg_3s = 37.33 t/s
1.19.826.985 I slot print_timing: id 0 | task 0 | n_decoded = 1926, tg = 37.61 t/s, tg_3s = 37.48 t/s
1.22.836.027 I slot print_timing: id 0 | task 0 | n_decoded = 2038, tg = 37.59 t/s, tg_3s = 37.22 t/s
1.25.850.307 I slot print_timing: id 0 | task 0 | n_decoded = 2150, tg = 37.56 t/s, tg_3s = 37.16 t/s
1.28.865.548 I slot print_timing: id 0 | task 0 | n_decoded = 2262, tg = 37.54 t/s, tg_3s = 37.14 t/s
1.31.882.072 I slot print_timing: id 0 | task 0 | n_decoded = 2375, tg = 37.54 t/s, tg_3s = 37.46 t/s
1.34.890.801 I slot print_timing: id 0 | task 0 | n_decoded = 2487, tg = 37.53 t/s, tg_3s = 37.22 t/s
1.37.914.823 I slot print_timing: id 0 | task 0 | n_decoded = 2600, tg = 37.52 t/s, tg_3s = 37.37 t/s
1.40.933.866 I slot print_timing: id 0 | task 0 | n_decoded = 2713, tg = 37.51 t/s, tg_3s = 37.43 t/s
1.43.937.162 I slot print_timing: id 0 | task 0 | n_decoded = 2825, tg = 37.51 t/s, tg_3s = 37.29 t/s
1.46.946.557 I slot print_timing: id 0 | task 0 | n_decoded = 2937, tg = 37.49 t/s, tg_3s = 37.22 t/s
1.49.970.579 I slot print_timing: id 0 | task 0 | n_decoded = 3049, tg = 37.48 t/s, tg_3s = 37.04 t/s
1.52.991.436 I slot print_timing: id 0 | task 0 | n_decoded = 3162, tg = 37.48 t/s, tg_3s = 37.41 t/s
1.56.002.830 I slot print_timing: id 0 | task 0 | n_decoded = 3274, tg = 37.47 t/s, tg_3s = 37.19 t/s
1.59.009.598 I slot print_timing: id 0 | task 0 | n_decoded = 3386, tg = 37.46 t/s, tg_3s = 37.25 t/s
1.59.171.756 I slot print_timing: id 0 | task 0 | prompt eval time = 1898.09 ms / 2903 tokens ( 0.65 ms per token, 1529.43 tokens per second)
1.59.171.762 I slot print_timing: id 0 | task 0 | eval time = 90556.30 ms / 3392 tokens ( 26.70 ms per token, 37.46 tokens per second)
1.59.171.763 I slot print_timing: id 0 | task 0 | total time = 92454.39 ms / 6295 tokens
1.59.171.766 I slot print_timing: id 0 | task 0 | graphs reused = 3378
1.59.172.030 I slot release: id 0 | task 0 | stop processing: n_tokens = 6294, truncated = 0
Name and Version
version: 10814 (9b9277f)
built with MSVC 19.44.35228.0 for Windows AMD64
Operating systems
Windows
Which llama.cpp modules do you know to be affected?
llama-server
Command line
Problem description & steps to reproduce
Compared with not using speculative decoding, the generation speed is even worst.
(no spec v0.4.0) 37 t/s
(DFlash v0.4.0) 34 t/s
(DFlash v0.3.2) 88 t/s
no spec rerun is tested with just removing:
Relevant log output
Logs - DFlash
Logs - No Spec