淺聊PrismaQuant, INT4 Autoround以及單純NVFP4量化在RTX Pro 4500下的表現
-
Lorbus/Qwen3.6-27B-int4-AutoRound
測試咒語
cd /home/rw/llama-benchy && source .venv/bin/activate && llama-benchy --base-url "http://localhost:7380/v1" --model "Qwen3.6-27B-int4-AutoRound" --tokenizer "/home/rw/vllm/models/Lorbus/Qwen3.6-27B-int4-AutoRound" --pp 2048 --tg 480 --depth 0 1000 5000 10000 20000 50000 100000 150000 200000 210000 --latency-mode generation --skip-coherence --concurrency 1 --save-result "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/lorbus-qwen3.6.md" --format md --emit-progress "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/lorbus-qwen3.6.progress.jsonl"model test t/s peak t/s ttfr (ms) est_ppt (ms) e2e_ttft (ms) Qwen3.6-27B-int4-AutoRound pp2048 2265.07 ± 142.42 979.34 ± 59.72 908.36 ± 59.72 979.34 ± 59.72 Qwen3.6-27B-int4-AutoRound tg480 75.79 ± 0.56 98.33 ± 5.44 Qwen3.6-27B-int4-AutoRound pp2048 @ d1000 2398.37 ± 34.22 1342.38 ± 18.09 1271.40 ± 18.09 1342.38 ± 18.09 Qwen3.6-27B-int4-AutoRound tg480 @ d1000 75.11 ± 1.80 93.33 ± 4.19 Qwen3.6-27B-int4-AutoRound pp2048 @ d5000 2358.05 ± 15.15 3060.58 ± 19.20 2989.60 ± 19.20 3060.58 ± 19.20 Qwen3.6-27B-int4-AutoRound tg480 @ d5000 77.10 ± 2.79 92.33 ± 2.62 Qwen3.6-27B-int4-AutoRound pp2048 @ d10000 2224.06 ± 9.02 5488.80 ± 22.23 5417.81 ± 22.23 5488.80 ± 22.23 Qwen3.6-27B-int4-AutoRound tg480 @ d10000 76.34 ± 5.30 90.00 ± 4.90 Qwen3.6-27B-int4-AutoRound pp2048 @ d20000 2092.20 ± 2.93 10609.36 ± 14.97 10538.38 ± 14.97 10610.69 ± 15.00 Qwen3.6-27B-int4-AutoRound tg480 @ d20000 74.81 ± 1.35 95.33 ± 3.68 Qwen3.6-27B-int4-AutoRound pp2048 @ d50000 1795.50 ± 0.72 29059.50 ± 11.27 28988.52 ± 11.27 29061.94 ± 11.21 Qwen3.6-27B-int4-AutoRound tg480 @ d50000 71.68 ± 3.47 93.00 ± 2.16 Qwen3.6-27B-int4-AutoRound pp2048 @ d100000 1468.22 ± 0.25 69576.21 ± 11.65 69505.23 ± 11.65 69581.18 ± 12.20 Qwen3.6-27B-int4-AutoRound tg480 @ d100000 72.35 ± 1.91 90.33 ± 4.50 Qwen3.6-27B-int4-AutoRound pp2048 @ d150000 1240.97 ± 1.12 122595.75 ± 110.82 122524.77 ± 110.82 122603.15 ± 110.62 Qwen3.6-27B-int4-AutoRound tg480 @ d150000 67.56 ± 0.87 84.33 ± 4.78 Qwen3.6-27B-int4-AutoRound pp2048 @ d200000 1076.24 ± 0.13 187806.97 ± 23.37 187735.98 ± 23.37 187815.98 ± 23.34 Qwen3.6-27B-int4-AutoRound tg480 @ d200000 61.65 ± 1.78 76.33 ± 1.89 Qwen3.6-27B-int4-AutoRound pp2048 @ d210000 1047.66 ± 0.10 202474.09 ± 19.62 202403.11 ± 19.62 202483.65 ± 19.76 Qwen3.6-27B-int4-AutoRound tg480 @ d210000 59.64 ± 1.94 76.67 ± 3.30 -
rdtand/Qwen3.6-27B-PrismaAURA-5.5bit-vllm
cd /home/rw/llama-benchy && source .venv/bin/activate && llama-benchy --base-url "http://localhost:7380/v1" --model "Qwen3.6-27B-PrismaAURA-5.5bit-vllm" --tokenizer "/home/rw/vllm/models/rdtand/Qwen3.6-27B-PrismaAURA-5.5bit-vllm" --pp 2048 --tg 480 --depth 0 1000 5000 10000 20000 50000 100000 130000 --latency-mode generation --skip-coherence --concurrency 1 --save-result "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/rdtand-prismaaura.md" --format md --emit-progress "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/rdtand-prismaaura.progress.jsonl"model test t/s peak t/s ttfr (ms) est_ppt (ms) e2e_ttft (ms) Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 5063.05 ± 507.65 513.19 ± 43.93 409.10 ± 43.93 513.19 ± 43.93 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 63.48 ± 2.54 76.67 ± 5.25 Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 @ d1000 4938.83 ± 600.47 730.08 ± 72.32 625.99 ± 72.32 730.08 ± 72.32 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 @ d1000 69.03 ± 2.59 83.67 ± 2.87 Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 @ d5000 4905.40 ± 59.54 1541.22 ± 17.49 1437.13 ± 17.49 1541.22 ± 17.49 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 @ d5000 65.70 ± 5.04 84.00 ± 3.74 Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 @ d10000 4905.96 ± 63.31 2560.56 ± 31.78 2456.47 ± 31.78 2560.56 ± 31.78 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 @ d10000 67.96 ± 2.99 85.33 ± 3.68 Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 @ d20000 4389.10 ± 17.37 5127.60 ± 19.84 5023.51 ± 19.84 5128.85 ± 19.71 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 @ d20000 66.39 ± 1.81 80.33 ± 3.09 Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 @ d50000 3278.48 ± 6.13 15980.09 ± 29.65 15876.00 ± 29.65 15982.46 ± 29.54 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 @ d50000 66.58 ± 2.41 81.67 ± 2.36 Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 @ d100000 2333.58 ± 0.11 43834.67 ± 2.13 43730.58 ± 2.13 43839.15 ± 2.17 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 @ d100000 54.77 ± 1.31 79.33 ± 1.25 Qwen3.6-27B-PrismaAURA-5.5bit-vllm pp2048 @ d130000 1986.20 ± 0.93 66587.42 ± 31.08 66483.33 ± 31.08 66593.71 ± 30.32 Qwen3.6-27B-PrismaAURA-5.5bit-vllm tg480 @ d130000 59.76 ± 2.42 74.33 ± 3.86 -
rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm
cd /home/rw/llama-benchy && source .venv/bin/activate && llama-benchy --base-url "http://localhost:7380/v1" --model "Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm" --tokenizer "/home/rw/vllm/models/rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm" --pp 2048 --tg 480 --depth 0 1000 5000 10000 20000 50000 100000 150000 200000 210000 --latency-mode generation --skip-coherence --concurrency 1 --save-result "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/rdtand-prismascout.md" --format md --emit-progress "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/rdtand-prismascout.progress.jsonl"model test t/s peak t/s ttfr (ms) est_ppt (ms) e2e_ttft (ms) Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 5856.51 ± 717.09 463.07 ± 47.30 355.65 ± 47.30 463.07 ± 47.30 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 75.13 ± 2.16 92.33 ± 2.05 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d1000 6657.31 ± 235.17 565.99 ± 16.34 458.57 ± 16.34 565.99 ± 16.34 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d1000 72.04 ± 7.78 88.67 ± 5.91 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d5000 6241.79 ± 108.18 1237.18 ± 19.42 1129.77 ± 19.42 1237.18 ± 19.42 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d5000 70.70 ± 1.92 88.33 ± 2.62 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d10000 5324.02 ± 168.79 2372.81 ± 71.23 2265.40 ± 71.23 2372.81 ± 71.23 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d10000 71.88 ± 4.87 87.33 ± 1.25 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d20000 4831.87 ± 20.01 4670.67 ± 18.94 4563.26 ± 18.94 4671.97 ± 18.87 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d20000 68.59 ± 3.03 86.67 ± 3.30 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d50000 3541.51 ± 2.18 14804.18 ± 9.14 14696.76 ± 9.14 14806.50 ± 9.23 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d50000 69.34 ± 2.55 84.00 ± 5.72 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d100000 2460.18 ± 0.29 41587.44 ± 4.82 41480.03 ± 4.82 41591.90 ± 5.14 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d100000 63.53 ± 4.67 79.67 ± 3.68 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d150000 1883.43 ± 0.66 80837.23 ± 28.22 80729.81 ± 28.22 80845.53 ± 29.14 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d150000 63.47 ± 2.92 78.67 ± 3.30 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d200000 1525.83 ± 0.26 132525.85 ± 22.57 132418.43 ± 22.57 132535.59 ± 23.55 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d200000 61.81 ± 1.34 74.00 ± 1.41 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm pp2048 @ d210000 1469.93 ± 1.33 144365.98 ± 130.60 144258.56 ± 130.60 144372.64 ± 135.28 Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm tg480 @ d210000 56.70 ± 2.40 74.67 ± 1.70 -
sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP
cd /home/rw/llama-benchy && source .venv/bin/activate && llama-benchy --base-url "http://localhost:7380/v1" --model "Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP" --tokenizer "/home/rw/vllm/models/sakamakismile/Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP" --pp 2048 --tg 480 --depth 0 1000 5000 10000 20000 50000 100000 150000 200000 210000 --latency-mode generation --skip-coherence --concurrency 1 --save-result "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/sakamakismile-qwen3.6.md" --format md --emit-progress "/home/rw/vllm/benchmark-results/2026-07-10-vllm-generation/sakamakismile-qwen3.6.progress.jsonl"model test t/s peak t/s ttfr (ms) est_ppt (ms) e2e_ttft (ms) Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 8586.53 ± 13.35 336.52 ± 0.32 238.67 ± 0.32 336.52 ± 0.32 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 75.96 ± 2.83 92.00 ± 2.83 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d1000 7766.18 ± 35.67 490.46 ± 1.80 392.61 ± 1.80 490.46 ± 1.80 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d1000 67.58 ± 4.10 89.33 ± 4.19 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d5000 6792.75 ± 10.01 1135.57 ± 1.41 1037.73 ± 1.41 1135.57 ± 1.41 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d5000 70.23 ± 3.33 88.00 ± 3.56 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d10000 5970.12 ± 6.61 2116.07 ± 2.24 2018.22 ± 2.24 2116.07 ± 2.24 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d10000 72.55 ± 6.17 92.67 ± 2.87 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d20000 5155.64 ± 2.62 4374.46 ± 2.13 4276.61 ± 2.13 4375.80 ± 2.06 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d20000 76.50 ± 0.36 94.00 ± 4.32 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d50000 3711.32 ± 0.59 14122.15 ± 2.27 14024.31 ± 2.27 14124.25 ± 2.20 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d50000 72.70 ± 2.83 92.33 ± 3.77 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d100000 2540.93 ± 0.52 40259.90 ± 7.98 40162.05 ± 7.98 40264.76 ± 7.79 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d100000 65.96 ± 2.32 78.67 ± 2.87 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d150000 1932.50 ± 0.48 78777.95 ± 19.55 78680.11 ± 19.55 78784.65 ± 19.37 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d150000 61.13 ± 0.73 77.67 ± 1.25 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d200000 1558.91 ± 0.24 129706.83 ± 19.70 129608.98 ± 19.70 129715.93 ± 19.79 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d200000 59.02 ± 2.99 70.33 ± 2.36 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP pp2048 @ d210000 1500.01 ± 0.12 141463.13 ± 11.03 141365.28 ± 11.03 141472.49 ± 11.21 Huihui-Qwen3.6-27B-abliterated-NVFP4-MTP tg480 @ d210000 57.59 ± 0.42 73.00 ± 3.56 -
@566656661 你把它复制到帖子开头吧,效果会更好一点。
-
我感觉vllm对mtp的支持始终有问题,打开mtp会造成tool调用出错、loop、边界识别错误等问题,num_speculative_tokens降低到1凑合着能用,彻底关闭mtp就好不少,但是速度又太慢。
我让codex给vllm0.24.0打了几个还没merge的pr,似乎好点了。
哦,对了,jinja chat template模板也有影响,llama.cpp似乎在脚本方面比较稳定,但是存在prompt fill巨慢的问题,总之,各种mtp、dflash加速确实很快,很爽,但是真要用起来,还需要很多调教。目前見到的情況是因爲MTP跟enable-prefix-caching一起用會導致推論精度下降, 導致Tool Call出現問題, 有人報告Tool Call精準度下降到50%
-
,系统 取消固定了此主题
-
,5 566656661 引用了 此主题