<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/3bd2d0d5-f6eb-4bb2-bc96-11d4ed0a28fb.jpg" alt="807750669_28464673439809579_8935132091605666435_n.jpg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/36932ee4-b95b-40f7-bbae-2d8fb5e2f193.jpg" alt="808521371_28465044669772456_1491874540721602469_n.jpg" class=" img-fluid img-markdown" /></p>
<h2>雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果<br />
硬體：<br />
Radeon AI PRO R9700 32GB ×2<br />
總 VRAM：64GB<br />
System RAM：224GB<br />
RDNA4 / gfx1201<br />
模型：<br />
Qwen3.8-Flash-Next IQ3_M<br />
48 層 MoE<br />
Context：128K<br />
KV Cache：Q4_0<br />
Windows + Vulkan<br />
這次是在原本 Vulkan b10751 配置上繼續調整。<br />
實測結果<br />
測試 原配置 調整後<br />
Coding 56.67 tok/s ----＞63.19 tok/s<br />
Fibonacci 56.54 tok/s ---＞62.73 tok/s<br />
繁中散文 56.27 tok/s ----＞61.21 tok/s<br />
8K Context 44.85 tok/s----＞ 60.50 tok/s<br />
32K Context 42.90 tok/s ----＞57.46 tok/s<br />
64K Context 43.19 tok/s----＞ 54.50 tok/s<br />
128K Context 45.90 tok/s ----＞50.80 tok/s<br />
64K Context 的速度保持率由 80.0% → 87.4%。<br />
48 層 MoE Expert 維持 Zero Spill。</h2>
<h2>這次速度上升的主要調整<br />
GGML_VK_RM_KQ=1<br />
原本 Vulkan kernel 使用 rm_kq=2。<br />
改成：<br />
GGML_VK_RM_KQ=1<br />
這項調整的目的，是降低單個 Workgroup 的寄存器壓力，讓 RDNA4 Compute Unit 可以維持更多 active wavefront。<br />
在這次測試紀錄中，這是短文本 Decode 提升幅度最大的單項調整，紀錄約 +15.28%。<br />
改用 Graphics Queue<br />
加入：<br />
GGML_VK_ALLOW_GRAPHICS_QUEUE=1<br />
原本 Vulkan 主要走 Compute Queue。<br />
這次測試中，改用 Graphics Queue 後，Windows AMD Driver 下的 queue submission / synchronization 開銷降低，紀錄約再增加 1.15 tok/s。<br />
Tensor Split 從 0.48,0.52 改成 0.47,0.53<br />
原本：<br />
--tensor-split 0.48,0.52<br />
調整成：<br />
--tensor-split 0.47,0.53<br />
原因是兩張 GPU 的實際負載並不完全對稱。<br />
主卡同時承擔 Windows 顯示相關負載，所以把稍多模型權重分給第二張卡，讓兩張 GPU 每層運算完成時間更接近。<br />
這次測試紀錄約增加：<br />
+0.76 tok/s。<br />
CPU Threads：4 → 16<br />
原本：<br />
--threads 4<br />
改成：<br />
--threads 16<br />
因為目前 PLE 是透過：<br />
-ot "per_layer_token_embd.weight=CPU"<br />
放在 System RAM。<br />
所以 Decode 並不是完全只有 GPU 工作，CPU 仍要處理 PLE 對應的 Host memory 存取。<br />
增加 CPU thread 數後，這部分延遲下降。<br />
本次紀錄約增加：<br />
+0.50 tok/s。<br />
長 Context 的改善來自 QSA Pooled-Key Cache<br />
短文本從約 56 → 63 tok/s，主要是前面幾項 Vulkan / scheduling 調整。<br />
但是 8K～64K 的提升幅度更大，不只是 RM_KQ。<br />
這次另外加入 QSA incremental pooled-key cache，避免長 Context decode 時重複計算部分 pooled-key 特徵。<br />
所以長 Context 的改善比較明顯：<br />
8K：44.85 → 60.50<br />
32K：42.90 → 57.46<br />
64K：43.19 → 54.50<br />
128K：45.90 → 50.80 tok/s<br />
同時維持：<br />
-ncmoe 0<br />
讓 48 層 MoE Expert 都留在 GPU。</h2>
<p dir="auto">目前主要配置<br />
GGML_VK_RM_KQ=1<br />
GGML_VK_ALLOW_GRAPHICS_QUEUE=1<br />
--tensor-split 0.47,0.53<br />
--threads 16<br />
-ngl 999<br />
-ncmoe 0<br />
-ot "per_layer_token_embd.weight=CPU"<br />
--ctx-size 131072<br />
-fa on<br />
-ctk q4_0<br />
-ctv q4_0<br />
所以這次不是換模型、降量化精度或使用 MTP 得到的提升。<br />
主要差異是：<br />
<strong>Vulkan kernel 調整<br />
Queue 選擇<br />
雙 GPU 負載重新分配<br />
CPU PLE 路徑調整<br />
長 Context cache 優化</strong><br />
最後的結果是：<br />
短文本約 63 tok/s<br />
64K 約 54.5 tok/s<br />
128K 約 50.8 tok/s<br />
48 層 MoE Zero Spill</p>
]]></description><link>https://lcz.me/topic/1692</link><generator>RSS for Node</generator><lastBuildDate>Tue, 15 Sep 2026 01:28:24 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1692.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 14 Sep 2026 04:53:10 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Tue, 15 Sep 2026 01:02:10 GMT]]></title><description><![CDATA[<p dir="auto">两者不是一个量级，也不是同一种用法，别只盯量化位数。</p>
<p dir="auto">Flash-Next 是 MoE，激活参数远小于总量，IQ3_M 的损失主要在推理/代码这类精度敏感任务；但它的总容量和长上下文能力，通常还是高于 27B dense 的 Q8。显存上：27B Q8 约 28–29G，单卡 32G 放得下（留点 KV 池）；Flash-Next IQ3_M 体积大得多，单卡 32G 基本要 offload，双卡 64G 才舒服。</p>
<p dir="auto">实用分配：长上下文/agent 类任务走 Flash-Next IQ3_M；数学、代码、最终定稿用 27B Q4/Q5 或 Q8。若显存放得下，IQ4_XS 比 IQ3_M 更值，别为省几 G 牺牲推理质量。</p>
]]></description><link>https://lcz.me/post/18244</link><guid isPermaLink="true">https://lcz.me/post/18244</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 15 Sep 2026 01:02:10 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Tue, 15 Sep 2026 01:00:51 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E7%AE%97%E7%9B%A4%E4%BF%A0" aria-label="Profile: 算盤俠">@<bdi>算盤俠</bdi></a> 寫個2048遊戲就有差了，hermes agent 的使用體感也是明顯解決能力比較強</p>
]]></description><link>https://lcz.me/post/18242</link><guid isPermaLink="true">https://lcz.me/post/18242</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Tue, 15 Sep 2026 01:00:51 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 14:51:11 GMT]]></title><description><![CDATA[<p dir="auto">Qwen3.8-Flash-Next IQ3_M 对比27b的Q8差距大吗？</p>
]]></description><link>https://lcz.me/post/18172</link><guid isPermaLink="true">https://lcz.me/post/18172</guid><dc:creator><![CDATA[算盤俠]]></dc:creator><pubDate>Mon, 14 Sep 2026 14:51:11 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 14:17:23 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a> 我GIGABYTE技嘉 Z890 AERO G/ATX/1851腳位主板是能直插的，沒有問題。</p>
]]></description><link>https://lcz.me/post/18161</link><guid isPermaLink="true">https://lcz.me/post/18161</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 14:17:23 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 14:15:22 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stormaround" aria-label="Profile: stormaround">@<bdi>stormaround</bdi></a> 因為平常前端還是要windows拿來跑arc gis pro等等專業工具，背景跑hermes不影響使用。</p>
]]></description><link>https://lcz.me/post/18160</link><guid isPermaLink="true">https://lcz.me/post/18160</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 14:15:22 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 14:12:20 GMT]]></title><description><![CDATA[<p dir="auto">怎么用windwos跑，速度不太快，github上有个现成的，我用e5 洋垃圾 + 128g ddr3 + 双r9700,一般decode速度都有70以上，zx-bench 4并发decode平均也有50多</p>
]]></description><link>https://lcz.me/post/18155</link><guid isPermaLink="true">https://lcz.me/post/18155</guid><dc:creator><![CDATA[stormaround]]></dc:creator><pubDate>Mon, 14 Sep 2026 14:12:20 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 13:33:35 GMT]]></title><description><![CDATA[<p dir="auto">这个硬件配置不用照抄。知道芯片组 就可以了。有俩个买完发现显卡不能直接插上的帖子了。针对实际情况和商家探讨好。再入手才是正确的。插不上商家给你换或退也不会有任何问题。<br />
并且9700 主要还是当前硬件价格造成的问题。组上后并不是很舒适。现在 市场流通的 A100 开始增加了。当然我不推荐A100 。只是分析 硬件换代潮还需要等多久。</p>
]]></description><link>https://lcz.me/post/18136</link><guid isPermaLink="true">https://lcz.me/post/18136</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Mon, 14 Sep 2026 13:33:35 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 13:22:25 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 行</p>
]]></description><link>https://lcz.me/post/18129</link><guid isPermaLink="true">https://lcz.me/post/18129</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 13:22:25 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 13:20:24 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/geekyang" aria-label="Profile: Geekyang">@<bdi>Geekyang</bdi></a> GIGABYTE技嘉 Z890 AERO G/ATX/1851腳位，270kplus,1200w電源。</p>
]]></description><link>https://lcz.me/post/18128</link><guid isPermaLink="true">https://lcz.me/post/18128</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 13:20:24 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 13:18:15 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/nami-ryuu" aria-label="Profile: nami-ryuu">@<bdi>nami-ryuu</bdi></a> 夠用</p>
]]></description><link>https://lcz.me/post/18125</link><guid isPermaLink="true">https://lcz.me/post/18125</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 13:18:15 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 11:19:21 GMT]]></title><description><![CDATA[<p dir="auto">把主板和详细配置贴一下。</p>
]]></description><link>https://lcz.me/post/18095</link><guid isPermaLink="true">https://lcz.me/post/18095</guid><dc:creator><![CDATA[Geekyang]]></dc:creator><pubDate>Mon, 14 Sep 2026 11:19:21 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 10:51:46 GMT]]></title><description><![CDATA[<p dir="auto">我弟你让AI整理下你的帖子，我帮你整成了markdown格式，下不为例。</p>
]]></description><link>https://lcz.me/post/18093</link><guid isPermaLink="true">https://lcz.me/post/18093</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Mon, 14 Sep 2026 10:51:46 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 09:49:18 GMT]]></title><description><![CDATA[<p dir="auto">@128Gram够用吗ddr4， pcie4X16</p>
]]></description><link>https://lcz.me/post/18082</link><guid isPermaLink="true">https://lcz.me/post/18082</guid><dc:creator><![CDATA[nami ryuu]]></dc:creator><pubDate>Mon, 14 Sep 2026 09:49:18 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 08:01:26 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bunsei" aria-label="Profile: Bunsei">@<bdi>Bunsei</bdi></a> 我RAM夠大就不用放到SSD上面</p>
]]></description><link>https://lcz.me/post/18062</link><guid isPermaLink="true">https://lcz.me/post/18062</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 08:01:26 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 07:40:23 GMT]]></title><description><![CDATA[<p dir="auto">可以试试看R9V？ 这个项目好像针对于2*R9700跑Qwen3.8Flash有专门的优化，而且可以把N-gram 表放在固态上。</p>
]]></description><link>https://lcz.me/post/18058</link><guid isPermaLink="true">https://lcz.me/post/18058</guid><dc:creator><![CDATA[Bunsei]]></dc:creator><pubDate>Mon, 14 Sep 2026 07:40:23 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 06:55:47 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E5%BC%A0%E5%85%89%E7%92%9E" aria-label="Profile: 张光璞">@<bdi>张光璞</bdi></a> PP是只預填充速度而已，TP=2設置只能在Linux環境下跑~</p>
]]></description><link>https://lcz.me/post/18047</link><guid isPermaLink="true">https://lcz.me/post/18047</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 06:55:47 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 06:52:42 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/neo" aria-label="Profile: neo">@<bdi>neo</bdi></a> 只能跑到ddr5 4400 頻率，因為是美光ddr5 5600 128G(64+64)跟 ddr5 5600 96G(48+48)混搭一共四張會降頻。</p>
]]></description><link>https://lcz.me/post/18046</link><guid isPermaLink="true">https://lcz.me/post/18046</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Mon, 14 Sep 2026 06:52:42 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 05:31:18 GMT]]></title><description><![CDATA[<p dir="auto">感谢分享，prefill速度不错，内存是DDR几的呀？</p>
]]></description><link>https://lcz.me/post/18030</link><guid isPermaLink="true">https://lcz.me/post/18030</guid><dc:creator><![CDATA[neo]]></dc:creator><pubDate>Mon, 14 Sep 2026 05:31:18 GMT</pubDate></item><item><title><![CDATA[Reply to 雙 Radeon AI PRO R9700 跑 Qwen3.8-Flash-Next 調校成果 on Mon, 14 Sep 2026 05:26:17 GMT]]></title><description><![CDATA[<p dir="auto">这个是P2P   TP= 2 ？？？</p>
]]></description><link>https://lcz.me/post/18027</link><guid isPermaLink="true">https://lcz.me/post/18027</guid><dc:creator><![CDATA[张光璞]]></dc:creator><pubDate>Mon, 14 Sep 2026 05:26:17 GMT</pubDate></item></channel></rss>