<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[R9700 + Qwen3.8-27B：128K、MTP、Q4/Q6 都折腾了一遍]]></title><description><![CDATA[<p dir="auto">最近常在油管刷到老特每日一发的视频（辛苦啦 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f604.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--smile" style="height:23px;width:auto;vertical-align:middle" title="😄" alt="😄" />），看着看着一路顺藤摸瓜找到了这个论坛，又看到站内这篇 R9700 + Qwen3.8-27B 的实测。看完原帖和评论区各种折腾之后手痒了，刚好自己也有一张 R9700，那就跟着折腾一遍。</p>
<p dir="auto">先感谢原帖作者的实测和分享，我这次不少测试思路也是受到这篇帖子的启发：<br />
<a href="https://lcz.me/topic/1218/r9700-32gb-%E8%B7%91-qwen3.8-27b-llama.cpp-vulkan-mtp-%E5%AE%9E%E6%B5%8B%E4%B8%8E%E8%B8%A9%E5%9D%91%E8%AE%B0%E5%BD%95/6?_=1787701951422">[R9700 32GB 跑 Qwen3.8-27B，llama.cpp Vulkan + MTP 实测与踩坑记录]</a></p>
<p dir="auto">第一次在论坛发帖，本来只想简单记录一下，结果折腾着折腾着就啰里八嗦写了这么长一篇 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f602.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--joy" style="height:23px;width:auto;vertical-align:middle" title="😂" alt="😂" />。<br />
排版和阅读体验如果有什么不太舒服的地方，还请大家多多包涵。</p>
<p dir="auto">我的平台和原帖有一个比较大的区别：我不是标准桌面 PCIe x16 平台，而是 <strong>铭凡 (Minisforum) AI X1 Pro + OCuLink + DEG1 + R9700</strong>，操作系统也是 <strong>Windows 11 + AMD 官方驱动 + llama.cpp Vulkan</strong>。</p>
<p dir="auto">所以这次主要想看几个问题：</p>
<ul>
<li>R9700 通过 OCuLink 跑 Qwen3.8-27B 的实际表现如何？</li>
<li><code>UD-Q4_K_XL</code> 和 <code>UD-Q6_K_XL</code> 在 R9700 上速度差多少？</li>
<li>Q8 KV + 128K Context 是否实用？</li>
<li>Qwen3.8 的 MTP 在 Windows / Vulkan 下能带来多少实际收益？</li>
<li><code>spec-draft-n-max</code> 应该设多少？</li>
</ul>
<p dir="auto">先说结论：</p>
<blockquote>
<p dir="auto"><strong>在我这套环境和测试 workload 下，UD-Q4_K_XL + Q8 KV + 128K Context + MTP 的甜点是 <code>n-max=3</code>。</strong></p>
<p dir="auto">长代码生成实测约 <strong>43.2–43.4 t/s</strong>，而且复测结果比较稳定。</p>
</blockquote>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th style="text-align:right">实测</th>
</tr>
</thead>
<tbody>
<tr>
<td>Q4 Raw Decode（tg128）</td>
<td style="text-align:right"><strong>29.46 t/s</strong></td>
</tr>
<tr>
<td>Q6 Raw Decode（tg128）</td>
<td style="text-align:right"><strong>22.72 t/s</strong></td>
</tr>
<tr>
<td>Q4 MTP 最佳长输出</td>
<td style="text-align:right"><strong>43.44 t/s</strong></td>
</tr>
<tr>
<td>Q4 MTP 复测长输出</td>
<td style="text-align:right"><strong>43.20 t/s</strong></td>
</tr>
<tr>
<td>MTP 最佳参数</td>
<td style="text-align:right"><strong>n-max=3</strong></td>
</tr>
<tr>
<td>Prompt Prefill，8443 tokens</td>
<td style="text-align:right"><strong>584.39 t/s</strong></td>
</tr>
<tr>
<td>Context</td>
<td style="text-align:right"><strong>128K</strong></td>
</tr>
<tr>
<td>KV Cache</td>
<td style="text-align:right"><strong>Q8_0 / Q8_0</strong></td>
</tr>
</tbody>
</table>
<hr />
<h2>一、测试平台</h2>
<p dir="auto"><img src="https://upload.lcz.me/uploads/9652f16a-a5ec-42de-b238-b3aac18b083f.jpg" alt="llm-test.jpg" class=" img-fluid img-markdown" /></p>
<blockquote>
<p dir="auto">Minisforum AI X1 Pro + Radeon AI PRO R9700 32GB，通过 OCuLink + Minisforum DEG1 外接，独立 ATX 电源供电。</p>
</blockquote>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>配置</th>
</tr>
</thead>
<tbody>
<tr>
<td>主机</td>
<td><strong>Minisforum AI X1 Pro</strong></td>
</tr>
<tr>
<td>CPU</td>
<td><strong>AMD Ryzen AI 9 HX 470</strong></td>
</tr>
<tr>
<td>系统内存</td>
<td><strong>96 GB</strong></td>
</tr>
<tr>
<td>iGPU</td>
<td><strong>AMD Radeon 890M</strong></td>
</tr>
<tr>
<td>UMA Frame Buffer</td>
<td><strong>8 GB</strong></td>
</tr>
<tr>
<td>GPU</td>
<td><strong>ASRock Radeon AI PRO R9700 Creator</strong></td>
</tr>
<tr>
<td>VRAM</td>
<td><strong>32 GB</strong></td>
</tr>
<tr>
<td>GPU 连接</td>
<td><strong>OCuLink</strong></td>
</tr>
<tr>
<td>eGPU Dock</td>
<td><strong>Minisforum DEG1</strong></td>
</tr>
<tr>
<td>PSU</td>
<td><strong>be quiet! Power Zone 2 850W</strong></td>
</tr>
<tr>
<td>OS</td>
<td><strong>Windows 11 Pro</strong></td>
</tr>
</tbody>
</table>
<p dir="auto">这次 LLM 推理明确指定：</p>
<pre><code class="language-text">Vulkan0 = AMD Radeon AI PRO R9700
</code></pre>
<p dir="auto">890M 没有参与模型 GPU offload。</p>
<p dir="auto"><code>llama-bench --list-devices</code>：</p>
<pre><code class="language-text">Vulkan0: AMD Radeon AI PRO R9700
         32624 MiB, 31704 MiB free

Vulkan1: AMD Radeon(TM) 890M Graphics
         79994 MiB, 75994 MiB free
</code></pre>
<p dir="auto">R9700 Vulkan backend 同时识别到：</p>
<pre><code class="language-text">fp16: 1
bf16: 1
int dot: 1
matrix cores: KHR_coopmat
</code></pre>
<hr />
<h2>二、软件与推理环境</h2>
<p dir="auto">这次除了跑 <code>llama-bench</code> 做标准化性能测试之外，我的实际使用环境是 <strong>DeepSeek Harness（DSH）+ llama.cpp</strong>。</p>
<p dir="auto">DSH 通过本机 OpenAI-compatible API 连接 <code>llama-server</code>，整体链路大致如下：</p>
<pre><code class="language-text">DSH
 ↓
OpenAI-compatible API
 ↓
llama-server
 ↓
Vulkan0
 ↓
R9700
</code></pre>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>配置</th>
</tr>
</thead>
<tbody>
<tr>
<td>Agent / 前端</td>
<td><strong>DeepSeek Harness（DSH）</strong></td>
</tr>
<tr>
<td>API</td>
<td><strong>llama.cpp OpenAI-compatible API</strong></td>
</tr>
<tr>
<td>API 地址</td>
<td><strong>127.0.0.1:8080/v1</strong></td>
</tr>
<tr>
<td>llama.cpp</td>
<td><strong>build: 758443071 (10612)</strong></td>
</tr>
<tr>
<td>Backend</td>
<td><strong>Vulkan</strong></td>
</tr>
<tr>
<td>GPU</td>
<td><strong>Vulkan0 / R9700</strong></td>
</tr>
<tr>
<td>GPU Offload</td>
<td><strong>全部</strong></td>
</tr>
<tr>
<td>Parallel</td>
<td><strong>1</strong></td>
</tr>
<tr>
<td>Flash Attention</td>
<td><strong>ON</strong></td>
</tr>
<tr>
<td>KV Cache</td>
<td><strong>Q8_0 / Q8_0</strong></td>
</tr>
<tr>
<td>Context</td>
<td><strong>131072（128K）</strong></td>
</tr>
<tr>
<td>Reasoning</td>
<td><strong>ON</strong></td>
</tr>
<tr>
<td>Speculative Decoding</td>
<td><strong>MTP</strong></td>
</tr>
<tr>
<td>MTP Device</td>
<td><strong>Vulkan0</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><code>llama-server</code> 启动后，DSH 通过下面的模型 alias 连接：</p>
<pre><code class="language-text">qwen3.8-27b@q4_k_xl
</code></pre>
<p dir="auto">主要测试模型：</p>
<pre><code class="language-text">Qwen3.8-27B-UD-Q4_K_XL.gguf
Qwen3.8-27B-UD-Q6_K_XL.gguf
</code></pre>
<p dir="auto">这篇里其实有两种不同性质的测试，后面会分开写：</p>
<ol>
<li>
<p dir="auto"><strong><code>llama-bench</code></strong><br />
用来测 Q4 / Q6 的标准化 Prompt Processing（pp）和 Raw Token Generation（tg），方便做比较。</p>
</li>
<li>
<p dir="auto"><strong><code>llama-server + MTP</code> 实际长输出测试</strong><br />
用实际 Coding / 长文本 workload 测试 MTP，包括不同 <code>spec-draft-n-max</code> 设置；部分测试由 DSH 作为前端连接本机 <code>llama-server</code>。</p>
</li>
</ol>
<p dir="auto">所有最终的 Prompt、Generation、Draft Acceptance、Mean Len 等性能数据，都是直接取自 <strong><code>llama-server</code> 的 timing log</strong>，不是 DSH 显示的估算速度。</p>
<p dir="auto">所以后面看到的 <strong>29.46 t/s</strong> 和 <strong>43.2～43.4 t/s</strong> 代表的是两种不同测试：</p>
<pre><code class="language-text">29.46 t/s    = llama-bench Q4 tg128（Raw Decode）

43.2~43.4 t/s
             = llama-server + MTP 实际长输出
</code></pre>
<p dir="auto">两者可以用来观察 Raw Decode 与实际 MTP workload 的差异，但不是完全相同的 benchmark，不能直接当成严格的 apples-to-apples MTP 加速率。</p>
<h2>三、先跑 llama-bench：Q4 vs Q6</h2>
<p dir="auto">为了把模型本身的基础性能和 MTP 后的实际 server generation 分开看，我先用 <code>llama-bench</code> 跑 Q4 和 Q6。</p>
<p dir="auto">两者使用相同参数：</p>
<pre><code class="language-text">-dev Vulkan0
-ngl 999
-fa on
-ctk q8_0
-ctv q8_0
-p 512,2048,8192
-n 128
-r 3
</code></pre>
<h3>Benchmark 结果</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>测试</th>
<th style="text-align:right">UD-Q4_K_XL</th>
<th style="text-align:right">UD-Q6_K_XL</th>
</tr>
</thead>
<tbody>
<tr>
<td>模型大小</td>
<td style="text-align:right"><strong>16.34 GiB</strong></td>
<td style="text-align:right">23.55 GiB</td>
</tr>
<tr>
<td>pp512</td>
<td style="text-align:right"><strong>680.55 ± 0.46 t/s</strong></td>
<td style="text-align:right">655.10 ± 0.53 t/s</td>
</tr>
<tr>
<td>pp2048</td>
<td style="text-align:right"><strong>662.92 ± 8.77 t/s</strong></td>
<td style="text-align:right">646.17 ± 1.21 t/s</td>
</tr>
<tr>
<td>pp8192</td>
<td style="text-align:right"><strong>629.81 ± 2.33 t/s</strong></td>
<td style="text-align:right">607.60 ± 3.55 t/s</td>
</tr>
<tr>
<td>tg128</td>
<td style="text-align:right"><strong>29.46 ± 0.06 t/s</strong></td>
<td style="text-align:right">22.72 ± 0.04 t/s</td>
</tr>
</tbody>
</table>
<p dir="auto">这里最明显的是 decode。</p>
<p dir="auto">Q4：</p>
<blockquote>
<p dir="auto"><strong>29.46 t/s</strong></p>
</blockquote>
<p dir="auto">Q6：</p>
<blockquote>
<p dir="auto"><strong>22.72 t/s</strong></p>
</blockquote>
<p dir="auto">以 Q6 为基准，Q4 的 <code>tg128</code> 大约快 <strong>29.7%</strong>。</p>
<p dir="auto">Prefill 的差距反而只有几个百分点。</p>
<p dir="auto">这也让我最后比较倾向把 <strong>UD-Q4_K_XL 当成 R9700 32GB 的 daily driver</strong>：不仅 decode 更快，而且模型少占约 7 GiB VRAM，对大 Context / KV cache 也更友好。</p>
<hr />
<h2>四、接下来测试 MTP</h2>
<p dir="auto"><code>llama-bench tg128</code> 是很好的标准化 raw decode reference，但我的实际用途不是只生成 128 tokens。</p>
<p dir="auto">我主要用本地模型做 coding / agent，所以接下来使用实际的长代码生成 workload 测 <code>llama-server + MTP</code>。</p>
<p dir="auto">最终 Q4 server 配置：</p>
<pre><code class="language-powershell">&amp; "D:\AI\Apps\llama.cpp\b10612-vulkan\llama-server.exe" `
  --model "D:\AI\Models\LLM\unsloth\Qwen3.8-27B-GGUF\Qwen3.8-27B-UD-Q4_K_XL.gguf" `
  --mmproj "D:\AI\Models\LLM\unsloth\Qwen3.8-27B-GGUF\mmproj-F16.gguf" `
  --device Vulkan0 `
  --gpu-layers all `
  --ctx-size 131072 `
  --parallel 1 `
  --flash-attn on `
  --cache-type-k q8_0 `
  --cache-type-v q8_0 `
  --spec-type draft-mtp `
  --spec-draft-device Vulkan0 `
  --spec-draft-ngl all `
  --spec-draft-n-max 3 `
  --spec-draft-n-min 0 `
  --spec-draft-p-min 0.75 `
  --reasoning on `
  --host 127.0.0.1 `
  --port 8080 `
  --alias "qwen3.8-27b@q4_k_xl"
</code></pre>
<hr />
<h2>五、MTP <code>n-max</code> 调优</h2>
<p dir="auto">这里是这轮测试里我觉得最有意思的部分。</p>
<p dir="auto">同一类长代码生成 workload 下，我测试了不同的 <code>spec-draft-n-max</code>。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>MTP 设置</th>
<th style="text-align:right">Generation</th>
<th style="text-align:right">Acceptance</th>
<th style="text-align:right">Mean Len</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>n-max=2</code></td>
<td style="text-align:right">39.50 t/s</td>
<td style="text-align:right"><strong>90.51%</strong></td>
<td style="text-align:right">2.62</td>
</tr>
<tr>
<td><code>n-max=3</code></td>
<td style="text-align:right"><strong>43.44 t/s</strong></td>
<td style="text-align:right">88.59%</td>
<td style="text-align:right">3.13</td>
</tr>
<tr>
<td><code>n-max=4</code></td>
<td style="text-align:right">43.19 t/s</td>
<td style="text-align:right">86.58%</td>
<td style="text-align:right">3.49</td>
</tr>
<tr>
<td><code>n-max=5</code></td>
<td style="text-align:right">33.40 t/s</td>
<td style="text-align:right">82.80%</td>
<td style="text-align:right"><strong>3.70</strong></td>
</tr>
</tbody>
</table>
<p dir="auto">结果很明显：</p>
<pre><code class="language-text">2 → 3：明显变快
3 → 4：基本持平，3 略快
4 → 5：性能明显下降
</code></pre>
<p dir="auto"><img src="https://upload.lcz.me/uploads/2708187d-2f2a-48f6-bf1e-f5922b29a60e.png" alt="dsh.png" class=" img-fluid img-markdown" /><br />
所以在我这里：</p>
<blockquote>
<p dir="auto"><strong><code>n-max=3</code> 是 sweet spot。</strong></p>
</blockquote>
<p dir="auto"><code>n-max=5</code> 虽然平均 draft length 更长，但 acceptance 已经掉到 82.8%，额外 speculative work 的成本反而把实际 throughput 拉低到：</p>
<blockquote>
<p dir="auto"><strong>33.40 t/s</strong></p>
</blockquote>
<p dir="auto">所以 MTP 的 <code>n-max</code> 看来并不是越高越好。</p>
<hr />
<h2>六、<code>n-max=3</code> 再跑一次确认</h2>
<p dir="auto">第一次比较好的结果：</p>
<pre><code class="language-text">Generation     43.44 t/s
Acceptance     88.59%
Mean Len        3.13
</code></pre>
<p dir="auto">后来重新启动 server，用相同的核心配置又跑了一次：</p>
<pre><code class="language-text">prompt eval time = 14447.55 ms / 8443 tokens
                 = 584.39 tokens per second

eval time        = 238937.76 ms / 10324 tokens
                 = 43.20 tokens per second

total time       = 253385.31 ms / 18767 tokens

draft acceptance = 0.87117
                 = 6147 accepted / 7056 generated

mean len         = 3.11
</code></pre>
<p dir="auto">整理一下：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th style="text-align:right">实测</th>
</tr>
</thead>
<tbody>
<tr>
<td>Prompt Prefill，8443 tokens</td>
<td style="text-align:right"><strong>584.39 t/s</strong></td>
</tr>
<tr>
<td>长输出 Generation，10324 tokens</td>
<td style="text-align:right"><strong>43.20 t/s</strong></td>
</tr>
<tr>
<td>MTP Draft Acceptance</td>
<td style="text-align:right"><strong>87.12%</strong></td>
</tr>
<tr>
<td>MTP Mean Len</td>
<td style="text-align:right"><strong>3.11</strong></td>
</tr>
<tr>
<td>总处理 tokens</td>
<td style="text-align:right"><strong>18767</strong></td>
</tr>
<tr>
<td>总耗时</td>
<td style="text-align:right"><strong>约 253.4 秒</strong></td>
</tr>
<tr>
<td>Context</td>
<td style="text-align:right"><strong>128K</strong></td>
</tr>
<tr>
<td>KV Cache</td>
<td style="text-align:right"><strong>Q8_0 / Q8_0</strong></td>
</tr>
<tr>
<td>MTP</td>
<td style="text-align:right"><strong>n-max 3 / p-min 0.75</strong></td>
</tr>
</tbody>
</table>
<p dir="auto">两次 <code>n-max=3</code>：</p>
<pre><code class="language-text">43.44 t/s
43.20 t/s
</code></pre>
<p dir="auto"><img src="https://upload.lcz.me/uploads/da703597-6a97-42e0-9c26-78257000d503.png" alt="dsh43.png" class=" img-fluid img-markdown" /><br />
平均：</p>
<blockquote>
<p dir="auto"><strong>43.32 t/s</strong></p>
</blockquote>
<p dir="auto">两次只差约 <strong>0.6%</strong>，所以我觉得把这套配置的实际长输出性能记成：</p>
<blockquote>
<p dir="auto"><strong>约 43.3 t/s</strong></p>
</blockquote>
<p dir="auto">是比较合理的。</p>
<p dir="auto">而且第二次实际生成了 <strong>10,324 tokens</strong>，不是很短的 benchmark。</p>
<hr />
<h2>七、Raw decode 29.46 t/s，MTP workload 约 43.3 t/s</h2>
<p dir="auto">把几个数字放在一起看：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>测试</th>
<th style="text-align:right">Generation</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>llama-bench tg128</code></td>
<td style="text-align:right"><strong>29.46 t/s</strong></td>
</tr>
<tr>
<td>MTP 长代码生成 #1</td>
<td style="text-align:right"><strong>43.44 t/s</strong></td>
</tr>
<tr>
<td>MTP 长代码生成 #2</td>
<td style="text-align:right"><strong>43.20 t/s</strong></td>
</tr>
<tr>
<td>MTP 两次平均</td>
<td style="text-align:right"><strong>43.32 t/s</strong></td>
</tr>
</tbody>
</table>
<p dir="auto">从数值上看：</p>
<pre><code class="language-text">29.46 → 43.32 t/s
</code></pre>
<p dir="auto">约高 <strong>47%</strong>。</p>
<p dir="auto">不过这里必须说明：</p>
<blockquote>
<p dir="auto"><strong>这不是严格 apples-to-apples 的 +47% MTP benchmark。</strong></p>
</blockquote>
<p dir="auto"><code>llama-bench tg128</code> 和实际长代码生成是不同 workload，所以我把 <strong>29.46 t/s 当成 standardized raw decode reference</strong>，而不是宣称“开启 MTP 必定提升 47%”。</p>
<p dir="auto">真正比较有意义的是：在我的实际 coding workload 中，MTP 调好后可以长期维持在 <strong>43 t/s 左右</strong>。</p>
<hr />
<h2>八、还试了 <code>-b 16384 -ub 2048</code></h2>
<p dir="auto">看到一些 llama.cpp 调优讨论后，我也测试了：</p>
<pre><code class="language-text">-b 16384
-ub 2048
</code></pre>
<p dir="auto">结果：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>配置</th>
<th style="text-align:right">Generation</th>
<th style="text-align:right">Prompt</th>
<th style="text-align:right">Acceptance</th>
</tr>
</thead>
<tbody>
<tr>
<td>默认 batch / MTP3</td>
<td style="text-align:right"><strong>43.44 t/s</strong></td>
<td style="text-align:right">583.45 t/s</td>
<td style="text-align:right">88.59%</td>
</tr>
<tr>
<td><code>-b 16384 -ub 2048</code> / MTP3</td>
<td style="text-align:right">42.84 t/s</td>
<td style="text-align:right">583.03 t/s</td>
<td style="text-align:right">87.33%</td>
</tr>
</tbody>
</table>
<p dir="auto">在我的 workload 上没有提升，反而慢了一点。</p>
<p dir="auto">所以最后没有保留这两个参数。</p>
<hr />
<h2>九、为什么最后选择 Q4 而不是 Q6？</h2>
<p dir="auto">目前对我来说：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th style="text-align:right">UD-Q4_K_XL</th>
<th style="text-align:right">UD-Q6_K_XL</th>
</tr>
</thead>
<tbody>
<tr>
<td>模型大小</td>
<td style="text-align:right"><strong>16.34 GiB</strong></td>
<td style="text-align:right">23.55 GiB</td>
</tr>
<tr>
<td>tg128</td>
<td style="text-align:right"><strong>29.46 t/s</strong></td>
<td style="text-align:right">22.72 t/s</td>
</tr>
<tr>
<td>相对 Q4 decode</td>
<td style="text-align:right"><strong>100%</strong></td>
<td style="text-align:right">77.1%</td>
</tr>
<tr>
<td>VRAM 余量</td>
<td style="text-align:right"><strong>较多</strong></td>
<td style="text-align:right">较少</td>
</tr>
<tr>
<td>128K Context</td>
<td style="text-align:right"><strong>更宽松</strong></td>
<td style="text-align:right">更紧</td>
</tr>
<tr>
<td>日常 Coding / Agent</td>
<td style="text-align:right"><strong>我的选择</strong></td>
<td style="text-align:right">偏质量时考虑</td>
</tr>
</tbody>
</table>
<p dir="auto">Q6 的量化质量理论上当然更好。</p>
<p dir="auto">但 R9700 是 32GB VRAM。</p>
<p dir="auto">我的实际目标又包括：</p>
<ul>
<li>128K Context</li>
<li>Q8 KV</li>
<li>MTP</li>
<li>mmproj</li>
<li>coding agent</li>
<li>长 session</li>
</ul>
<p dir="auto">所以多出来的 VRAM headroom 对我来说很有价值。</p>
<p dir="auto">目前我的 daily driver 会选择：</p>
<blockquote>
<p dir="auto"><strong>UD-Q4_K_XL</strong></p>
</blockquote>
<hr />
<h2>十、128K Context + Q8 KV</h2>
<p dir="auto">这次最终配置使用：</p>
<pre><code class="language-text">--ctx-size 131072
--parallel 1
--cache-type-k q8_0
--cache-type-v q8_0
</code></pre>
<p dir="auto">我没有为了 benchmark 把 context 降到很小。</p>
<p dir="auto">原因很简单：我的目标是实际接 coding agent 使用，而不是只追求一个最高 t/s 数字。</p>
<p dir="auto">Q4 模型约 16.34 GiB，因此在 R9700 32GB 上仍然可以给 KV cache、MTP 和其他 runtime allocation 留出不少空间。</p>
<p dir="auto">这也是我觉得 Q4 在 32GB 卡上特别合适的原因之一。</p>
<hr />
<h2>十一、关于 OCuLink</h2>
<p dir="auto">我的 R9700 并不是插在桌面 PCIe x16 主板上，而是：</p>
<pre><code class="language-text">Minisforum AI X1 Pro
        │
      OCuLink
        │
Minisforum DEG1
        │
Radeon AI PRO R9700 32GB
</code></pre>
<p dir="auto">OCuLink 的 PCIe host bandwidth 显然不等于桌面 PCIe x16。</p>
<p dir="auto">不过对于模型已经完整驻留 VRAM 的 LLM inference，decode 阶段的大部分权重访问发生在 GPU 本地显存，因此 PCIe 链路带宽不一定会像某些需要频繁 CPU <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2194.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--left_right_arrow" style="height:23px;width:auto;vertical-align:middle" title="↔" alt="↔" /> GPU 数据交换的 workload 那样成为主要瓶颈。</p>
<p dir="auto">这次 OCuLink 环境下实测：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th style="text-align:right">实测</th>
</tr>
</thead>
<tbody>
<tr>
<td>Q4 pp512</td>
<td style="text-align:right"><strong>680.55 t/s</strong></td>
</tr>
<tr>
<td>Q4 pp2048</td>
<td style="text-align:right"><strong>662.92 t/s</strong></td>
</tr>
<tr>
<td>Q4 pp8192</td>
<td style="text-align:right"><strong>629.81 t/s</strong></td>
</tr>
<tr>
<td>Q4 tg128</td>
<td style="text-align:right"><strong>29.46 t/s</strong></td>
</tr>
<tr>
<td>Q4 MTP 长输出</td>
<td style="text-align:right"><strong>约 43.2～43.4 t/s</strong></td>
</tr>
</tbody>
</table>
<p dir="auto">从目前结果来看，这套 OCuLink 配置下的 llama.cpp Vulkan 推理表现正常，没有观察到明显异常的性能瓶颈。</p>
<p dir="auto">不过这里要特别说明：<strong>我没有同一张 R9700 在 PCIe x16 下的直接 A/B 测试数据。</strong></p>
<p dir="auto">所以这些结果只能证明这套 <strong>OCuLink + R9700</strong> 配置能够达到上述性能，不能据此得出“OCuLink 相比 PCIe x16 没有性能损失”的结论。</p>
<p dir="auto">如果以后有机会用同一张 R9700、同一个 llama.cpp build、同一个 GGUF 和完全相同参数分别测试 OCuLink 与 PCIe x16，再来量化 OCuLink 对 Prefill 和 Decode 的实际影响会比较有意义。</p>
<hr />
<h2>十二、BIOS / ReBAR</h2>
<p dir="auto">测试期间也检查了一下 BIOS。</p>
<p dir="auto">确认：</p>
<pre><code class="language-text">Resizable BAR Support = Enabled
UMA Frame Buffer = 8 GB
</code></pre>
<p dir="auto">Windows 正常识别 R9700：</p>
<pre><code class="language-text">AMD Radeon AI PRO R9700
PCI bus 200, device 0, function 0
</code></pre>
<p dir="auto">我这版 Minisforum BIOS 没有找到一个独立可见的 <strong>Above 4G Decoding</strong> 开关，因此这里不写成“已确认开启”。</p>
<p dir="auto">没有为了跑分去修改隐藏 BIOS 设置。</p>
<hr />
<h2>十三、和原帖结果的关系</h2>
<p dir="auto">这也是我觉得比较有意思的地方。</p>
<p dir="auto">原帖作者的环境、quant、KV 设置等和我并不完全一样，所以两边数字不能直接当作同一 benchmark 排名。</p>
<p dir="auto">但一个值得继续研究的差异是 <strong>MTP <code>n-max</code></strong>。</p>
<p dir="auto">我的环境：</p>
<pre><code class="language-text">Windows 11
AMD proprietary Vulkan driver
llama.cpp build 758443071 (10612)
UD-Q4_K_XL
Q8 KV
128K
R9700 / OCuLink
</code></pre>
<p dir="auto">在这个组合下：</p>
<blockquote>
<p dir="auto"><strong><code>n-max=3</code> 最合适。</strong></p>
</blockquote>
<p dir="auto">而原帖以及评论区的其他配置有不同的 sweet spot。</p>
<p dir="auto">这说明 MTP 最佳参数很可能跟以下因素都有关系：</p>
<ul>
<li>quant</li>
<li>KV cache type</li>
<li>backend</li>
<li>driver</li>
<li>llama.cpp build</li>
<li>workload</li>
<li>draft acceptance rate</li>
</ul>
<p dir="auto">所以我现在的感觉是：</p>
<blockquote>
<p dir="auto"><strong>不要直接照抄别人的 <code>n-max</code>，最好自己跑 2 / 3 / 4 / 5。</strong></p>
</blockquote>
<p dir="auto">R9700 跑一次这种 sweep 成本并不高，但最后可能差很多。</p>
<hr />
<h2>十四、目前的 Daily Driver 配置</h2>
<p dir="auto">经过这轮测试，我准备先把下面这套作为日常配置：</p>
<pre><code class="language-text">Model: Qwen3.8-27B UD-Q4_K_XL

GPU: R9700 / Vulkan0
GPU offload: all

Context: 131072
Parallel: 1

Flash Attention: ON

KV:
K = q8_0
V = q8_0

MTP:
n-max = 3
n-min = 0
p-min = 0.75

Reasoning: ON
</code></pre>
<p dir="auto">实际长代码生成：</p>
<blockquote>
<p dir="auto"><strong>≈ 43.3 t/s</strong></p>
</blockquote>
<p dir="auto">完整启动命令：</p>
<pre><code class="language-powershell">&amp; "D:\AI\Apps\llama.cpp\b10612-vulkan\llama-server.exe" `
  --model "D:\AI\Models\LLM\unsloth\Qwen3.8-27B-GGUF\Qwen3.8-27B-UD-Q4_K_XL.gguf" `
  --mmproj "D:\AI\Models\LLM\unsloth\Qwen3.8-27B-GGUF\mmproj-F16.gguf" `
  --device Vulkan0 `
  --gpu-layers all `
  --ctx-size 131072 `
  --parallel 1 `
  --flash-attn on `
  --cache-type-k q8_0 `
  --cache-type-v q8_0 `
  --spec-type draft-mtp `
  --spec-draft-device Vulkan0 `
  --spec-draft-ngl all `
  --spec-draft-n-max 3 `
  --spec-draft-n-min 0 `
  --spec-draft-p-min 0.75 `
  --reasoning on `
  --host 127.0.0.1 `
  --port 8080 `
  --alias "qwen3.8-27b@q4_k_xl"
</code></pre>
<hr />
<h2>十五、总结</h2>
<p dir="auto">这轮测试下来，我自己的几个结论：</p>
<h3>1. R9700 32GB 跑 Qwen3.8-27B Q4 很舒服</h3>
<p dir="auto">模型完整 GPU offload 后，还有足够空间留给大 Context 和 KV。</p>
<h3>2. Q4 和 Q6 的 decode 差距比 prefill 明显得多</h3>
<pre><code class="language-text">Q4 tg128 = 29.46 t/s
Q6 tg128 = 22.72 t/s
</code></pre>
<p dir="auto">对我的 coding / agent 用途，Q4 的综合平衡更好。</p>
<h3>3. MTP 很值得开，但需要调</h3>
<p dir="auto">我的实际长代码 workload：</p>
<pre><code class="language-text">n-max 2 → 39.50 t/s
n-max 3 → 43.44 t/s
n-max 4 → 43.19 t/s
n-max 5 → 33.40 t/s
</code></pre>
<p dir="auto"><code>n-max=3</code> 是明显的 sweet spot。</p>
<h3>4. <code>-b 16384 -ub 2048</code> 在我的环境没有帮助</h3>
<p dir="auto">所以没有保留。</p>
<h3>5. 128K + Q8 KV 是可以实际使用的</h3>
<p dir="auto">我更愿意牺牲一点理论最高 benchmark，换取真正适合 coding agent 的配置。</p>
<h3>6. OCuLink 环境下的推理表现正常</h3>
<p dir="auto">目前没有观察到明显异常的性能瓶颈，但因为缺少同卡 PCIe x16 的直接 A/B 测试，所以这里只作为 OCuLink 环境下的实测数据点，不对 OCuLink 相比 PCIe x16 的性能损失做定量结论。</p>
<hr />
<p dir="auto">最后感谢原帖作者和评论区参与测试的网友。</p>
<p dir="auto">这轮测试很大程度上是受到原帖的启发。我这里主要补充一个：</p>
<blockquote>
<p dir="auto"><strong>Windows + AMD 官方 Vulkan + OCuLink + R9700 + UD-Q4_K_XL / Q6_K_XL</strong></p>
</blockquote>
<p dir="auto">的数据点。</p>
<p dir="auto">如果其他 R9700 用户有 Linux/RADV、ROCm、PCIe x16，或者不同 llama.cpp build 下同一个 Q4_K_XL 的数据，也欢迎一起对比。</p>
<p dir="auto">尤其想看看大家的：</p>
<pre><code class="language-text">llama-bench tg128
MTP n-max 2 / 3 / 4 / 5
draft acceptance
</code></pre>
<p dir="auto">到底差多少。</p>
]]></description><link>https://lcz.me/topic/1345</link><generator>RSS for Node</generator><lastBuildDate>Mon, 07 Sep 2026 17:34:07 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1345.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 26 Aug 2026 22:23:14 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to R9700 + Qwen3.8-27B：128K、MTP、Q4/Q6 都折腾了一遍 on Fri, 28 Aug 2026 09:13:36 GMT]]></title><description><![CDATA[<p dir="auto">格式工整，数据翔实，图文并茂，置顶</p>
]]></description><link>https://lcz.me/post/14592</link><guid isPermaLink="true">https://lcz.me/post/14592</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 28 Aug 2026 09:13:36 GMT</pubDate></item><item><title><![CDATA[Reply to R9700 + Qwen3.8-27B：128K、MTP、Q4/Q6 都折腾了一遍 on Thu, 27 Aug 2026 05:39:21 GMT]]></title><description><![CDATA[<p dir="auto">Q4 的智力够不够 ，我现在已经不再用豆包。以前问他一个霍尔是线性的还是开关的。丫告诉我是开关霍尔，误导了我特别长时间 ，浪费了我大量时间 。这玩意方向错了越努力越离谱。</p>
]]></description><link>https://lcz.me/post/14269</link><guid isPermaLink="true">https://lcz.me/post/14269</guid><dc:creator><![CDATA[张光璞]]></dc:creator><pubDate>Thu, 27 Aug 2026 05:39:21 GMT</pubDate></item><item><title><![CDATA[Reply to R9700 + Qwen3.8-27B：128K、MTP、Q4/Q6 都折腾了一遍 on Thu, 27 Aug 2026 03:24:19 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/skyrocker" aria-label="Profile: skyrocker">@<bdi>skyrocker</bdi></a> 我也是 R9700，看完你的帖子忍不住跟着跑了一遍。我是 Linux 平台，跟你的 Windows 正好互补，把两边数据放一起挺有意思的。</p>
<h2>我的环境</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>配置</th>
</tr>
</thead>
<tbody>
<tr>
<td>主机</td>
<td>HP Z420 工作站</td>
</tr>
<tr>
<td>CPU</td>
<td>Intel Xeon E5-2667 v2（8C16T @ 3.3GHz）</td>
</tr>
<tr>
<td>内存</td>
<td>64 GB DDR3</td>
</tr>
<tr>
<td>GPU</td>
<td>AMD Radeon AI PRO R9700 32GB</td>
</tr>
<tr>
<td>GPU 连接</td>
<td>PCIe 3.0 x16（板载插槽）</td>
</tr>
<tr>
<td>OS</td>
<td>Ubuntu 24.04 + RADV（Mesa 26.1.5）</td>
</tr>
<tr>
<td>llama.cpp</td>
<td>build 10618（master 最新）</td>
</tr>
<tr>
<td>Backend</td>
<td>Vulkan</td>
</tr>
<tr>
<td>KV Cache</td>
<td>Q8_0 / Q8_0</td>
</tr>
<tr>
<td>MTP</td>
<td>draft-mtp，p-min 0.75，reasoning on</td>
</tr>
<tr>
<td>模型</td>
<td>同一个 unsloth UD-Q4_K_XL</td>
</tr>
</tbody>
</table>
<p dir="auto">模型、KV、MTP 参数跟你保持一致，唯一差别就是平台本身（Linux/RADV/PCIe x16 vs Windows/官方驱动/OCuLink）。</p>
<h2>llama-bench 对比（同款命令 -p 512,2048,8192 -n 128 -r 3）</h2>
<h3>UD-Q4_K_XL</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>测试</th>
<th style="text-align:right">你（Windows/OCuLink）</th>
<th style="text-align:right">我（Linux/PCIe x16）</th>
<th style="text-align:right">差距</th>
</tr>
</thead>
<tbody>
<tr>
<td>pp512</td>
<td style="text-align:right">680.55 ± 0.46</td>
<td style="text-align:right"><strong>909.60 ± 1.10</strong></td>
<td style="text-align:right">+33.7%</td>
</tr>
<tr>
<td>pp2048</td>
<td style="text-align:right">662.92 ± 8.77</td>
<td style="text-align:right"><strong>856.50 ± 0.52</strong></td>
<td style="text-align:right">+29.2%</td>
</tr>
<tr>
<td>pp8192</td>
<td style="text-align:right">629.81 ± 2.33</td>
<td style="text-align:right"><strong>840.52 ± 0.91</strong></td>
<td style="text-align:right">+33.5%</td>
</tr>
<tr>
<td>tg128</td>
<td style="text-align:right"><strong>29.46 ± 0.06</strong></td>
<td style="text-align:right">16.00 ± 0.03</td>
<td style="text-align:right">-45.7%</td>
</tr>
</tbody>
</table>
<h3>UD-Q6_K_XL</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>测试</th>
<th style="text-align:right">你（Windows/OCuLink）</th>
<th style="text-align:right">我（Linux/PCIe x16）</th>
<th style="text-align:right">差距</th>
</tr>
</thead>
<tbody>
<tr>
<td>pp512</td>
<td style="text-align:right">655.10 ± 0.53</td>
<td style="text-align:right"><strong>870.92 ± 1.12</strong></td>
<td style="text-align:right">+32.9%</td>
</tr>
<tr>
<td>pp2048</td>
<td style="text-align:right">646.17 ± 1.21</td>
<td style="text-align:right"><strong>827.91 ± 0.35</strong></td>
<td style="text-align:right">+28.1%</td>
</tr>
<tr>
<td>pp8192</td>
<td style="text-align:right">607.60 ± 3.55</td>
<td style="text-align:right"><strong>811.36 ± 0.45</strong></td>
<td style="text-align:right">+33.5%</td>
</tr>
<tr>
<td>tg128</td>
<td style="text-align:right"><strong>22.72 ± 0.04</strong></td>
<td style="text-align:right">13.36 ± 0.02</td>
<td style="text-align:right">-41.2%</td>
</tr>
</tbody>
</table>
<p dir="auto">Q4 和 Q6 的 pattern 完全一致：prefill 我这边快 30% 左右，tg 你那边快 40%+。Q6/Q4 的相对比例两边也差不多（你 77.1%，我 83.5%），说明这个差距跟量化关系不大，是平台层面的系统性差异。</p>
<h2>MTP n-max sweep（同款长代码生成 workload）</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>n-max</th>
<th style="text-align:right">你（Windows）</th>
<th style="text-align:right">我（Linux）</th>
<th style="text-align:right">差距</th>
</tr>
</thead>
<tbody>
<tr>
<td>2</td>
<td style="text-align:right">39.50</td>
<td style="text-align:right">33.5</td>
<td style="text-align:right">-15%</td>
</tr>
<tr>
<td>3</td>
<td style="text-align:right"><strong>43.44</strong></td>
<td style="text-align:right"><strong>34.5</strong></td>
<td style="text-align:right">-21%</td>
</tr>
<tr>
<td>4</td>
<td style="text-align:right">43.19</td>
<td style="text-align:right">34.5</td>
<td style="text-align:right">-20%</td>
</tr>
<tr>
<td>5</td>
<td style="text-align:right">33.40</td>
<td style="text-align:right">32.6</td>
<td style="text-align:right">-2%</td>
</tr>
</tbody>
</table>
<p dir="auto">sweet spot 一致，都是 n=3/4。我这边 n=3 的 draft acceptance 85.6%（你 88.6%），mean len 2.86（你 3.13），差距不大。</p>
<h2>几个推测（注意：不是结论，缺 A/B 数据）</h2>
<ol>
<li>
<p dir="auto"><strong>prefill 差距可能跟 OCuLink 有关，但没法证实</strong>：OCuLink 是 PCIe 4.0 x4（~8GB/s），我这边是 PCIe 3.0 x16（~16GB/s），prefill 吃 CPU<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2194.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--left_right_arrow" style="height:23px;width:auto;vertical-align:middle" title="↔" alt="↔" />GPU 传输，理论上 x16 宽就是快。但你帖子里自己也说了没有同卡 PCIe x16 的 A/B，我这边也没有 OCuLink 环境，所以这只是基于带宽的推测，不是实测结论。</p>
</li>
<li>
<p dir="auto"><strong>tg 差距可能跟驱动有关，同样是推测</strong>：decode 是 GPU 内部运算，跟 PCIe 无关，同一个 R9700 同一个模型，唯一平台变量是 RADV vs AMD 官方驱动。GitHub #26663 里 9070 XT 的讨论也提到 RADV 的 MUL_MAT_VEC 优化不如官方驱动。但严格说我没有在同一张卡上做过 RADV vs 官方驱动的 A/B，不能下死结论。</p>
</li>
<li>
<p dir="auto"><strong>MTP 确实把差距收窄了（这个是数据事实）</strong>：raw tg128 差 46%（29.46 vs 16.00），MTP 长生成差 15-21%（n=2~4），n=5 甚至几乎打平（-2%）。这个是从两边实测数字直接算出来的，不涉及推测。</p>
</li>
<li>
<p dir="auto"><strong>UD-Q4_K_XL 至少在我这几次测试里没遇到死循环</strong>：跑了 4 个 n-max × 2000 tokens 长代码生成，一次都没碰到 Qwen3.8 思考死循环。样本不大，不敢说"确实稳"，只能说"没遇到"。跟清风明月那个 67 t/s 长时间零崩的数据比，我这只能算个小的数据点。</p>
</li>
</ol>
<h2>想请教</h2>
<ol>
<li>你的 draft acceptance 是拿哪个 workload 统计的？我这边是长代码生成，可能跟你的 DSH 场景有点出入</li>
<li>OCuLink 下你试过 pp8192 以上的 prefill 吗？我想确认 x4 链路在高 pp 是不是瓶颈更明显</li>
<li><strong>tg 差距想听听你的判断</strong>：同样的 R9700 + 同样模型 + 同样 llama.cpp 系，tg 差 40%+ 这个幅度挺大的。我目前怀疑 RADV 的 MUL_MAT_VEC 路径不如 AMD 官方驱动（GitHub #26663 里 9070 XT 也有类似的讨论），但没做过同卡 A/B，不敢下结论。你 Windows 下有没有试过 RADV（WSL/msys 之类）？或者有没有见过其他 Linux R9700 用户的 tg 数据？很想确认这是驱动问题还是 llama.cpp Vulkan 后端在 Linux 下的普遍现象</li>
</ol>
<hr />
<p dir="auto">数据都记在本地 DB 了，有需要可以随时展开。</p>
]]></description><link>https://lcz.me/post/14255</link><guid isPermaLink="true">https://lcz.me/post/14255</guid><dc:creator><![CDATA[alan.lgv60]]></dc:creator><pubDate>Thu, 27 Aug 2026 03:24:19 GMT</pubDate></item><item><title><![CDATA[Reply to R9700 + Qwen3.8-27B：128K、MTP、Q4/Q6 都折腾了一遍 on Thu, 27 Aug 2026 02:17:43 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/skyrocker" aria-label="Profile: skyrocker">@<bdi>skyrocker</bdi></a> 不好意思，手滑发错了<br />
请把 Qwen3.8-27B 的完整测试 prompt、llama-server 参数和 llama-bench command 发给我。<br />
让我在我的环境测试一下</p>
]]></description><link>https://lcz.me/post/14231</link><guid isPermaLink="true">https://lcz.me/post/14231</guid><dc:creator><![CDATA[alan.lgv60]]></dc:creator><pubDate>Thu, 27 Aug 2026 02:17:43 GMT</pubDate></item><item><title><![CDATA[Reply to R9700 + Qwen3.8-27B：128K、MTP、Q4/Q6 都折腾了一遍 on Thu, 27 Aug 2026 00:49:41 GMT]]></title><description><![CDATA[<p dir="auto">还是严重推荐q6，这个是真正的甜点，显存留那么多也没用。</p>
]]></description><link>https://lcz.me/post/14217</link><guid isPermaLink="true">https://lcz.me/post/14217</guid><dc:creator><![CDATA[AGI]]></dc:creator><pubDate>Thu, 27 Aug 2026 00:49:41 GMT</pubDate></item><item><title><![CDATA[Reply to R9700 + Qwen3.8-27B：128K、MTP、Q4/Q6 都折腾了一遍 on Thu, 27 Aug 2026 00:08:42 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/alan.lgv60" aria-label="Profile: alan.lgv60">@<bdi>alan.lgv60</bdi></a> 哈哈，你这个应该是回错帖子了 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f604.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--smile" style="height:23px;width:auto;vertical-align:middle" title="😄" alt="😄" /></p>
<p dir="auto">我这篇测试的是 Qwen3.8-27B + llama.cpp Vulkan / MTP，主要是 LLM 的 Coding 长输出和 llama-bench</p>
<p dir="auto">你提到的「拿起铅笔书写 / 纸揉成团丢掉」、8 steps、CFG、sigma shift、175 帧、audioMode、seed 这些看起来都是视频生成相关的参数<br />
我这边这次测试完全没有用到。</p>
<p dir="auto">我这篇如果你想复现的话，我倒是可以把 Qwen3.8-27B 的完整测试 prompt、llama-server 参数和 llama-bench command 发给你</p>
]]></description><link>https://lcz.me/post/14210</link><guid isPermaLink="true">https://lcz.me/post/14210</guid><dc:creator><![CDATA[skyrocker]]></dc:creator><pubDate>Thu, 27 Aug 2026 00:08:42 GMT</pubDate></item></channel></rss>