<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G，本地部署 Qwen3.8-27B 怎么吐字最快？]]></title><description><![CDATA[<p dir="auto">配置如下：</p>
<ul>
<li>主板：微星 B660M 迫击炮（第一根 x16 槽是 CPU 的 PCIe 4.0 x16，第二根 x16 长槽是芯片组的 PCIe 3.0 x4）</li>
<li>显卡：RTX 4060Ti 16G 单卡</li>
<li>CPU：i7-12700K</li>
<li>内存：DDR4 32G</li>
</ul>
<p dir="auto">想本地跑 Qwen3.8-27B（稠密 27B），请教几个问题：</p>
<ol>
<li>
<p dir="auto">单卡 16G 显然装不下 27B 权重（Q4_K_XL 就 16G+）+ KV 缓存，必然 offload 到内存。这种单机单卡 + 32G DDR4 的 offload 方案，实测吐字大概能到多少 t/s？是不是基本就是“慢成乌龟”？</p>
</li>
<li>
<p dir="auto">想提速最划算的路子是不是再淘一张 16G 卡（第二张 4060Ti 16G 或 5060Ti 16G）组双卡 32G，走 llama.cpp 张量并行（-sm tensor -ts 1,1）+ MTP？B660M 副槽只有 PCIe 3.0 x4，跑 27B 推理带宽够不够用？</p>
</li>
<li>
<p dir="auto">量化与参数：27B 上 Q4_K_M / UD-Q4_K_XL 就行了吧？MTP 的 spec-draft-n-max 在双 16G 卡上设 1 还是 2 比较稳（怕 OOM）？上下文开到 64K 是不是就只能放弃 KV 量化？</p>
</li>
</ol>
<p dir="auto">论坛里看过双 3060(24G) 跑 43~50 t/s、双 5060Ti 65 t/s 的帖子，想确认下我这板子加第二张卡是否是最优解，还是单卡硬扛有更聪明的办法。谢谢各位大佬！</p>
]]></description><link>https://lcz.me/topic/1354</link><generator>RSS for Node</generator><lastBuildDate>Thu, 10 Sep 2026 00:44:08 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1354.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 27 Aug 2026 06:51:02 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G，本地部署 Qwen3.8-27B 怎么吐字最快？ on Sat, 29 Aug 2026 03:19:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/asd2667" aria-label="Profile: asd2667">@<bdi>asd2667</bdi></a> 试了一下，现在是不能开mtp的……速度太慢了……Found it. The truth:</p>
<ul>
<li>Warm TTFT is actually fine: your unique 37-token prompt just processed at 29.6 tok/s (1.25s).</li>
<li>The 2.5 tok/s you hit was the first request after the restart — one-time CUDA-graph + MTP draft compilation. That's a cold-start cost, not steady-state.<br />
But I also found MTP is hurting, not helping on this 16GB setup:</li>
<li>Decode: 12.8 tok/s with MTP vs 18.9 tok/s without it earlier (the draft compute overhead exceeds the speculative gain at these speeds; draft acceptance mean was only ~2.3).</li>
<li>MTP adds the cold-start compile overhead you felt.<br />
So per the official page, MTP was meant for the bigger M-model on larger VRAM. On your XS-Pro/16GB it's net-negative. I'll disable MTP, keep the rest of the recommended config (q8/q4 KV, flash-attn, 64k, checkpoints, sampling), and pre-warm so you never hit the cold start:</li>
</ul>
]]></description><link>https://lcz.me/post/14700</link><guid isPermaLink="true">https://lcz.me/post/14700</guid><dc:creator><![CDATA[胜威马]]></dc:creator><pubDate>Sat, 29 Aug 2026 03:19:58 GMT</pubDate></item><item><title><![CDATA[Reply to 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G，本地部署 Qwen3.8-27B 怎么吐字最快？ on Fri, 28 Aug 2026 11:10:24 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a> <a href="/post/14441">说</a>:</p>
<p dir="auto">只换显卡的话 还是 7900XTX 24G 。速度提升最快。<br />
考虑后续开发建议 R9700 32G。</p>
</blockquote>
<p dir="auto">有点纠结，7900xt好像涨价了。性价比感觉都不如9700了。想着要不要等一等，而且换了这个还要换电源，现在的电源是750w的</p>
]]></description><link>https://lcz.me/post/14609</link><guid isPermaLink="true">https://lcz.me/post/14609</guid><dc:creator><![CDATA[胜威马]]></dc:creator><pubDate>Fri, 28 Aug 2026 11:10:24 GMT</pubDate></item><item><title><![CDATA[Reply to 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G，本地部署 Qwen3.8-27B 怎么吐字最快？ on Fri, 28 Aug 2026 11:04:02 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/asd2667" aria-label="Profile: asd2667">@<bdi>asd2667</bdi></a> <a href="/post/14362">说</a>:</p>
<p dir="auto">zerodigest/Qwen3.8-27B-Uncensored-YMQ-MTP-GGUF</p>
</blockquote>
<p dir="auto">谢谢，我今天晚上试试</p>
]]></description><link>https://lcz.me/post/14608</link><guid isPermaLink="true">https://lcz.me/post/14608</guid><dc:creator><![CDATA[胜威马]]></dc:creator><pubDate>Fri, 28 Aug 2026 11:04:02 GMT</pubDate></item><item><title><![CDATA[Reply to 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G，本地部署 Qwen3.8-27B 怎么吐字最快？ on Thu, 27 Aug 2026 15:56:00 GMT]]></title><description><![CDATA[<p dir="auto">只换显卡的话 还是 7900XTX 24G 。速度提升最快。<br />
考虑后续开发建议 R9700 32G。</p>
]]></description><link>https://lcz.me/post/14441</link><guid isPermaLink="true">https://lcz.me/post/14441</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Thu, 27 Aug 2026 15:56:00 GMT</pubDate></item><item><title><![CDATA[Reply to 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G，本地部署 Qwen3.8-27B 怎么吐字最快？ on Thu, 27 Aug 2026 12:37:00 GMT]]></title><description><![CDATA[<p dir="auto">zerodigest/Qwen3.8-27B-Uncensored-YMQ-MTP-GGUF上这个模型 XS-PRO版</p>
]]></description><link>https://lcz.me/post/14362</link><guid isPermaLink="true">https://lcz.me/post/14362</guid><dc:creator><![CDATA[asd2667]]></dc:creator><pubDate>Thu, 27 Aug 2026 12:37:00 GMT</pubDate></item><item><title><![CDATA[Reply to 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G，本地部署 Qwen3.8-27B 怎么吐字最快？ on Thu, 27 Aug 2026 07:10:11 GMT]]></title><description><![CDATA[<p dir="auto">三个问题分开答，都有论坛实测数据支撑：</p>
<ol>
<li>
<p dir="auto">单卡 offload 速度：4060Ti 16G 跑 27B Q4_K_XL（约 17.5GB），权重得有一小半放 DDR4。DDR4 双通道有效带宽约 35-40GB/s，offload 部分每 token 都要过内存，实测量级大约 8-15 tok/s——比"乌龟"强一点，但长文本纠错/分析这种要读大量上下文的场景会等得很难受。不是不能跑，是体验问题。</p>
</li>
<li>
<p dir="auto">加第二张 16G 卡：最划算的路子，值得。32G 显存把 Q4 权重全放下（16.5-17.5GB）+ KV 富余，速度直接跳到 35-50 tok/s 量级（论坛实测：双 3060 24G 43-50、双 5060Ti 65；4060Ti 比 5060Ti 慢一档）。PCIe 3.0 x4 带宽约 4GB/s 够不够：tensor split 每 token 跨卡传的激活量只有几百 KB，x4 损失个位数百分比，27B 这个规模完全没问题。同型号卡用 tensor split（-ts 1,1），异构才考虑 layer split——TID:1350 有实测。</p>
</li>
<li>
<p dir="auto">量化与参数：</p>
</li>
</ol>
<ul>
<li>Q4_K_M 和 UD-Q4_K_XL 都行，UD 版对敏感张量保更高精度，纠错这类吃细节的任务推荐 UD-Q4_K_XL，体积大不了多少。</li>
<li>MTP：Qwen3.8-27B 的 MTP draft head 偏弱，接受率实测只有约 40%（TID:1350），双卡 tensor split 下跨卡同步开销还会吃掉收益——建议先关掉跑基线，想试就从 spec-draft-n-max 1 起步，2 大概率是负优化。terry 也说过多会话时投机解码收益很小。</li>
<li>64K 上下文 + KV 量化：恰恰相反，64K 时 KV 是大头，更该开量化。Qwen3.8-27B 是 GQA-4，q8_0 KV 实测约 37KB/token（TID:1350 David Chen 生产数据），64K 约 2.4GB，双卡 32G 完全没压力。Q8 KV 长上下文质量损失可忽略，放心开。</li>
</ul>
]]></description><link>https://lcz.me/post/14288</link><guid isPermaLink="true">https://lcz.me/post/14288</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 27 Aug 2026 07:10:11 GMT</pubDate></item></channel></rss>