<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。]]></title><description><![CDATA[<p dir="auto">我是一个新手，还在摸索阶段，只知道复制别人的东西过来修修改改。</p>
<p dir="auto">huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF<br />
这个huihui的模型里面，有一个来自于 ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF模型的去审查版本，只有11G左右，但是质量非常高，由于是 ISTA-DASLab这个全球开源AI研究的顶尖实验室的作品，其Q3S的部分能力比起BF16近乎无损。</p>
<p dir="auto">目前从官方评测来看，它应该是量化的qwen3.8 27B模型里面又小又强的一个典范。</p>
<p dir="auto">我的配置是i5 14400 B760主板 32G内存 5060TI 16G（穷，从4060 8G升级过来的），这是载入后的内存和显存状况，基本快满了，但是跑起来速度还可以，30-40T是有的，上下文设置看启动参数，目前不知道怎么计算多少合适。<br />
<img src="https://upload.lcz.me/uploads/4da67f02-828d-4f05-a209-19f66a449924.jpeg" alt="3f93969e-acff-4849-97e8-01a71c58a545-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/5930a393-c09e-4c79-af76-11930d3fef7b.jpeg" alt="4fc253e6-c6fc-41ec-b18c-149b84c9c764-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">下面是我的启动参数：</p>
<p dir="auto">llama-server.exe ^<br />
--models-dir D:\llamacpp\models ^<br />
--mmproj "%MODEL_DIR%%MODEL_MMP%" ^<br />
--models-max 1 ^<br />
--host 127.0.0.1 ^<br />
--port 8080 ^<br />
-c 65536 ^<br />
-np 1 ^<br />
-ngl 99 ^<br />
--fit off ^<br />
-fa on ^<br />
--jinja ^<br />
--reasoning auto ^<br />
--reasoning-format auto ^<br />
--ui-mcp-proxy ^<br />
--cache-type-k q8_0 ^<br />
--cache-type-v q8_0 ^<br />
--spec-type draft-mtp ^<br />
--spec-draft-n-max 2 ^<br />
-t 8</p>
<p dir="auto">看看下一步能不能跑hermes看看，目前还在学，不知道怎么弄。<br />
有没有大神指点一下这个还有什么优化改进空间。</p>
]]></description><link>https://lcz.me/topic/1817</link><generator>RSS for Node</generator><lastBuildDate>Sun, 20 Sep 2026 16:01:46 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1817.rss" rel="self" type="application/rss+xml"/><pubDate>Sat, 19 Sep 2026 06:02:07 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sun, 20 Sep 2026 12:54:18 GMT]]></title><description><![CDATA[<p dir="auto">不知道有没有人用T10的显卡。。。。。</p>
]]></description><link>https://lcz.me/post/19540</link><guid isPermaLink="true">https://lcz.me/post/19540</guid><dc:creator><![CDATA[PENG XU]]></dc:creator><pubDate>Sun, 20 Sep 2026 12:54:18 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sun, 20 Sep 2026 12:07:51 GMT]]></title><description><![CDATA[<p dir="auto">用 RTX 5070 Ti 測試了 Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf (9.6 GB)</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/b63e4c98-83c9-4b77-90d1-3aac7339e451.jpeg" alt="image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/19538</link><guid isPermaLink="true">https://lcz.me/post/19538</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Sun, 20 Sep 2026 12:07:51 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sun, 20 Sep 2026 01:23:46 GMT]]></title><description><![CDATA[<p dir="auto">我认为5060ti 16G真是一个入门摸的神卡，便宜能学</p>
]]></description><link>https://lcz.me/post/19410</link><guid isPermaLink="true">https://lcz.me/post/19410</guid><dc:creator><![CDATA[byronduck-jpg]]></dc:creator><pubDate>Sun, 20 Sep 2026 01:23:46 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sat, 19 Sep 2026 16:03:09 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> 补充 DFlash2 那条：它和模型自带的 MTP 是两条路子，一般不能简单叠加。</p>
<ul>
<li>MTP 是模型自带的多 token 预测头，和主模型共用权重，不额外加载 draft。</li>
<li>DFlash2 是外挂的 draft 模型做投机解码，要单独一份小模型，显存和时间都多花。</li>
<li>多数运行时只允许一个 draft 源，先确认能不能同时开，多半是二选一。</li>
</ul>
<p dir="auto">要试就分三组测：关投机 / 只 MTP / 只 DFlash2，比接受长度、接受率和每 token 墙钟时间，别只看瞬时 t/s。另外你标题带 mtp，先确认它真在跑（看接受率），否则 40 t/s 就是 448÷11 的带宽上限，换 draft 也提不了多少。</p>
]]></description><link>https://lcz.me/post/19370</link><guid isPermaLink="true">https://lcz.me/post/19370</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sat, 19 Sep 2026 16:03:09 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sat, 19 Sep 2026 14:24:42 GMT]]></title><description><![CDATA[<p dir="auto">5060Ti 16G CP值很高 漲價前 NT$19500, 是入門款很不錯的測試平台 也是很熱門的遊戲卡</p>
<p dir="auto">omg =&gt; Downloads last month 1,154,265 感謝分享<br />
它裡面有個mtp 模型可以試試,<br />
不知道能不能用DFlash2, 速度可能會更快, 但據說輸出品質會下降</p>
]]></description><link>https://lcz.me/post/19355</link><guid isPermaLink="true">https://lcz.me/post/19355</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Sat, 19 Sep 2026 14:24:42 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sat, 19 Sep 2026 13:02:13 GMT]]></title><description><![CDATA[<p dir="auto">结论：这样加不会更快，大概率还更慢。你的场景不缺显存，缺的是带宽，而第二张卡带宽只有一半。</p>
<ul>
<li>RTX PRO 4500 32G 带宽约 896 GB/s，5060Ti 16G 约 448 GB/s，正好一半。Qwen3.8-27B 的 NVFP4 权重 13~14G 级，你现在的 50~75 t/s 基本贴着 896/13.5≈66 的上限，单卡已经没浪费。</li>
<li>TP=2 是每层按维度对半切，快卡算完自己那半要等慢卡，单层耗时由 5060Ti 决定，等效带宽约 2×448=896，和单卡一样；再叠加消费卡没有 NVLink、all-reduce 走 PCIe（第二槽常是 x8 甚至芯片组 x4），净结果通常低于单卡。</li>
<li>显存也不缺：13~14G 权重加 KV，32G 够用，不存在卸载到内存被拖慢的问题，所以没有“补容量换速度”的空间。</li>
</ul>
<p dir="auto">5060Ti 有价值的用法：</p>
<ol>
<li>容量卡：跑更大的模型（70B 级量化）或超长 ctx，这时才值得上 TP/层切；同一个模型不会加速。</li>
<li>副卡：embedding/rerank、ComfyUI，或小 draft 模型做投机解码——这是这套硬件唯一可能真提 t/s 的路子，但它靠提高接受长度，不改带宽上限。</li>
<li>想单卡更快：降 ctx、确认 KV 走 fp8/q8、确认 NVFP4 真走的 FP4 kernel（没被反量化回 bf16）、把 MTP 打开。</li>
</ol>
<p dir="auto">验一下再决定：nvidia-smi topo -m 看两卡是否 CPU 直连、有无 P2P；然后单卡与 TP=2 各跑一次 pp/tg 对比。若 5060Ti 插在芯片组 x4 槽，all-reduce 会明显拖后腿，结论会更差。</p>
]]></description><link>https://lcz.me/post/19352</link><guid isPermaLink="true">https://lcz.me/post/19352</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sat, 19 Sep 2026 13:02:13 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sat, 19 Sep 2026 10:24:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a></p>
<p dir="auto">我现在有一张rtxpro 4500 32gb, 我加入另外一张5060ti 16gb, 在qwen3.8 27b nvfp4 bf16, 速度上可以在加速吗？现在是 50ts - 75ts 之间</p>
]]></description><link>https://lcz.me/post/19333</link><guid isPermaLink="true">https://lcz.me/post/19333</guid><dc:creator><![CDATA[imbiplaza ASUS]]></dc:creator><pubDate>Sat, 19 Sep 2026 10:24:58 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sat, 19 Sep 2026 07:02:29 GMT]]></title><description><![CDATA[<p dir="auto">启动参数里的 ctx 不是拍脑袋填的，它直接吃显存，按下面的账算就不会爆。</p>
<p dir="auto">先给个定心丸：5060Ti 16G 的显存带宽约 448 GB/s，这个 IQ3_S 权重 11G 左右，单请求 decode 的理论上限大约是 448 ÷ 11 ≈ 40 t/s。你现在 30–40 t/s 基本贴着上限跑了，再挤也就几个点，别为了速度把上下文开爆。标题里带 mtp，如果真开了投机解码，接受长度够的话能超过这个数，但 MTP 不改带宽上限，只是每步多吐几个 token。</p>
<p dir="auto">上下文怎么估——KV cache 只跟层数、KV 头数、头维、KV 精度有关，跟权重多大无关：<br />
每 token 字节 = 2(K+V) × n_layer × n_kv_head × head_dim × dtype 字节<br />
Qwen3.8-27B 这类的 n_layer / n_head_kv / n_embd_head 会在 llama-server 启动日志里打印出来，照着代：</p>
<ul>
<li>KV 用 fp16：约 0.25 MiB/token，32k 上下文约 8G；</li>
<li>KV 用 q8_0：约 0.13 MiB/token，32k 约 4G。</li>
</ul>
<p dir="auto">你这张卡：权重 11G + 计算缓冲 1–1.5G，留给 KV 的只有 3–4G。所以建议：</p>
<ul>
<li>先 --ctx-size 16384 跑，日志会打印 KV self size，确认余量再往上加；</li>
<li>想上 32k 就加 --cache-type-k q8_0 --cache-type-v q8_0（质量损失很小）；</li>
<li>--flash-attn on 必开，--n-gpu-layers 99 尽量全上卡，-ub 512 -b 2048 起步。</li>
</ul>
<p dir="auto">32G 内存“快满”是 mmap 把权重算进了 page cache，不是泄漏，只要不开始 swap 就没事；真被内存卡住就减 ctx 或临时 --no-mmap。要 64k+ 的上下文，优先把内存加到 64G，别硬压这张 16G 卡。</p>
<p dir="auto">楼上说让 agent 帮你配参数可以，但显存/上下文的账得自己核一遍，agent 很爱给一个直接 OOM 的 --ctx-size。</p>
]]></description><link>https://lcz.me/post/19309</link><guid isPermaLink="true">https://lcz.me/post/19309</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sat, 19 Sep 2026 07:02:29 GMT</pubDate></item><item><title><![CDATA[Reply to 新手贴：配置32G内存+5060TI 16G跑Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp，在40上下tokens/s。 on Sat, 19 Sep 2026 06:07:38 GMT]]></title><description><![CDATA[<p dir="auto">跑hermes 讓anitigravity 來配就行~直接用講的</p>
]]></description><link>https://lcz.me/post/19308</link><guid isPermaLink="true">https://lcz.me/post/19308</guid><dc:creator><![CDATA[鍾子揚]]></dc:creator><pubDate>Sat, 19 Sep 2026 06:07:38 GMT</pubDate></item></channel></rss>