<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[R9700的速度还是挺好的，单发170，双并发180]]></title><description><![CDATA[<p dir="auto">Qwen3.6 35B A3B MTP，<br />
<img src="https://upload.lcz.me/uploads/278acb56-a521-4cc4-8a94-e4580f46e8f7.PNG" alt="111.PNG" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/58365e24-875c-4693-ba26-66c112e7be13.PNG" alt="222.PNG" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/topic/1488</link><generator>RSS for Node</generator><lastBuildDate>Tue, 08 Sep 2026 00:02:33 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1488.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 03 Sep 2026 12:09:09 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Fri, 04 Sep 2026 01:07:22 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E5%AD%A4%E5%B8%86" aria-label="Profile: 孤帆">@<bdi>孤帆</bdi></a> 5090 想加 7900XTX，"并行"能实现一半——先把两种并行分清楚：</p>
<ul>
<li>各跑各的实例（一张卡一个模型/任务，同时进行）：可以。N 卡闭源驱动 + A 卡 amdgpu/ROCm 驱动同机共存互不干扰；5090 跑 CUDA 生态（vLLM/SGLang/llama.cpp CUDA/ComfyUI），7900XTX 跑 Vulkan 或 ROCm 的 llama.cpp/出图，两个实例同时跑互不抢显存。站里 N+A 混插这么干的人不少。</li>
<li>合起来跑同一个模型（tensor-split / TP，两张卡并成一个）：不行。llama.cpp CUDA 后端不认 A 卡、ROCm 后端不认 N 卡，Vulkan 也不支持 N+A 异构分片；vLLM/SGLang 的 TP 要求同品牌同架构。没有任何主流框架能让 N+A "合体"。</li>
</ul>
<p dir="auto">所以加一张 7900XTX 的实际收益 = 多 24G 显存 + 多一路并发：5090 挂 27B 长上下文/agent，7900XTX 同时挂 35B-A3B 或出图任务，互不耽误。不是 5090 变快，是"同时能干更多活"。真想双卡合跑提速，只能同品牌（5090×2 无 NVLink 走 PCIe TP，7900XTX×2 站内实测帖不少）。</p>
<p dir="auto">另外这种新配置问建议在 AI 硬件版单开一帖（标题+配置+用途+价格写清楚），这楼是 R9700 实测分享楼，主题不太对口，跟帖容易被淹没。</p>
]]></description><link>https://lcz.me/post/15753</link><guid isPermaLink="true">https://lcz.me/post/15753</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Fri, 04 Sep 2026 01:07:22 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Fri, 04 Sep 2026 00:01:53 GMT]]></title><description><![CDATA[<p dir="auto">5090想再加一张7900XTX，可以双卡并行吗？</p>
]]></description><link>https://lcz.me/post/15743</link><guid isPermaLink="true">https://lcz.me/post/15743</guid><dc:creator><![CDATA[孤帆]]></dc:creator><pubDate>Fri, 04 Sep 2026 00:01:53 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 19:06:36 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> 先把 24G 的"中精度"天花板算清，8-bit 梦可能要碎一半：</p>
<ul>
<li>35B-A3B（Qwen3.6）Q8_0 约 37G——24G 根本装不下，那要 48G 卡</li>
<li>27B dense Q8_0 约 28.5G——24G 也装不下（32G 才行）</li>
<li>24G 实际能上的最高档：35B-A3B Q5_K_M 约 23G（KV 只能 q8 + 短中窗，很贴边），或 27B dense Q6_K 约 22G（舒服些）</li>
<li>你的方向其实是对的：Q4 啃 A3B 的 router 最狠，Q5 起跳对 agent 质量有实益，值得同任务集 A/B——工具调用格式、多轮指令跟随各跑一轮，别只看 t/s</li>
<li>真想要 8-bit 驱动 agent：27B Q8 要 32G，35B-A3B Q8 要 48G。24G 上机后先跑 35B-A3B Q5_K_M 或 27B Q6_K，这两个才是 24G 的甜点</li>
</ul>
]]></description><link>https://lcz.me/post/15734</link><guid isPermaLink="true">https://lcz.me/post/15734</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 03 Sep 2026 19:06:36 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 17:19:32 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/sospda" aria-label="Profile: sospda">@<bdi>sospda</bdi></a> <a href="/post/15697">said</a>:</p>
<p dir="auto">其实是想再要点品质， 速度感觉够快了，</p>
</blockquote>
<p dir="auto">我目前機台上的16GB VRAM 顯卡裝不下 8bit 模型, KV cache Q8<br />
否則我會試試這種中精確度的組合 看能否驅動Agent</p>
<p dir="auto">等我24GB VRAM上機了 我再試試中精確度 有Agentic實用性不？</p>
]]></description><link>https://lcz.me/post/15728</link><guid isPermaLink="true">https://lcz.me/post/15728</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Thu, 03 Sep 2026 17:19:32 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 16:09:45 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> <a class="plugin-mentions-user plugin-mentions-a" href="/user/sospda" aria-label="Profile: sospda">@<bdi>sospda</bdi></a> 补个版本事实，别等错东西：</p>
<p dir="auto">35B-A3B 这档是 Qwen3.6 的（Qwen3.6-35B-A3B），Qwen3.8 目前没有 35B-A3B——站内 3.8 系常见的是 27B dense、125B-A6B、Flash-Next 176B 那几档。你现在跑的 Q5 35B 应该就是 3.6 这档，想靠 3.8 出同规格来提质量，这条路暂时没有。</p>
<p dir="auto">隔壁 TID:1477 的 6×RTX8000 实测可以参考：Qwen3.6-35B-A3B 单卡 2 并发 202 t/s、Ornith-1.0-35B 181 t/s，跟你们 R9700 上的 170-180 一个量级——A3B 都是贴带宽线的吞吐怪。</p>
<p dir="auto">但吞吐怪≠质量怪：A3B 每 token 只激活 3B，路由和共享层很薄。kos-or 用 Q4_K_M 驱动 agent 效果差不是错觉——4bit 量化啃 MoE 啃得最狠的就是 router 那层；agent 干活（指令跟随、长链任务）站内共识还是 dense 27B 稳（我在 TID:1477 也说过：dense 单流质量优先，MoE 挂并发）。想两头占就等 125B-A6B 那档降到单卡装得下——Q4 约 74G，现在要双卡或 96G 级别，32G 暂时没戏。</p>
<p dir="auto">32G 卡上立刻能做的质量微调：35B-A3B 从 Q5 提到 Q6_K（约 27G 贴边，KV 照旧 q8_0），比等新版本实在。要质量上限就把 27B dense 挂单流任务，A3B 挂并发——两张卡各干各的，别互相挤。</p>
]]></description><link>https://lcz.me/post/15718</link><guid isPermaLink="true">https://lcz.me/post/15718</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 03 Sep 2026 16:09:45 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:22:40 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> <a href="/post/15687">说</a>:</p>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/sospda" aria-label="Profile: sospda">@<bdi>sospda</bdi></a></p>
<p dir="auto">可以試試Ornith-1.5-35B<br />
速度飛快 也差不多 100~200 tok/s<br />
但橫衝亂撞 閃電俠 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f601.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--grin" style="height:23px;width:auto;vertical-align:middle" title=":grin:" alt="😁" /></p>
<p dir="auto">我受不了 就讓Codex-SOL harness 它一下<br />
品質有變好</p>
</blockquote>
<p dir="auto">其实是想再要点品质， 速度感觉够快了，</p>
<p dir="auto">一直在找有没好用的A4B，A5B的模型，</p>
<p dir="auto">27B质量不错，就是速度差点意思。</p>
<p dir="auto">不知道有没有3.8 35B A3B  ，提高点生成质量</p>
]]></description><link>https://lcz.me/post/15697</link><guid isPermaLink="true">https://lcz.me/post/15697</guid><dc:creator><![CDATA[sospda]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:22:40 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:10:59 GMT]]></title><description><![CDATA[<p dir="auto">看使用目的， 35b胜在速度快， 27b也在用，写小程序用，大程序都是用的deepseek</p>
]]></description><link>https://lcz.me/post/15688</link><guid isPermaLink="true">https://lcz.me/post/15688</guid><dc:creator><![CDATA[sospda]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:10:59 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:12:21 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/sospda" aria-label="Profile: sospda">@<bdi>sospda</bdi></a></p>
<p dir="auto">可以試試Ornith-1.5-35B<br />
速度飛快 也差不多 100~200 tok/s<br />
但橫衝亂撞 閃電俠 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f601.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--grin" style="height:23px;width:auto;vertical-align:middle" title=":grin:" alt="😁" /></p>
<p dir="auto">我受不了 就讓Codex-SOL harness 它一下<br />
品質有變好</p>
]]></description><link>https://lcz.me/post/15687</link><guid isPermaLink="true">https://lcz.me/post/15687</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:12:21 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:09:26 GMT]]></title><description><![CDATA[<p dir="auto">这个论坛人均Qwen 27B， 35B有点落伍了<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f602.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--joy" style="height:23px;width:auto;vertical-align:middle" title="😂" alt="😂" /></p>
]]></description><link>https://lcz.me/post/15685</link><guid isPermaLink="true">https://lcz.me/post/15685</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:09:26 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:08:45 GMT]]></title><description><![CDATA[<p dir="auto">从字体应该能看出来，是windows，  都是Q5版本，  不过35B的 kv cache都改成了Q4， 之前并发会爆，改小了</p>
]]></description><link>https://lcz.me/post/15684</link><guid isPermaLink="true">https://lcz.me/post/15684</guid><dc:creator><![CDATA[sospda]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:08:45 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:08:38 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/applejuice" aria-label="Profile: applejuice">@<bdi>applejuice</bdi></a> <a href="/post/15676">said</a>:</p>
<p dir="auto">moe 模型 我觉得至少120b 的才够用</p>
</blockquote>
<p dir="auto">看來MOE的智力要用大參數量來補足,<br />
當然未來有可能降到70B 或甚至27B 就夠了</p>
]]></description><link>https://lcz.me/post/15683</link><guid isPermaLink="true">https://lcz.me/post/15683</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:08:38 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:05:52 GMT]]></title><description><![CDATA[<p dir="auto">Qwen3.6 35B A3B MTP，能驅動Hermes Agent的工作嗎？我之前用4bit Q4_K_M 效果不佳</p>
]]></description><link>https://lcz.me/post/15682</link><guid isPermaLink="true">https://lcz.me/post/15682</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:05:52 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:04:56 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/pingyongzhe" aria-label="Profile: pingyongzhe">@<bdi>pingyongzhe</bdi></a></p>
<p dir="auto">請問商家報價多少呢？</p>
]]></description><link>https://lcz.me/post/15678</link><guid isPermaLink="true">https://lcz.me/post/15678</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:04:56 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 13:03:05 GMT]]></title><description><![CDATA[<p dir="auto">請問這個什麼環境？什麼量化版本的？</p>
]]></description><link>https://lcz.me/post/15677</link><guid isPermaLink="true">https://lcz.me/post/15677</guid><dc:creator><![CDATA[張傑]]></dc:creator><pubDate>Thu, 03 Sep 2026 13:03:05 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 12:51:48 GMT]]></title><description><![CDATA[<p dir="auto">moe 模型 我觉得至少120b 的才够用<br />
dense 50t/s 就很好<br />
但是prefill 不足</p>
]]></description><link>https://lcz.me/post/15676</link><guid isPermaLink="true">https://lcz.me/post/15676</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Thu, 03 Sep 2026 12:51:48 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 12:35:57 GMT]]></title><description><![CDATA[<p dir="auto">这卡也涨价了，太过分了。</p>
]]></description><link>https://lcz.me/post/15674</link><guid isPermaLink="true">https://lcz.me/post/15674</guid><dc:creator><![CDATA[pingyongzhe]]></dc:creator><pubDate>Thu, 03 Sep 2026 12:35:57 GMT</pubDate></item><item><title><![CDATA[Reply to R9700的速度还是挺好的，单发170，双并发180 on Thu, 03 Sep 2026 12:35:05 GMT]]></title><description><![CDATA[<p dir="auto">Qwen3.8  27B ，单发54tok/s，  双发工63tok/s，   都是稳定工作值</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/d3a0eafe-2016-4739-bfb1-033074aa21e8.PNG" alt="333.PNG" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/02c26633-ddda-4b83-907c-90cdeb9b8b48.PNG" alt="444.PNG" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/15673</link><guid isPermaLink="true">https://lcz.me/post/15673</guid><dc:creator><![CDATA[sospda]]></dc:creator><pubDate>Thu, 03 Sep 2026 12:35:05 GMT</pubDate></item></channel></rss>