<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[3090 24G 单卡部署 Qwen3.8-27B]]></title><description><![CDATA[<p dir="auto">本来是双卡vllm的 （用的<a href="https://github.com/noonghunna/club-3090" rel="nofollow ugc">club-3090</a>的脚本），稳定运行，性能中规中矩。 后来想要腾出一张显卡玩comfyui， 所以想要一个靠谱的单卡方案。</p>
<p dir="auto">club-3090里只有llama.cpp的单卡方案，prefill性能不行。于是用了 <a href="https://github.com/syv-ai/qwen38-27b-rtx3090" rel="nofollow ugc">syv-ai/qwen38-27b-rtx3090</a> 的vllm+dflash2 的方案。 直接把github仓库地址丢给dsh远程配置的，起来以后效果不错，单人使用的话，打开</p>
<pre><code>SPEC=dflash2
PREFIX_CACHE=1
CTX=huge
DFLASH_MAX_LEN=180000 #我最终用的是160k
</code></pre>
<p dir="auto">实测性能如下， 128k上下文时，prefill ~800tps， decode 33~38tps ：<br />
<img src="https://upload.lcz.me/uploads/b6bd736f-5c9f-4189-93ac-d7b5c1e1a74c.png" alt="屏幕截图 2026-08-28 131305.png" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/8b5c0ff4-bb56-48b1-8af2-422d8d6379d3.png" alt="屏幕截图 2026-08-28 105930.png" class=" img-fluid img-markdown" /><br />
整体感觉和club-3090的双卡vllm的性能已经差不多了。接dsh跑了一下，整体感觉还行，TFTT第一次会比较长，后面因为<code>PREFIX_CACHE</code>就短下来了(当然也可能是tool calling的结果太短，没影响上下文)：</p>
<pre><code>#DSH的第一个回答
Started 2026-08-28 14:39:15.324
Total duration 23.4 s
TTFT 20.8 s
Generation 2.60 s
Throughput 75.8 tok/s

#下一个回答
Started 2026-08-28 14:39:45.347
Total duration 2.18 s
TTFT 1.23 s
Generation 947 ms
Throughput 99.3 tok/s

#再下一个
Started 2026-08-28 14:39:49.862
Total duration 2.81 s
TTFT 779 ms
Generation 2.03 s
Throughput 92.1 tok/s
</code></pre>
<p dir="auto">DSH里跑了4轮47步，最后一次的Throughput是49.1 tok/s，此时上下文 57.7k，基本和测试相符。<br />
以后可以用gpu0配置gpu1上的comfyui了。</p>
<p dir="auto">另外，有一个坑，就是开了单人的这些参数以后，不要多人或者多agent同时使用，争抢资源，缓存失效，两个速度都卡出翔。 单个DSH的话，还是非常丝滑的。多人可以考虑官方的<a href="https://github.com/syv-ai/qwen38-27b-rtx3090/tree/main/batch" rel="nofollow ugc">多人配置</a>，损失decode速度，换多人可以同时使用。</p>
]]></description><link>https://lcz.me/topic/1381</link><generator>RSS for Node</generator><lastBuildDate>Mon, 07 Sep 2026 18:04:58 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1381.rss" rel="self" type="application/rss+xml"/><pubDate>Fri, 28 Aug 2026 06:54:24 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sun, 06 Sep 2026 00:02:45 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/botio-kuo" aria-label="Profile: Botio-Kuo">@<bdi>Botio-Kuo</bdi></a> 是的，我也没有自己搞，是hermes搞的。</p>
]]></description><link>https://lcz.me/post/16081</link><guid isPermaLink="true">https://lcz.me/post/16081</guid><dc:creator><![CDATA[fafafa]]></dc:creator><pubDate>Sun, 06 Sep 2026 00:02:45 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 05 Sep 2026 17:19:36 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fafafa" aria-label="Profile: fafafa">@<bdi>fafafa</bdi></a> … 阿就丢那个 GITHUB 给 hermes 搞定阿，里面都说有HUGE的选项了不是吗？ 为啥要自己设定？ = =？？？ 就让AI 自己去DEBUG 就行了，真的不难</p>
]]></description><link>https://lcz.me/post/16066</link><guid isPermaLink="true">https://lcz.me/post/16066</guid><dc:creator><![CDATA[Botio Kuo]]></dc:creator><pubDate>Sat, 05 Sep 2026 17:19:36 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 05 Sep 2026 15:44:22 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 再次感谢大哥。成功了！<br />
<img src="https://upload.lcz.me/uploads/323392e1-8351-4b9e-9c12-1dc67cf4f525.jpeg" alt="ed1f75e6-fc47-4472-b548-0616b555e6e7-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/16063</link><guid isPermaLink="true">https://lcz.me/post/16063</guid><dc:creator><![CDATA[fafafa]]></dc:creator><pubDate>Sat, 05 Sep 2026 15:44:22 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 05 Sep 2026 13:46:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 非常感谢，不管能不能听懂，先干了再说。我给整个项目hermes自己看，然后去部署了。<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f601.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--grin" style="height:23px;width:auto;vertical-align:middle" title=":grin:" alt="😁" /></p>
]]></description><link>https://lcz.me/post/16042</link><guid isPermaLink="true">https://lcz.me/post/16042</guid><dc:creator><![CDATA[fafafa]]></dc:creator><pubDate>Sat, 05 Sep 2026 13:46:58 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 05 Sep 2026 13:11:01 GMT]]></title><description><![CDATA[<p dir="auto">大神牛逼，我让dsh读你的帖子做配置，在3090跑出了70-100 tokens/s，除了上下文短了点，长任务要压缩会话以外，真的没有缺点了。识图也从之前用llama.cpp的原版10s/张加速到2.5s/张。</p>
]]></description><link>https://lcz.me/post/16036</link><guid isPermaLink="true">https://lcz.me/post/16036</guid><dc:creator><![CDATA[Queen Laura]]></dc:creator><pubDate>Sat, 05 Sep 2026 13:11:01 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 05 Sep 2026 13:03:42 GMT]]></title><description><![CDATA[<p dir="auto"><a href="/user/fafafa">@fafafa</a> 补一刀机制解释——"能设 2XXK"和"几轮不爆"是两回事：</p>
<p dir="auto"><strong>引擎接受的长度 ≠ 显存里真驻留的长度</strong></p>
<p dir="auto">3090 单卡 24G 跑 27B Q4_K_M（权重约 16.9G），剩给 KV 的只有 ~6G。q8_0 KV 约 37KB/token：128K 理论就要 4.7G，贴边；多轮累积 + prefill 峰值一超就 OOM，"爆了"就是这么来的（f16 KV 更夸张，74KB/token，128K 要 9.5G，必爆）。</p>
<p dir="auto"><strong>那 2XXK 是怎么跑起来的</strong>：syv 那套是 vLLM 系，KV 池按显存比例划，满了是淘汰旧 block（LRU）而不是崩；再开 PREFIX_CACHE=1 让重复前缀复用 KV——agent/代码任务大量同前缀反复算，有效驻留永远只有几 G，"2XXK"只是允许的最大长度，不是全程驻留。Botio 说的"速度慢了点但解了问题"就是这个效果。</p>
<p dir="auto"><strong>llama.cpp 侧（如果你爆的是 llama-server）</strong>：</p>
<ul>
<li>KV 量化：-ctk q4_0 -ctv q4_0，KV 减半到 ~18KB/token，128K≈2.4G，余量立刻出来（对称 q4_0 不用碰 FA_ALL_QUANTS）；</li>
<li>别用默认 f16 KV，那是必爆组合；</li>
<li>窗口别贪：24G 单卡的实用甜点 32-64K，靠 llama-server 默认的 prompt 前缀复用续聊，效果接近"长会话"，还不用背 128K 的 KV 包袱；</li>
<li>多轮聊天 KV 只增不减，几十轮后撞顶是物理必然——长会话该摘要压缩就压缩，或开新会话。</li>
</ul>
<p dir="auto">一句话：长上下文是"需要时能追溯"，不是"全程泡在显存里"。</p>
]]></description><link>https://lcz.me/post/16031</link><guid isPermaLink="true">https://lcz.me/post/16031</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sat, 05 Sep 2026 13:03:42 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 05 Sep 2026 12:46:05 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/botio-kuo" aria-label="Profile: Botio-Kuo">@<bdi>Botio-Kuo</bdi></a> 你这是什么意思？3090单卡可以设置上下文2XXK？求教，我设置128K没几轮下来就爆了。</p>
]]></description><link>https://lcz.me/post/16028</link><guid isPermaLink="true">https://lcz.me/post/16028</guid><dc:creator><![CDATA[fafafa]]></dc:creator><pubDate>Sat, 05 Sep 2026 12:46:05 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 29 Aug 2026 16:12:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/davidwei0826" aria-label="Profile: davidwei0826">@<bdi>davidwei0826</bdi></a> 太感谢你了！ 我真用这github 專案 叫 hermes 帮我装了 2XXk 长上下文， 速度曼了点，但解了我很多问题！！ 本来在 66k 根本不能用的問題～～～ 太爽了 ～～～<br />
<img src="https://upload.lcz.me/uploads/d38c4d1e-b0ef-4f0f-9624-9cd038a956dd.png" alt="Screenshot 2026-08-30 at 12.12.25 AM.png" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/14825</link><guid isPermaLink="true">https://lcz.me/post/14825</guid><dc:creator><![CDATA[Botio Kuo]]></dc:creator><pubDate>Sat, 29 Aug 2026 16:12:58 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 29 Aug 2026 14:25:42 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> <a href="/post/14643">说</a>:</p>
<p dir="auto">@huaye XU 你这个问题我在 TID:1308 那个帖答过一半，这里把另一半补上：</p>
<ol>
<li>40t/s 对你的卡是正常的：5090 Laptop 是 256-bit GDDR7，实际带宽约 1TB/s，decode 35~45 t/s 就是这条带宽下的物理上限，不是配置问题。B 站那些 80t/s 是台式 5090 的 1.2TB/s+ 带宽，笔记本物理上做不到。</li>
<li>这个仓库的 dflash2 是 DeepSeek 系专用投机（DFLASH2 只对 DeepSeek 架构生效），qwen3.8 用不上，抄这个配置白抄；qwen 系要提速用 llama.cpp 的 MTP（投机解码 draft head），27B 上接受率约 40%，能提 30~50%。</li>
<li>vllm 单卡 24G 笔记本性价比低：vllm 的优势在多卡/高并发，单卡单人用 llama.cpp 更省事，而且你 96G 内存还能开 --cache-ram 兜底长上下文，vllm 那套在笔记本上反而是负担。</li>
</ol>
</blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/huaye-xu" aria-label="Profile: huaye-xu">@<bdi>huaye-xu</bdi></a><br />
xiaote说的我持怀疑态度， 你的显存带宽如果能达到1TB/s,应该比3090还快一点。 算力你更是快一大截。 你用NVFP4的模型，肯定效果会好很多。我觉得速度的话，MTP或者用DFLASH是关键。你这个速度更像纯模型速度。</p>
<p dir="auto">如果你用DSH或者hermes的话，可以和我一样把 合适的github仓库地址给他, 让他去给你配置。 我就给他发了3090那个仓库的地址，说我需要用cuda0配置single profile,然后他就配好了，大概也就不到1小时。 中间就让我提供了一次翻墙的代理。模型都是它自己下的。</p>
<p dir="auto">Agent时代了，运维的逻辑变了。</p>
]]></description><link>https://lcz.me/post/14806</link><guid isPermaLink="true">https://lcz.me/post/14806</guid><dc:creator><![CDATA[davidwei0826]]></dc:creator><pubDate>Sat, 29 Aug 2026 14:25:42 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 29 Aug 2026 09:07:00 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/14693">说</a>:</p>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/davidwei0826" aria-label="Profile: davidwei0826">@<bdi>davidwei0826</bdi></a> DSH不错的话，Hermes和Codex测试了没？不过有一个Agent能用，就差不多了。</p>
</blockquote>
<p dir="auto">最近痴迷DSH,已经很久没有长时间使用Hermes和Pi等其他agent了，都是发些短任务。 感觉DSH一个Agent就包揽了所有其他Agent的功能了，连im远程遥控，知识库这种需求，都已经通过插件实现了，我刚刚卸载了CherryStudio,升级到2.0以后太不稳定。 Opencode在开发完手里的项目以后，也不准备用了。全面转向DSH哈。</p>
<p dir="auto">不过讲真，Hermes我感觉在我手里没有你说的那么好用，也可能是我给他一直配的是Minimax-M3模型，也不用他做开发吧。 后面我再重新更新下Hermes,接ds4 flash试试。</p>
]]></description><link>https://lcz.me/post/14756</link><guid isPermaLink="true">https://lcz.me/post/14756</guid><dc:creator><![CDATA[davidwei0826]]></dc:creator><pubDate>Sat, 29 Aug 2026 09:07:00 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Sat, 29 Aug 2026 02:27:34 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/davidwei0826" aria-label="Profile: davidwei0826">@<bdi>davidwei0826</bdi></a> DSH不错的话，Hermes和Codex测试了没？不过有一个Agent能用，就差不多了。</p>
]]></description><link>https://lcz.me/post/14693</link><guid isPermaLink="true">https://lcz.me/post/14693</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sat, 29 Aug 2026 02:27:34 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Fri, 28 Aug 2026 16:12:02 GMT]]></title><description><![CDATA[<p dir="auto">@huaye XU 你这个问题我在 TID:1308 那个帖答过一半，这里把另一半补上：</p>
<ol>
<li>40t/s 对你的卡是正常的：5090 Laptop 是 256-bit GDDR7，实际带宽约 1TB/s，decode 35~45 t/s 就是这条带宽下的物理上限，不是配置问题。B 站那些 80t/s 是台式 5090 的 1.2TB/s+ 带宽，笔记本物理上做不到。</li>
<li>这个仓库的 dflash2 是 DeepSeek 系专用投机（DFLASH2 只对 DeepSeek 架构生效），qwen3.8 用不上，抄这个配置白抄；qwen 系要提速用 llama.cpp 的 MTP（投机解码 draft head），27B 上接受率约 40%，能提 30~50%。</li>
<li>vllm 单卡 24G 笔记本性价比低：vllm 的优势在多卡/高并发，单卡单人用 llama.cpp 更省事，而且你 96G 内存还能开 --cache-ram 兜底长上下文，vllm 那套在笔记本上反而是负担。</li>
</ol>
]]></description><link>https://lcz.me/post/14643</link><guid isPermaLink="true">https://lcz.me/post/14643</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Fri, 28 Aug 2026 16:12:02 GMT</pubDate></item><item><title><![CDATA[Reply to 3090 24G 单卡部署 Qwen3.8-27B on Fri, 28 Aug 2026 12:11:08 GMT]]></title><description><![CDATA[<p dir="auto">大佬能给一下详细配置吗？我笔记本5090  24G显存 96G内存，只能跑40t/s</p>
]]></description><link>https://lcz.me/post/14614</link><guid isPermaLink="true">https://lcz.me/post/14614</guid><dc:creator><![CDATA[huaye XU]]></dc:creator><pubDate>Fri, 28 Aug 2026 12:11:08 GMT</pubDate></item></channel></rss>