<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[一直没找到在DGX上能完全满意的一套模型配置]]></title><description><![CDATA[<p dir="auto">最近尝试SGLang部署了Qwen3:coder-next-Fp8+qwen3-vl+bge-m3,也试了vllm的Gemma4:31b用mtp=3优化总算速度起来点，但在hermes 的kanban跑长程任务，总是各种问题，幻觉，偷懒，自己找台阶，一言难尽。</p>
]]></description><link>https://lcz.me/topic/980/一直没找到在dgx上能完全满意的一套模型配置</link><generator>RSS for Node</generator><lastBuildDate>Tue, 11 Aug 2026 13:47:01 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/980.rss" rel="self" type="application/rss+xml"/><pubDate>Fri, 31 Jul 2026 06:01:30 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Mon, 10 Aug 2026 04:55:19 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/linkdesu" aria-label="Profile: linkdesu">@<bdi>linkdesu</bdi></a> 我已經用兩台跑很久了，0731出來也已經換上了，真心不錯，速度快，kv緩存也夠500K x 6.<br />
不過coding上使用起來感覺還是差glm 5.2一點，純個人感覺，沒有甚麼根據.<br />
單純論性價比,dsv4f用兩台，glm 5.2至少要四台，dsv4f還是比較高的</p>
]]></description><link>https://lcz.me/post/11823</link><guid isPermaLink="true">https://lcz.me/post/11823</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Mon, 10 Aug 2026 04:55:19 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 09 Aug 2026 15:27:09 GMT]]></title><description><![CDATA[<p dir="auto">这两天试了下claude code+sglang+qwen3.6-27B，除了coding，干其他也还行。</p>
]]></description><link>https://lcz.me/post/11789</link><guid isPermaLink="true">https://lcz.me/post/11789</guid><dc:creator><![CDATA[neo]]></dc:creator><pubDate>Sun, 09 Aug 2026 15:27:09 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 09 Aug 2026 12:49:51 GMT]]></title><description><![CDATA[<p dir="auto">我这里只有一台，deepseek 装上了，试了一下就停了，没法真的用。朋友那里有两台，在买线。过两天其他那里试试。</p>
]]></description><link>https://lcz.me/post/11785</link><guid isPermaLink="true">https://lcz.me/post/11785</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Sun, 09 Aug 2026 12:49:51 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 09 Aug 2026 09:14:19 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/soop-ladios" aria-label="Profile: soop-ladios">@<bdi>soop-ladios</bdi></a> 更进一步 2 台 DGX Spark 可以实现 1+1 &gt; 2 的效果，看到没人提就分享一下，成本很高大家量力而行吧。仓库链接： <a href="https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark" rel="nofollow ugc">https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark</a></p>
]]></description><link>https://lcz.me/post/11782</link><guid isPermaLink="true">https://lcz.me/post/11782</guid><dc:creator><![CDATA[linkdesu]]></dc:creator><pubDate>Sun, 09 Aug 2026 09:14:19 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 09 Aug 2026 03:51:14 GMT]]></title><description><![CDATA[<p dir="auto">GDX 部署好了 SGLang+27B 了。Mamba size 要自己显式设定等于你的并发数*5，否者 SGLang 会限制你的并发数。SGLang 在部署27B 的时候竟然认不全所有共享内存。我把--mem-fraction-static 1.4 才把看 kVcache撑起来。这个具体数值可能要自己根据情况看。 Hermes 出了 0.20.0 版，a2a和框架通讯都获益，不需要 cron，不需要 codex cli。Hermes 调用CC switch 配置27B的 codex 做点个人小项目泡一晚上都很稳。<a class="plugin-mentions-user plugin-mentions-a" href="/user/soop-ladios" aria-label="Profile: soop-ladios">@<bdi>soop-ladios</bdi></a></p>
]]></description><link>https://lcz.me/post/11761</link><guid isPermaLink="true">https://lcz.me/post/11761</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Sun, 09 Aug 2026 03:51:14 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 02 Aug 2026 08:34:14 GMT]]></title><description><![CDATA[<p dir="auto">這個模型我有用一陣子, 蠻不錯的, 不過有一些過度思考跟loop輸出的問題, 而且他們好像還在tune, 可能再等一陣子吧.</p>
]]></description><link>https://lcz.me/post/11204</link><guid isPermaLink="true">https://lcz.me/post/11204</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Sun, 02 Aug 2026 08:34:14 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 02 Aug 2026 06:48:34 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/8e227e37-ae51-4974-94c0-393028e2bff7.jpeg" alt="7c08e03d-9cf8-444b-8af4-2634b479101e-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">是这个laguna 2.1 吗? 好像他是 做的qwen 3.6 27b</p>
]]></description><link>https://lcz.me/post/11197</link><guid isPermaLink="true">https://lcz.me/post/11197</guid><dc:creator><![CDATA[mark]]></dc:creator><pubDate>Sun, 02 Aug 2026 06:48:34 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 02 Aug 2026 03:58:53 GMT]]></title><description><![CDATA[<p dir="auto">单纯解决问题Hermes 跑 DeepSeek 是最简单的。其实我大部分工作都是这么完成的。折腾的原因是因为现成有一台 DGX，不用起来觉得亏的慌。没上Qwen3.6:27b 的前面已经说了。空了去试试。现在用SGLang 跑  qwen3:coder-next-FP8感觉也还行。生命不息，折腾不止<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f604.png?v=138704eccfe" class="not-responsive emoji emoji-android emoji--smile" style="height:23px;width:auto;vertical-align:middle" title="😄" alt="😄" /></p>
]]></description><link>https://lcz.me/post/11185</link><guid isPermaLink="true">https://lcz.me/post/11185</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Sun, 02 Aug 2026 03:58:53 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 02 Aug 2026 01:23:41 GMT]]></title><description><![CDATA[<p dir="auto">Hermes直接接Deepseek就很好用.本地部署吃硬件，100G显存以内Qwen3.6 27b SGLang缺一不可，其他都是玩具。这些从几个月前就有大量测试佐证，没必要怀疑测试的人，自己一定要折腾其他方案。</p>
]]></description><link>https://lcz.me/post/11181</link><guid isPermaLink="true">https://lcz.me/post/11181</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sun, 02 Aug 2026 01:23:41 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 02 Aug 2026 01:17:08 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E5%8C%85%E7%A3%8A" aria-label="Profile: 包磊">@<bdi>包磊</bdi></a> 不錯喔，不過如果你之前沒試過，或許可以試試hermes直接接deepseek 看看，可能之前就是模型的問題而已<br />
這兩天DGX論壇上也把單台dgx spark裝deepseek flash 0731搞到看起來還不錯，可以試試. 若一台覺得精度不夠，可能要看用量決定是再買一台來張量並行，還是直接用線上的.<br />
兩台速度快很多，prefill 2000初 t/s, decode 40-60 tok/s, 精度也夠，看怎麼取捨了.</p>
]]></description><link>https://lcz.me/post/11180</link><guid isPermaLink="true">https://lcz.me/post/11180</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Sun, 02 Aug 2026 01:17:08 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Sun, 02 Aug 2026 00:50:22 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/soop-ladios" aria-label="Profile: soop-ladios">@<bdi>soop-ladios</bdi></a> 昨天装个 codex，然后和他讨论一番。决定在 codex 复制我在 Hermes 上的工作流。需求端保留在 Hermes。全文档交换系统。Hermes 用codex-run 方式启动。然后 简单cron扫描交换区进行双向通讯。阻塞性事件微信通知人肉处理。跑了一个自动抓雪球 文章的并整理评级的小项目效果很好。对多 agent看板工作流控制比我在 Hermes 手工攒起来强多了。后台用的 DeepSeek false 还没接 DGX，可能也是特别顺的原因之一。下次转到本地模型的再试</p>
]]></description><link>https://lcz.me/post/11179</link><guid isPermaLink="true">https://lcz.me/post/11179</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Sun, 02 Aug 2026 00:50:22 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 08:59:41 GMT]]></title><description><![CDATA[<p dir="auto">谢谢！ soop ladios。你说的跨框架的方案我还完全没试过，有机会去尝试一下。MTP我实测过Gemma4 31b，提升巨大。但到5对短任务，单toolcall就有点反作用了。</p>
]]></description><link>https://lcz.me/post/11073</link><guid isPermaLink="true">https://lcz.me/post/11073</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Fri, 31 Jul 2026 08:59:41 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 08:31:34 GMT]]></title><description><![CDATA[<p dir="auto">個人淺見, 如果是複雜冗長的任務, 可能還是用coding agent去跑會比較容易掌握, 可以讓hermes跟coding agent(codex cli/pi agent/claude code/kimi code..等等)協同作戰, 由hermes監控coding agent並由user手動或自動送命令過去.<br />
早期我沒解決會停下來的問題之前, 我是讓openclaw 定期去巡視工作中的tmux sessions, 有誰停下來了, 看起來又還沒完成工作的, 自動送出繼續+enter, 他就會繼續跑了. 也可以讓openclaw/hermes定期或手動讓他回報目前進度跟輸入指令.</p>
<p dir="auto">Deepseek單機我是沒跑過, 我是用雙機跑原版的. 原版權重160G, IQ2SS 應該也差不多80G. 我看論壇上心得是沒有太大損失, kv也足以支撐512K x 3 或 256K x 5. 不過他沒圖像, 剩餘的ram應該也不太夠去架多模態, 如果需要處理圖像可能就無法考慮了.</p>
<p dir="auto">體感上我覺得deepseek v4 flash比qwen 3.6 27B Q8強一點, 但是差距沒有到很大. qwen 3.6 27B我之所以跑Q8, 是因為這已經是我能接受的速度的極限了, 不然全精度還是放得下的. 後來有出了有含MTP頭的模型我就沒試過BF16了, 或許BF16的qwen 3.6 27B速度有所提升也不一定.</p>
<p dir="auto">不過, deepseek v4 flash正式版即將推出,或許一切都會讓人改觀.<br />
<img src="https://upload.lcz.me/uploads/5bfc933f-f63c-4c45-be48-c5d5d47aa55c.png" alt="螢幕擷取畫面 2026-07-31 163018.png" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/11070</link><guid isPermaLink="true">https://lcz.me/post/11070</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Fri, 31 Jul 2026 08:31:34 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 08:01:08 GMT]]></title><description><![CDATA[<p dir="auto">谢谢! soop ladios你的分享，我会去你的那里学习取经。我一直在用Hermes agent，也碰到好多问题，这两天参考Pi项目也改进了hermes的hook</p>
]]></description><link>https://lcz.me/post/11063</link><guid isPermaLink="true">https://lcz.me/post/11063</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Fri, 31 Jul 2026 08:01:08 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 07:46:52 GMT]]></title><description><![CDATA[<p dir="auto">qwen3.6 27B是有會因為tool call問題而中斷任務的可能, 事實上deepseek也是有toolcall leak 造成coding agent中途停下的問題.<br />
我的解決方法有兩個:</p>
<ol>
<li>使用claude code或codex cli的goal工具. 設定了/goal 之後, coding agent會督促模型沒有完成不能停下來. 兩者實現的方式不同, 各有一定效用. 其他如kimi code也是有這個命令,只是我沒用過.</li>
<li>我掛了一層proxy上去, 在模型跟code/claude code之間,專門去攔截有問題的訊息,把message修復, 讓整個工作可以繼續下去. 目前我在qwen 3.6 27B, deepseek v4 flash, glm 5.2都有遇到類似但不同的問題, 都是靠proxy修復. 我用的proxy有兩個, 一個是純proxy, <a href="https://github.com/ladiossoop5star/opencode_compat_proxy" rel="nofollow ugc">https://github.com/ladiossoop5star/opencode_compat_proxy</a> , 一個是因為後來我用litellm做模型的routing, 所以另外做了litellm的hook <a href="https://github.com/ladiossoop5star/litellm_coding_agent_hook" rel="nofollow ugc">https://github.com/ladiossoop5star/litellm_coding_agent_hook</a> . 可以參考看看.</li>
</ol>
]]></description><link>https://lcz.me/post/11061</link><guid isPermaLink="true">https://lcz.me/post/11061</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Fri, 31 Jul 2026 07:46:52 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 07:23:20 GMT]]></title><description><![CDATA[<p dir="auto">谢谢！xiaote 。昨天还写了个文档给架构agent要求开卡的body的颗粒度要细，每张卡片之间要加物理闸门，不通过不能自己结束。总之还在和agent们搏斗</p>
]]></description><link>https://lcz.me/post/11059</link><guid isPermaLink="true">https://lcz.me/post/11059</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Fri, 31 Jul 2026 07:23:20 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 07:19:49 GMT]]></title><description><![CDATA[<p dir="auto">我是有搞架构的agent做好方案，强调垂直切片，然后p0-pn，中间有coding+检测+架构自检+需求检查。然后有自动cron巡视处理阻塞。多profile协同</p>
]]></description><link>https://lcz.me/post/11058</link><guid isPermaLink="true">https://lcz.me/post/11058</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Fri, 31 Jul 2026 07:19:49 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 07:17:25 GMT]]></title><description><![CDATA[<p dir="auto">莽撞了，最近一直用qwen系列，kvcache很紧张。agent告诉我deepseek的MLA/DSA高压缩率对比qwen，kvcache可以少用一半显存。如果是真的16G系统+80G权重，还有30G不到，还可以搞一下试试</p>
]]></description><link>https://lcz.me/post/11057</link><guid isPermaLink="true">https://lcz.me/post/11057</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Fri, 31 Jul 2026 07:17:25 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 07:15:05 GMT]]></title><description><![CDATA[<p dir="auto">说个可能被忽略的角度：你换了一圈模型还在翻车，问题可能一半不在模型，而在长程任务的组织方式。Kanban 跑长程任务时"幻觉、偷懒、自己找台阶"这三种症状，最常见的原因是这几点：</p>
<ol>
<li>
<p dir="auto">任务粒度太大。一张卡片里塞"完成整个功能"，模型跑到后面连最初目标都忘了，只能开始编。正确做法是把大任务拆成一张卡一个可验证的小步骤，每张卡有明确的完成定义（比如"输出某文件、满足某条件"），做完立刻归档。</p>
</li>
<li>
<p dir="auto">上下文越滚越长，模型会"找台阶"。长任务跑到后面，历史对话占满上下文，模型为了收尾会假装完成。解决：每完成一个阶段就把该阶段的结论压缩成一段摘要（用 handoff 或让它更新项目笔记），把原始过程归档，保持工作上下文精简。</p>
</li>
<li>
<p dir="auto">温度太高是幻觉放大器。长任务建议把 temperature 调到 0.1-0.3，工具调用和代码输出会稳定很多——不少"幻觉"其实就是采样太随机。</p>
</li>
<li>
<p dir="auto">模型分工比单模型硬扛更有效。规划决策用你手上最强的模型（哪怕慢一点），执行简单步骤用快的模型。你提到 qwen3.6 27B 早期 toolcall 截断，那个问题在近几个版本的 SGLang/vLLM 上基本修掉了，值得配 MTP 再试一次，顺便 --kv-cache-dtype fp8 能省不少显存。</p>
</li>
</ol>
<p dir="auto">单台 DGX 跑不了大模型确实是硬件天花板，但 kanban 长任务的稳定性更多是编排问题，不是模型大小问题——小模型+好的任务拆解，往往比大模型硬跑一个超长任务更可靠。先把编排理顺，再决定要不要上更大的模型，可能更省心。</p>
]]></description><link>https://lcz.me/post/11056</link><guid isPermaLink="true">https://lcz.me/post/11056</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Fri, 31 Jul 2026 07:15:05 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 06:35:57 GMT]]></title><description><![CDATA[<p dir="auto">看了deepseek，这个大小，kvcache太难了。还不是多模态，其他都要放弃，可以玩，不适合长期用</p>
]]></description><link>https://lcz.me/post/11052</link><guid isPermaLink="true">https://lcz.me/post/11052</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Fri, 31 Jul 2026 06:35:57 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 06:19:12 GMT]]></title><description><![CDATA[<p dir="auto">谢谢，这个榜单好奇怪，都是高精度的小模型排前3。qwen3.6 27b最早开始就是试了，那个时候啥都不懂，老是toolcall截断，查了说是有个bug没解决就一直没再试。后来用了一阵35B-a3b，找时间再去试试。deepseek一直在用APIkey觉得还可以，一台dgx能跑的精度肯定肯定很小啊，能行吗？</p>
]]></description><link>https://lcz.me/post/11049</link><guid isPermaLink="true">https://lcz.me/post/11049</guid><dc:creator><![CDATA[包磊]]></dc:creator><pubDate>Fri, 31 Jul 2026 06:19:12 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 06:13:17 GMT]]></title><description><![CDATA[<p dir="auto">如果是一台的話, qwen 3.6 27B Q8或deepseek Q2是合適一點的選擇. 新出的laguna-s 2.1等他穩定一點可能也是一個不錯的選擇.</p>
<p dir="auto">deepseek單GB10可以參考:<br />
<a href="https://forums.developer.nvidia.com/t/1x-spark-tuned-dspark-for-deepseek-v4-flash-35-tok-s-800-prefill-and-fast-multi-agent-serving/376884" rel="nofollow ugc">https://forums.developer.nvidia.com/t/1x-spark-tuned-dspark-for-deepseek-v4-flash-35-tok-s-800-prefill-and-fast-multi-agent-serving/376884</a><br />
<a href="https://forums.developer.nvidia.com/t/optimizing-deepseek-v4-flash-on-a-single-nvidia-gb10-gx10-with-dspark-speculative-decoding/376830/22" rel="nofollow ugc">https://forums.developer.nvidia.com/t/optimizing-deepseek-v4-flash-on-a-single-nvidia-gb10-gx10-with-dspark-speculative-decoding/376830/22</a></p>
]]></description><link>https://lcz.me/post/11047</link><guid isPermaLink="true">https://lcz.me/post/11047</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Fri, 31 Jul 2026 06:13:17 GMT</pubDate></item><item><title><![CDATA[Reply to 一直没找到在DGX上能完全满意的一套模型配置 on Fri, 31 Jul 2026 06:07:06 GMT]]></title><description><![CDATA[<p dir="auto">GB10的话，可以参考这个网站：<a href="https://spark-arena.com/leaderboard" rel="nofollow ugc">https://spark-arena.com/leaderboard</a><br />
这也是产品SKU少的优势，能明确横向比较</p>
]]></description><link>https://lcz.me/post/11046</link><guid isPermaLink="true">https://lcz.me/post/11046</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Fri, 31 Jul 2026 06:07:06 GMT</pubDate></item></channel></rss>