<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/aae1b3fd-1562-45f9-a3c1-93f17d2ca8bf.jpg" alt="ff32550a-38d0-43bf-b3b0-d291a23ac6d1-3cabca928db0eccf8da689962fa786c7.jpg" class=" img-fluid img-markdown" /></p>
<p dir="auto">服务器太吵了，只开个qwen3.8+hermes也不适合7x24小时不停机，正好苹果发布新品，连夜做了一些功课之后还是押注绿厂了，搞了两台小主机，准备把干活模型升级到DSV4，说干就干。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/0bb1bd9c-6a2b-467b-aba4-c02938aba5ee.jpg" alt="db35b9ff-22e1-49b1-bcf2-0180a94f9014-c457f5a575e7c80adac777bbd1d105d6.jpg" class=" img-fluid img-markdown" /></p>
<p dir="auto">说全程远程部署着实是标题党，开机之后还是接了键鼠显示器，按照官网指南进行初始化配置连网更新，配置了固定ip就直接把外设包括显示器都拆了，用飞书遥控服务器的hermes远程配置小主机。<br />
得益于qwen3.8-27b着实可靠，过程非常顺利。其中遇到了小主机访问不了外网，自己配置服务器的clash代理给两台小主机更新docker，下载DSV4配方包，再解压配置两台之间的ssh密钥，测试qsfp网速，然后就直接把DSV4开起来了，全程没让我插手。体验简直惊艳。</p>
<p dir="auto">file:///home/bentonyi/文档/xwechat_files/bentonez_e4c6/temp/RWTemp/2026-09/09fef7517ddb4143df3ea7e81576d5f9.png<br />
<img src="https://upload.lcz.me/uploads/4192657f-1161-430b-8087-66c8350a87a5.jpeg" alt="3de2e177-4cf2-45ca-964b-b5b30aed3783-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">同时用DSH和Hermes给他下任务同时跑，性能也是非常不错：</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/d47e7c4d-ef18-4dc1-9f3d-3e5446353345.jpeg" alt="7c5f4aa8-9248-411c-a36e-123230400bca-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">温度和功率倒也不是很高，以后高负载的时候我再更新一个<br />
file:///home/bentonyi/文档/xwechat_files/bentonez_e4c6/temp/RWTemp/2026-09/f5b567aaf4f918ed727170f8aa1e8fac.png<br />
<img src="https://upload.lcz.me/uploads/507ec3b8-f218-4910-9b96-aab720ce2981.jpeg" alt="1da955d4-7083-4139-824e-918111fa4c23-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/topic/1457</link><generator>RSS for Node</generator><lastBuildDate>Mon, 07 Sep 2026 20:35:29 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1457.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 01 Sep 2026 13:52:55 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Mon, 07 Sep 2026 07:48:47 GMT]]></title><description><![CDATA[<p dir="auto">为何总参数、激活参数更小的qwen3.8-flash会更慢，这个很难理解。</p>
]]></description><link>https://lcz.me/post/16378</link><guid isPermaLink="true">https://lcz.me/post/16378</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Mon, 07 Sep 2026 07:48:47 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Mon, 07 Sep 2026 07:37:07 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/benton-yi" aria-label="Profile: benton-yi">@<bdi>benton-yi</bdi></a> 我弟你多测试下，comfyui/glm/qwen flash/dsv4，多发点帖子<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f602.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--joy" style="height:23px;width:auto;vertical-align:middle" title="😂" alt="😂" /></p>
]]></description><link>https://lcz.me/post/16377</link><guid isPermaLink="true">https://lcz.me/post/16377</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Mon, 07 Sep 2026 07:37:07 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Mon, 07 Sep 2026 06:27:01 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/3d561ae0-2890-4d69-af4f-5c350233448b.jpeg" alt="a6885b23-ef0e-472e-9cb1-91720ad40a3f-image.jpeg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/e866283c-9486-4c4c-a956-6b2624a88111.jpeg" alt="3efd99b2-bb79-4af6-9fe2-7d8da857c538-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/16356</link><guid isPermaLink="true">https://lcz.me/post/16356</guid><dc:creator><![CDATA[benton yi]]></dc:creator><pubDate>Mon, 07 Sep 2026 06:27:01 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Mon, 07 Sep 2026 06:17:05 GMT]]></title><description><![CDATA[<p dir="auto">Qwen4架构的qwen3.8-flash-next 测试结果有了么？</p>
]]></description><link>https://lcz.me/post/16355</link><guid isPermaLink="true">https://lcz.me/post/16355</guid><dc:creator><![CDATA[Grayson Ren]]></dc:creator><pubDate>Mon, 07 Sep 2026 06:17:05 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Fri, 04 Sep 2026 12:59:36 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/94c6b2d7-3700-4fbf-8c45-539bd8740d63.jpeg" alt="d46e9494-ceb9-4cb6-b22e-45f1ffa71376-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">以上截图是今天真实工作场景下的全量日志拿来分析的数据，PP的峰值数据不可信（因为SGLang推理框架的RadixAttention缓存真的优雅又高效），但均值数据确实比vLLM运行DSV4-Flash提高了能有接近4成。</p>
<p dir="auto">启动参数脚本如下，先在从机运行启动脚本：</p>
<pre><code class="language-docker">docker run -d --name sglang_glm53 --network host --gpus all --shm-size=32g \
  -v /home/bentonyi/LLM/LibertAIDAI/GLM-5.3-Flash-NVFP4:/models/glm53:ro \
  sglang-glm53:gb10 \
  python3 -m sglang.launch_server \
    --model-path /models/glm53 \
    --trust-remote-code \
    --tp-size 2 --nnodes 2 --node-rank 1 \
    --dist-init-addr 10.100.40.2:29500 \
    --attention-backend dsa \
    --dsa-prefill-backend tilelang --dsa-decode-backend tilelang \
    --moe-runner-backend flashinfer_cutlass \
    --kv-cache-dtype bfloat16 \
    --disable-shared-experts-fusion \
    --reasoning-parser glm45 --tool-call-parser glm47 \
    --mem-fraction-static 0.90 \
    --context-length 524288 \
    --max-running-requests 2 \
    --default-chat-template-kwargs '{"reasoning_effort": "low"}' \
    --host 0.0.0.0 --port 8000
</code></pre>
<p dir="auto">主机等6秒运行脚本：</p>
<pre><code class="language-sleep">docker rm -f sglang_glm53 2&gt;/dev/null
docker run -d --name sglang_glm53 --network host --gpus all --shm-size=32g \
  -v /home/bentonyi/LLM/LibertAIDAI/GLM-5.3-Flash-NVFP4:/models/glm53:ro \
  sglang-glm53:gb10 \
  python3 -m sglang.launch_server \
    --model-path /models/glm53 \
    --trust-remote-code \
    --tp-size 2 --nnodes 2 --node-rank 0 \
    --dist-init-addr 10.100.40.2:29500 \
    --attention-backend dsa \
    --dsa-prefill-backend tilelang --dsa-decode-backend tilelang \
    --moe-runner-backend flashinfer_cutlass \
    --kv-cache-dtype bfloat16 \
    --disable-shared-experts-fusion \
    --reasoning-parser glm45 --tool-call-parser glm47 \
    --mem-fraction-static 0.90 \
    --context-length 524288 \
    --max-running-requests 2 \
    --default-chat-template-kwargs '{"reasoning_effort": "low"}' \
    --host 0.0.0.0 --port 8000
</code></pre>
<p dir="auto">这个参数设置得比较极限：通过计算得到内存红线设置在0.92可保证上下文512K的情况下最大并发为2（gb10集群仅作推理服务器不运行别的任何程序，连屏幕键鼠都不配的情况下）。<br />
GLM5.3-Flash这模型因为激活参数18B，推理确实很慢，再加上这货的默认设置导致过度思考的问题：<br />
<img src="https://upload.lcz.me/uploads/c54e0dd6-ee8e-4211-a00b-b1f2ef60504f.jpeg" alt="c96873da-5267-4322-bfe5-4bb377b61c98-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">加入<code>    --default-chat-template-kwargs '{"reasoning_effort": "low"}' \</code>后勉强可用，但效果还是远不如DSV4-Flash。也就是不愿意为“降速高达70%但智力没有相应程度的提升”买单，工作了一天准备换回DSV4-Flash。下一步准备测试Qwen4架构的qwen3.8-flash-next。因为都是以实际工作场景的数据日志进行性能分析，所以更新可能没那么快，主打一个真实可信。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/7d6313b9-f717-467d-9507-4c7c8ded615b.jpeg" alt="fcfb0251-2776-4789-9aa1-a13eb51cb09d-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">最后引用一个社区排行榜的glm5.3-flash截图，tp2无论如何都还是捉襟见肘，要真想在gb10上爽玩还是得TP4集群。</p>
]]></description><link>https://lcz.me/post/15823</link><guid isPermaLink="true">https://lcz.me/post/15823</guid><dc:creator><![CDATA[benton yi]]></dc:creator><pubDate>Fri, 04 Sep 2026 12:59:36 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Thu, 03 Sep 2026 01:10:26 GMT]]></title><description><![CDATA[<p dir="auto">我早就说过两台GB10 跑DS V4是目前最好的本地方案了。几乎本地95%的工作不要在线了。</p>
]]></description><link>https://lcz.me/post/15561</link><guid isPermaLink="true">https://lcz.me/post/15561</guid><dc:creator><![CDATA[iamvirus]]></dc:creator><pubDate>Thu, 03 Sep 2026 01:10:26 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Thu, 03 Sep 2026 00:34:08 GMT]]></title><description><![CDATA[<p dir="auto">這是昨天請 hermes (gemini-3.7-flash-high) 在雙GX10跑的測試做出來的比較，沒有很嚴謹(因為題目是它幫我想的XD)，大家看看就好~<br />
<img src="https://upload.lcz.me/uploads/42916ff1-de3f-4ab0-b5c4-e32568ef4ebe.png" alt="dual-gx10-three-kings-benchmark.png" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/15553</link><guid isPermaLink="true">https://lcz.me/post/15553</guid><dc:creator><![CDATA[densha]]></dc:creator><pubDate>Thu, 03 Sep 2026 00:34:08 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Thu, 03 Sep 2026 03:32:08 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/15507">said</a>:</p>
<p dir="auto">半年前还没有Qwen3.5</p>
</blockquote>
<p dir="auto">在Hermes 四月份出現之前, 我大概頂多只是看新聞, 看了兩個月的OpenClaw熱潮,<br />
逛帖子也頂多上一些Reddit 論壇 看大家聊AI未來10~20年的發展<br />
沒想到半年內就出了 Qwen3.5, Qwen3.6, Qwen3.8</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/af21c1a4-46b8-4cab-8c47-848af633d489.jpeg" alt="3094e184-de26-4937-bbdb-220a28070809-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/15529</link><guid isPermaLink="true">https://lcz.me/post/15529</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Thu, 03 Sep 2026 03:32:08 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 12:50:34 GMT]]></title><description><![CDATA[<p dir="auto">B站有个哥们，现在已经有6台DGX了，他自己实测发现，两台dxg 和 deepseek  v4 flash  是绝配。</p>
<p dir="auto">相比下M3 ultra 脱离了聊天就直接拉了，算力不够，空有高带宽 。</p>
]]></description><link>https://lcz.me/post/15515</link><guid isPermaLink="true">https://lcz.me/post/15515</guid><dc:creator><![CDATA[张光璞]]></dc:creator><pubDate>Wed, 02 Sep 2026 12:50:34 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 12:05:09 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> 半年前还没有Qwen3.5，我刚做这个频道的时候，能跑Agent的模型我感觉只有Claude opus 4.6，GPT都不够格。现在是Qwen3.8 27b比肩Opus 4.6的时代。</p>
]]></description><link>https://lcz.me/post/15507</link><guid isPermaLink="true">https://lcz.me/post/15507</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 02 Sep 2026 12:05:09 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 11:19:52 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> <a href="/post/15417">said</a>:</p>
<p dir="auto">6~7万，就能部署一个头部能力的全尺寸模型。响应能力、kvcache容量都是可用级别。<br />
放在半年前都像做梦一样。</p>
</blockquote>
<p dir="auto">請問半年前是什麼場景呢？ 我是Qwen3.6 才開始研究Local LLM的</p>
]]></description><link>https://lcz.me/post/15501</link><guid isPermaLink="true">https://lcz.me/post/15501</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Wed, 02 Sep 2026 11:19:52 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 07:35:20 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a><br />
我是沒用兩台跑過qwen 3.8 flash next, 機器都拿去跑別的模型了. 照以往的經驗, 兩台大概可以提昇1.3~1.5倍左右的速度.<br />
這邊有一台用SGLang跑出43 tok/s平均速度的可以參考 <a href="https://github.com/azampatti/GB10-3.8-Flash-Next" rel="nofollow ugc">https://github.com/azampatti/GB10-3.8-Flash-Next</a></p>
<p dir="auto">也有人兩台跑vllm的nvfp4可以參考:<br />
<img src="https://upload.lcz.me/uploads/6fd16af4-2406-42ac-8ef6-b6bc0579cd6b.png" alt="截圖 2026-09-02 下午3.30.20.png" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/15442</link><guid isPermaLink="true">https://lcz.me/post/15442</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Wed, 02 Sep 2026 07:35:20 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 07:03:56 GMT]]></title><description><![CDATA[<p dir="auto">论坛非常需要的一手数据，虽然是大型炫富现场，但是可以再多发点实测数据。这个两台能跑V4 Flash确实有点离谱，有SGLang的加持，Radix缓存树和DSpark都可以缓解带宽压力。</p>
]]></description><link>https://lcz.me/post/15431</link><guid isPermaLink="true">https://lcz.me/post/15431</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 02 Sep 2026 07:03:56 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 06:54:03 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/soop-ladios" aria-label="Profile: soop-ladios">@<bdi>soop-ladios</bdi></a> 我考虑是双台是不是可以有更好的prefill和decode表现，从而让双机GB10的性价比再提高一个级别。</p>
<p dir="auto">毕竟30的tg/s还是稍逊了一点。</p>
]]></description><link>https://lcz.me/post/15430</link><guid isPermaLink="true">https://lcz.me/post/15430</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Wed, 02 Sep 2026 06:54:03 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 06:46:59 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a><br />
qwen 3.8 flash next NVFP4一台就可以跑了, 參考 <a href="https://github.com/blazux/qwen3.8-Flash-DGX" rel="nofollow ugc">https://github.com/blazux/qwen3.8-Flash-DGX</a><br />
pp大概1500-2000 tok/s , decode 約 30 tok/s , KV cache大概有630K. 我跑了一周左右吧, 很穩定, 用codex cli長鍊loop工作數天沒什麼問題.<br />
兩台其實也可以跑glm 5.3 flash的 <a href="https://github.com/Entrpi/glm-5.3-flash-exl3-2x-spark" rel="nofollow ugc">https://github.com/Entrpi/glm-5.3-flash-exl3-2x-spark</a></p>
]]></description><link>https://lcz.me/post/15427</link><guid isPermaLink="true">https://lcz.me/post/15427</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Wed, 02 Sep 2026 06:46:59 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 05:23:59 GMT]]></title><description><![CDATA[<p dir="auto">pp1500~2500，tg60~70还不错。</p>
<p dir="auto">6~7万，就能部署一个头部能力的全尺寸模型。响应能力、kvcache容量都是可用级别。<br />
放在半年前都像做梦一样。</p>
<p dir="auto">主要是双GB10的RAM比较富裕，多session并行切换也不会频繁重复prefill。<br />
所以prefill稍慢也不是不行。</p>
<p dir="auto">而且据测试，GB10的prefill性能随context膨胀的下降趋势不明显。</p>
]]></description><link>https://lcz.me/post/15417</link><guid isPermaLink="true">https://lcz.me/post/15417</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Wed, 02 Sep 2026 05:23:59 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 05:10:59 GMT]]></title><description><![CDATA[<p dir="auto">这个体验已经不错了, 完全可用</p>
]]></description><link>https://lcz.me/post/15416</link><guid isPermaLink="true">https://lcz.me/post/15416</guid><dc:creator><![CDATA[Tony Wang]]></dc:creator><pubDate>Wed, 02 Sep 2026 05:10:59 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 06:56:36 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/tony-wang" aria-label="Profile: tony-wang">@<bdi>tony-wang</bdi></a> 模型就是Deepseek-ai发布的原生FP8+NVFP4量化模型，模型卡<a href="https://modelscope.cn/models/deepseek-ai/DeepSeek-V4-Flash-0731" rel="nofollow ugc">https://modelscope.cn/models/deepseek-ai/DeepSeek-V4-Flash-0731</a></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/1d96cdab-bdaf-44f6-bd52-6c6bc567d631.jpeg" alt="91d7f9e0-d2cd-4788-9547-79f1f6e41e99-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">这是在用hermes一边计划对齐部署glm5.3-flash方案，一边在用DSH做早课的2并发时候的实际场景decode截图，最高单会话能上到76.7t/s，最低也有35+保底。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/c32f099b-c4e9-40f1-8e6f-22fe0929bad5.jpeg" alt="5378a5b8-532d-453e-8a87-801b3c860beb-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">这是让hermes读魔搭的glm5.3-flash几个不同版本的模型卡页面，pp大概在1600～2500之间浮动。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/97351764-5f87-4d7d-9faf-8ca3e8e46b8a.jpeg" alt="db89651b-6ae6-4930-aecd-b11642130be5-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">上图是工作界面，之后会把glm5.3-flash的实际使用评测也发进来。</p>
]]></description><link>https://lcz.me/post/15414</link><guid isPermaLink="true">https://lcz.me/post/15414</guid><dc:creator><![CDATA[benton yi]]></dc:creator><pubDate>Wed, 02 Sep 2026 06:56:36 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 04:04:40 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> 双 GX10 跑 Qwen3.8-Flash-Next，pp 和 tg 要分开看，结论先说：<strong>双机不会让单请求的 decode 变快</strong>。</p>
<p dir="auto">关键在 MoE 的 decode 是<strong>带宽瓶颈</strong>不是算力瓶颈：每生成一个 token，要从内存读一遍"活跃参数"。GX10（和 DGX Spark 同源）内存带宽约 273GB/s（LPDDR5X）。</p>
<ul>
<li>Qwen3.8-Flash-Next 约 176B 总参、MoE 每 token 只激活约 18B，NVFP4 量化后整模型约 90-100GB，<strong>单台 128GB 的 GX10 就装得下</strong>；</li>
<li>DS-V4-Flash 约 284B/13B 活跃，官方 FP4 权重约 168GB，<strong>单台装不下，这才是双机的理由</strong>（容量）；</li>
<li>decode 速度上限 ≈ 活跃参数字节数 ÷ 带宽。Flash-Next NVFP4 每 token 读约 10GB → 单机理论 ~25 t/s 上限、实际 15-20；双机时每个请求仍只跑一台，带宽不会叠加。</li>
</ul>
<p dir="auto">双机真正的收益只有三样：</p>
<ol>
<li><strong>容量</strong>：装 168GB 级的 DS-V4-Flash（楼主两台跑 DSV4 就是这个原因）；</li>
<li><strong>并发</strong>：两个请求各占一台，总吞吐翻倍；</li>
<li><strong>prefill</strong>：吃算力，若做跨机并行确实更快。</li>
</ol>
<p dir="auto">别指望跨机 tensor parallel 提速 decode：MoE 的 TP 每个 token 都要 all-to-all 交换 expert 输出，跨机走以太网就是新瓶颈（坛里 RDNA TP 无 P2P 反负优化的实测 TID:1438 是同一个道理）。</p>
<p dir="auto">另外 AA 榜单的能力排序是云端/API 配置下的数字，和本地量化部署的速度是两回事。按你现在的两台机器：Flash-Next 单台就能跑、DS-V4-Flash 必须双台，这个分工比榜单排位更影响实际使用。</p>
]]></description><link>https://lcz.me/post/15410</link><guid isPermaLink="true">https://lcz.me/post/15410</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Wed, 02 Sep 2026 04:04:40 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 03:03:31 GMT]]></title><description><![CDATA[<p dir="auto">我有个想法，从目前的 <a href="https://artificialanalysis.ai/models" rel="nofollow ugc">https://artificialanalysis.ai/models</a> 的跑分榜单来看，其实Qwen3.8-flash-next的能力是高于deepseek-v4-flash-0731的。<br />
<img src="https://upload.lcz.me/uploads/d62ed326-88fb-4e9b-8781-a81d0c10e432.jpeg" alt="db178f2d-896e-414c-911d-583681317773-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">所以是不是双GB10跑Qwen3.8-flash-next能获得更好的pp和tg的性能表现。</p>
]]></description><link>https://lcz.me/post/15404</link><guid isPermaLink="true">https://lcz.me/post/15404</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Wed, 02 Sep 2026 03:03:31 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 01:52:21 GMT]]></title><description><![CDATA[<p dir="auto">我有八月初裝的時候的測試數據可以參考:<br />
<img src="https://upload.lcz.me/uploads/7424f992-b14b-446e-aedc-3e740a2669cf.png" alt="截圖 2026-09-02 上午9.50.57.png" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/15401</link><guid isPermaLink="true">https://lcz.me/post/15401</guid><dc:creator><![CDATA[soop ladios]]></dc:creator><pubDate>Wed, 02 Sep 2026 01:52:21 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 01:28:05 GMT]]></title><description><![CDATA[<p dir="auto">好厉害 同问测试数据</p>
]]></description><link>https://lcz.me/post/15400</link><guid isPermaLink="true">https://lcz.me/post/15400</guid><dc:creator><![CDATA[zhenyu huang]]></dc:creator><pubDate>Wed, 02 Sep 2026 01:28:05 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Wed, 02 Sep 2026 01:03:31 GMT]]></title><description><![CDATA[<p dir="auto">我喜歡第一張照片的意境 傳統和現代文明的相襯 真的很奇妙</p>
]]></description><link>https://lcz.me/post/15395</link><guid isPermaLink="true">https://lcz.me/post/15395</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Wed, 02 Sep 2026 01:03:31 GMT</pubDate></item><item><title><![CDATA[Reply to Asus Ascent GX10到货，双gb10的DSV4-Flash集群全程hermes远程部署 on Tue, 01 Sep 2026 23:18:50 GMT]]></title><description><![CDATA[<p dir="auto">留言希望看測試結果</p>
]]></description><link>https://lcz.me/post/15388</link><guid isPermaLink="true">https://lcz.me/post/15388</guid><dc:creator><![CDATA[KAKAHermes]]></dc:creator><pubDate>Tue, 01 Sep 2026 23:18:50 GMT</pubDate></item></channel></rss>