<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[准备好迎接Qwen3.8-Flash-Next 125B A6B]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/1477c4e5-5520-4837-be4e-5962704db63e.jpeg" alt="4d97ff5f-db32-475e-8a89-ad0fe9de3801-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/ec72c4a8-3ab4-40d1-88d1-05566ca4edfd.jpeg" alt="1bc82737-f228-4a64-b5cd-20e9854ffce3-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><a href="https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next" rel="nofollow ugc">https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next</a></p>
<p dir="auto">这会是DGX Spark的梦中情模吗？</p>
]]></description><link>https://lcz.me/topic/1324</link><generator>RSS for Node</generator><lastBuildDate>Tue, 08 Sep 2026 01:49:29 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1324.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 25 Aug 2026 17:18:42 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Thu, 27 Aug 2026 02:49:53 GMT]]></title><description><![CDATA[<p dir="auto"><a href="https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-GGUF" rel="nofollow ugc">https://huggingface.co/Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-GGUF</a><br />
這裡有大神在測試Gdx Spark，打算拿板凳等他。</p>
]]></description><link>https://lcz.me/post/14240</link><guid isPermaLink="true">https://lcz.me/post/14240</guid><dc:creator><![CDATA[Francis Cheng]]></dc:creator><pubDate>Thu, 27 Aug 2026 02:49:53 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 15:29:03 GMT]]></title><description><![CDATA[<p dir="auto">我剛剛丟給AI了....她生成好 llama啟動指令，我明早起來驗收</p>
]]></description><link>https://lcz.me/post/14163</link><guid isPermaLink="true">https://lcz.me/post/14163</guid><dc:creator><![CDATA[David Chen]]></dc:creator><pubDate>Wed, 26 Aug 2026 15:29:03 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 13:54:17 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bunsei" aria-label="Profile: Bunsei">@<bdi>Bunsei</bdi></a> <a href="/post/14133">说</a>:</p>
<p dir="auto">本来以为128G的模型，结果发现是一个快到200G的...只能等老哥们的实测效果了。</p>
</blockquote>
<p dir="auto">毕竟这么大的模型开放权重，我倒是觉得用不起不是别人的问题，是我的问题</p>
]]></description><link>https://lcz.me/post/14147</link><guid isPermaLink="true">https://lcz.me/post/14147</guid><dc:creator><![CDATA[vosrock]]></dc:creator><pubDate>Wed, 26 Aug 2026 13:54:17 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 13:30:13 GMT]]></title><description><![CDATA[<p dir="auto">下载中<br />
<img src="https://upload.lcz.me/uploads/cf3b23b8-2617-4648-8b81-004a279ed4f1.jpeg" alt="d6ae9ca1-7b30-43c9-859a-7593299d14ee-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/14144</link><guid isPermaLink="true">https://lcz.me/post/14144</guid><dc:creator><![CDATA[Arroyo Cheung]]></dc:creator><pubDate>Wed, 26 Aug 2026 13:30:13 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 13:29:17 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/applejuice" aria-label="Profile: applejuice">@<bdi>applejuice</bdi></a> <a href="/post/14135">说</a>:</p>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bunsei" aria-label="Profile: Bunsei">@<bdi>Bunsei</bdi></a> max5 ultra 你值得拥有</p>
</blockquote>
<p dir="auto">我在等第一波测评出来。 如果效果好的话，我打算换个平台。不用AM5了...</p>
]]></description><link>https://lcz.me/post/14143</link><guid isPermaLink="true">https://lcz.me/post/14143</guid><dc:creator><![CDATA[Bunsei]]></dc:creator><pubDate>Wed, 26 Aug 2026 13:29:17 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 13:02:18 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bunsei" aria-label="Profile: Bunsei">@<bdi>Bunsei</bdi></a> max5 ultra 你值得拥有</p>
]]></description><link>https://lcz.me/post/14135</link><guid isPermaLink="true">https://lcz.me/post/14135</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Wed, 26 Aug 2026 13:02:18 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 12:58:23 GMT]]></title><description><![CDATA[<p dir="auto">本来以为128G的模型，结果发现是一个快到200G的...只能等老哥们的实测效果了。</p>
]]></description><link>https://lcz.me/post/14133</link><guid isPermaLink="true">https://lcz.me/post/14133</guid><dc:creator><![CDATA[Bunsei]]></dc:creator><pubDate>Wed, 26 Aug 2026 12:58:23 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 11:01:57 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stxpnet" aria-label="Profile: stxpnet">@<bdi>stxpnet</bdi></a> <a href="/post/13977">说</a>:</p>
<p dir="auto">对啊，就看它优化得如何了。 我一直的设想是，这种MOE是文科生，回答问题速度快，多轮下来，直击一个问题的各个方面，然后最后进行一个大总结，速度快，角度全。应该和同等价值硬件能跑稠密模型有得一拼，然而事实却总是打脸。 我最后还是会选择稠密，宁愿久一点，都不要错一点。</p>
</blockquote>
<p dir="auto">应该是激活参数的大小决定了能力上限了，A6B我倒是觉得小了点，不过等出来看看吧，之前尝试过QWEN3 CODER NEXT，80BA3B跑内存也能到1000PR30TS，这个速度实际上也可用了，这个125B估计又要Q4量化才玩得动了</p>
]]></description><link>https://lcz.me/post/14121</link><guid isPermaLink="true">https://lcz.me/post/14121</guid><dc:creator><![CDATA[vosrock]]></dc:creator><pubDate>Wed, 26 Aug 2026 11:01:57 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 10:42:12 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stxpnet" aria-label="Profile: stxpnet">@<bdi>stxpnet</bdi></a> 395不像dgx spark有200g，不过听说有人用oculink外接rdma网卡实现互联的</p>
]]></description><link>https://lcz.me/post/14119</link><guid isPermaLink="true">https://lcz.me/post/14119</guid><dc:creator><![CDATA[GitNo2]]></dc:creator><pubDate>Wed, 26 Aug 2026 10:42:12 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 10:39:11 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 用的是oclulink连接，3090占用23.5g，其余的扔到395上</p>
]]></description><link>https://lcz.me/post/14118</link><guid isPermaLink="true">https://lcz.me/post/14118</guid><dc:creator><![CDATA[GitNo2]]></dc:creator><pubDate>Wed, 26 Aug 2026 10:39:11 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 09:13:30 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stxpnet" aria-label="Profile: stxpnet">@<bdi>stxpnet</bdi></a> 这两张卡连起来不是很正常吗？3090显卡坞，雷电4/5，Oculink都行，系统lInux，驱动Vulkan，很多框架可以直接认出两张卡，分层计算。</p>
<p dir="auto">况且他的帖子，根本没说协同工作，或许它就是分开跑的.....</p>
]]></description><link>https://lcz.me/post/14096</link><guid isPermaLink="true">https://lcz.me/post/14096</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 26 Aug 2026 09:13:30 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 09:11:32 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/gitno2" aria-label="Profile: GitNo2">@<bdi>GitNo2</bdi></a> 395和3090是怎样连接的呢？ 有点好奇。  不在同一个机器的两张显卡，有没有可能通过超高速光纤网络进行互联呢？</p>
]]></description><link>https://lcz.me/post/14095</link><guid isPermaLink="true">https://lcz.me/post/14095</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Wed, 26 Aug 2026 09:11:32 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 08:29:35 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/tony-wang" aria-label="Profile: Tony-Wang">@<bdi>Tony-Wang</bdi></a> 现在的主力是qwen3.8 27b-udq8xl，用了395+3090，等过几天llama.cpp支持新模型，我就试试双机和单机+3090两种方案</p>
]]></description><link>https://lcz.me/post/14086</link><guid isPermaLink="true">https://lcz.me/post/14086</guid><dc:creator><![CDATA[GitNo2]]></dc:creator><pubDate>Wed, 26 Aug 2026 08:29:35 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 08:25:28 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/gitno2" aria-label="Profile: GitNo2">@<bdi>GitNo2</bdi></a></p>
<p dir="auto">到时候分享一下，Q4也行啊，参数大有参数大的优势，如果一台395能跑Q4，那性价比是不错的</p>
]]></description><link>https://lcz.me/post/14084</link><guid isPermaLink="true">https://lcz.me/post/14084</guid><dc:creator><![CDATA[Tony Wang]]></dc:creator><pubDate>Wed, 26 Aug 2026 08:25:28 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 07:58:06 GMT]]></title><description><![CDATA[<p dir="auto">不知道我的双机aimax395跑起来会怎么样，感觉可以开q8</p>
]]></description><link>https://lcz.me/post/14080</link><guid isPermaLink="true">https://lcz.me/post/14080</guid><dc:creator><![CDATA[GitNo2]]></dc:creator><pubDate>Wed, 26 Aug 2026 07:58:06 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 00:43:06 GMT]]></title><description><![CDATA[<p dir="auto">8万5<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f602.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--joy" style="height:23px;width:auto;vertical-align:middle" title="😂" alt="😂" /></p>
]]></description><link>https://lcz.me/post/13982</link><guid isPermaLink="true">https://lcz.me/post/13982</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 26 Aug 2026 00:43:06 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 00:15:03 GMT]]></title><description><![CDATA[<p dir="auto">正好apple发布 M5 ultra了, 96G或256G 玩这个正合适.</p>
]]></description><link>https://lcz.me/post/13979</link><guid isPermaLink="true">https://lcz.me/post/13979</guid><dc:creator><![CDATA[Tony Wang]]></dc:creator><pubDate>Wed, 26 Aug 2026 00:15:03 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 00:13:42 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/quanta-magic" aria-label="Profile: Quanta-Magic">@<bdi>Quanta-Magic</bdi></a> 16G显存，64G空余DDR5内存应该可以跑得很欢了。</p>
]]></description><link>https://lcz.me/post/13978</link><guid isPermaLink="true">https://lcz.me/post/13978</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Wed, 26 Aug 2026 00:13:42 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Wed, 26 Aug 2026 00:12:50 GMT]]></title><description><![CDATA[<p dir="auto">对啊，就看它优化得如何了。 我一直的设想是，这种MOE是文科生，回答问题速度快，多轮下来，直击一个问题的各个方面，然后最后进行一个大总结，速度快，角度全。应该和同等价值硬件能跑稠密模型有得一拼，然而事实却总是打脸。 我最后还是会选择稠密，宁愿久一点，都不要错一点。</p>
]]></description><link>https://lcz.me/post/13977</link><guid isPermaLink="true">https://lcz.me/post/13977</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Wed, 26 Aug 2026 00:12:50 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Tue, 25 Aug 2026 22:06:38 GMT]]></title><description><![CDATA[<p dir="auto">哈哈确实，A6B 对显存小的场景太友好了——每 token 只激活 6B 参数，没激活的 expert 放内存按需加载，小主机 + 大内存就能玩 Q4 offload，慢但能跑。等官方 GGUF 发布，32G 内存的迷你主机都能试一把。要流畅还是 ~70G 显存全装（双 R9700 64G 差 6G 挺可惜，单张 96G 一步到位）。</p>
]]></description><link>https://lcz.me/post/13965</link><guid isPermaLink="true">https://lcz.me/post/13965</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 25 Aug 2026 22:06:38 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Tue, 25 Aug 2026 19:46:00 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 我靠，这个是小主机的春天，A6B绝对是值得玩的。</p>
]]></description><link>https://lcz.me/post/13952</link><guid isPermaLink="true">https://lcz.me/post/13952</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 25 Aug 2026 19:46:00 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Tue, 25 Aug 2026 19:11:42 GMT]]></title><description><![CDATA[<p dir="auto">按 125B 总参数量粗算权重体积（MoE 也一样按总参数量算）：</p>
<ul>
<li>FP16/BF16：~250GB</li>
<li>FP8：~125GB</li>
<li>Q6_K：~94GB</li>
<li>Q4_K_M：~70GB</li>
<li>IQ3 / IQ2：~55GB / ~40GB</li>
</ul>
<p dir="auto">所以"全量进显存"的门槛：Q4 也要 ~70GB——双 R9700（64GB）差一点，得单张 96GB（PRO 6000 这类）或 3×32GB 才行。</p>
<p dir="auto">但 A6B 真正的价值是 <strong>offload 友好</strong>：每 token 只激活 6B 参数，没激活的 expert 可以放内存按需加载。24GB 单卡 + 64GB 内存跑 Q4 完全能跑（比同体积 dense 125B 实用得多），代价是速度——激活 expert 的权重得不断从内存搬，t/s 会比较惨，但能用。想流畅还是得 ~70GB 显存全装。</p>
<p dir="auto">等有人出 GGUF 实测再报具体数，先按这个量级规划显存没错。</p>
]]></description><link>https://lcz.me/post/13950</link><guid isPermaLink="true">https://lcz.me/post/13950</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 25 Aug 2026 19:11:42 GMT</pubDate></item><item><title><![CDATA[Reply to 准备好迎接Qwen3.8-Flash-Next 125B A6B on Tue, 25 Aug 2026 19:04:24 GMT]]></title><description><![CDATA[<p dir="auto">这要多大的显存才能用？</p>
]]></description><link>https://lcz.me/post/13946</link><guid isPermaLink="true">https://lcz.me/post/13946</guid><dc:creator><![CDATA[Quanta Magic]]></dc:creator><pubDate>Tue, 25 Aug 2026 19:04:24 GMT</pubDate></item></channel></rss>