<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Qwen3.8-Flash-Next M5 Max实测]]></title><description><![CDATA[<p dir="auto">直接说重点，好，但是需要优化。</p>
<p dir="auto">因为明天要出差，就简短一点，给大家一点启发。</p>
<p dir="auto">ngram 50b 左右的部分可以放到ssd里直接读，不需要进内存，实际上就是一个120b a6b的模型，4 bit也就50多gb，还只要a6b，潜力巨大。</p>
<p dir="auto">M5 MAX 实测结果：<br />
128k上下文 没有额外配置</p>
<p dir="auto">最快30 tokens，最慢降速到十几tokens。均值26.2<br />
这部分感觉优化空间很大，理论上ngram应该属于秒读不会影响模型速度，30tokens对于a6b有点不正常了。</p>
<p dir="auto">另外补充：<br />
这个模型的reasoning好像设置跟其他模型不太一样，想太多，设置一下应该会好很多。</p>
]]></description><link>https://lcz.me/topic/1365</link><generator>RSS for Node</generator><lastBuildDate>Wed, 09 Sep 2026 21:48:02 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1365.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 27 Aug 2026 14:34:17 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Sat, 29 Aug 2026 10:42:19 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/johnnybegood" aria-label="Profile: johnnybegood">@<bdi>johnnybegood</bdi></a> 你想多了，3090还是更能打。</p>
]]></description><link>https://lcz.me/post/14769</link><guid isPermaLink="true">https://lcz.me/post/14769</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sat, 29 Aug 2026 10:42:19 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Sat, 29 Aug 2026 09:17:45 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/gary-pan" aria-label="Profile: Gary-Pan">@<bdi>Gary-Pan</bdi></a> 这么m5 max基本也就是3090的水平</p>
]]></description><link>https://lcz.me/post/14759</link><guid isPermaLink="true">https://lcz.me/post/14759</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Sat, 29 Aug 2026 09:17:45 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Fri, 28 Aug 2026 05:49:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/gary-pan" aria-label="Profile: Gary-Pan">@<bdi>Gary-Pan</bdi></a> 我弟回家后上图啊，就等着你这个做视频。</p>
]]></description><link>https://lcz.me/post/14546</link><guid isPermaLink="true">https://lcz.me/post/14546</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 28 Aug 2026 05:49:58 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Fri, 28 Aug 2026 00:37:41 GMT]]></title><description><![CDATA[<p dir="auto">看了一下，心仪的UD Q4 XL是111GB啊： Qwen3.8-Flash-Next-GGUF<br />
/UD-Q4_K_XL  111 GB   我的配置是48G显存 64G内存，我不想卸载到SSD还差得远。不过有条32G的想挪过来换掉 就刚刚踩线<br />
Qwen3.8-Flash-Next-GGUF  UD-IQ4_XS  93.7 GB  这个不换内存 似乎还刚好，有点心动了。</p>
<p dir="auto">不过我的决定是，死等 byteshape出 Qwen3.8-Flash-Next的GGUF ，这家的MOE模型调校非常好。同等体积下总是比别人的格式更均衡。</p>
]]></description><link>https://lcz.me/post/14477</link><guid isPermaLink="true">https://lcz.me/post/14477</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Fri, 28 Aug 2026 00:37:41 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Thu, 27 Aug 2026 14:48:23 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/gary-pan" aria-label="Profile: Gary-Pan">@<bdi>Gary-Pan</bdi></a> <a href="/post/14401">said</a>:</p>
<p dir="auto">ngram 50b 左右的部分可以放到ssd里直接读，不需要进内存，实际上就是一个120b a6b的模型，4 bit也就50多gb，还只要a6b，潜力巨大。</p>
</blockquote>
<p dir="auto">哇 謝謝分享 看樣子有機會可以試試看了 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f642.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--slightly_smiling_face" style="height:23px;width:auto;vertical-align:middle" title=":)" alt="🙂" /></p>
]]></description><link>https://lcz.me/post/14413</link><guid isPermaLink="true">https://lcz.me/post/14413</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Thu, 27 Aug 2026 14:48:23 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Thu, 27 Aug 2026 14:43:46 GMT]]></title><description><![CDATA[<p dir="auto">是的，感觉加上mtp，优化一下调用，速度应该能至少翻倍，就赢麻了，智力是完全够的</p>
]]></description><link>https://lcz.me/post/14409</link><guid isPermaLink="true">https://lcz.me/post/14409</guid><dc:creator><![CDATA[Gary Pan]]></dc:creator><pubDate>Thu, 27 Aug 2026 14:43:46 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Thu, 27 Aug 2026 14:42:08 GMT]]></title><description><![CDATA[<p dir="auto">已经达到可用的程度了。<br />
按 Mac的功耗。这个模型可用价值一下就上来了。</p>
]]></description><link>https://lcz.me/post/14407</link><guid isPermaLink="true">https://lcz.me/post/14407</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Thu, 27 Aug 2026 14:42:08 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-Flash-Next M5 Max实测 on Thu, 27 Aug 2026 14:37:11 GMT]]></title><description><![CDATA[<p dir="auto">这是Hermes给的 llama.cpp 的设置表，供参考。<br />
llama.cpp 服务端 reasoning 相关 flag 全表（从 master 源码 common/arg.cpp 实读，非记忆）：</p>
<p dir="auto">Flag: -rea / --reasoning<br />
参数值: on / off / auto<br />
默认: auto（从 chat template 自动检测）<br />
环境变量: LLAMA_ARG_REASONING<br />
作用: 总开关：thinking 开/关。on/off 会写入模板 kwarg enable_thinking<br />
能否被请求级覆盖: <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> body 里 chat_template_kwargs.enable_thinking 或 reasoning_effort:"none" 可覆盖（实测确认）<br />
────────────────────────────────────────<br />
Flag: --reasoning-effort<br />
参数值: default / minimal / low / medium / high / xhigh / max…<br />
默认: default（保持模板默认）<br />
环境变量: LLAMA_ARG_REASONING_EFFORT<br />
作用: 把档位塞进 chat template kwarg reasoning_effort——是否生效取决于模型模板认不认（Qwen3 系只认 on/off 布尔，不认档位）<br />
能否被请求级覆盖: <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> body 里 chat_template_kwargs.reasoning_effort<br />
────────────────────────────────────────<br />
Flag: --reasoning-budget<br />
参数值: N：-1 不限 / 0 立即结束 / &gt;0 思考 token 上限<br />
默认: -1<br />
环境变量: LLAMA_ARG_THINK_BUDGET<br />
作用: 采样层的思考 token 预算——服务端硬截断，与模板无关，对不认 effort 的模型这是唯一有效的"限流"手段<br />
能否被请求级覆盖: <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a0.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--warning" style="height:23px;width:auto;vertical-align:middle" title="⚠" alt="⚠" />️ body 有对应字段 thinking_budget_tokens（Anthropic 风格 thinking.budget_tokens 也会被翻译成它）<br />
────────────────────────────────────────<br />
Flag: --reasoning-budget-message<br />
参数值: 任意文本<br />
默认: 无<br />
环境变量: LLAMA_ARG_THINK_BUDGET_MESSAGE<br />
作用: 预算耗尽时，在 end-of-thinking 标签前注入的提示语（引导模型收尾）<br />
能否被请求级覆盖: <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 仅启动级<br />
────────────────────────────────────────<br />
Flag: --reasoning-preserve / --no-reasoning-preserve<br />
参数值: bool<br />
默认: 模板默认<br />
环境变量: —<br />
作用: 是否把思考痕迹保留在完整历史里（而非只有最后一条 assistant 消息）。需要模板有 supports_preserve_reasoning 能力<br />
能否被请求级覆盖: <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> body kwarg preserve_reasoning<br />
────────────────────────────────────────<br />
Flag: --reasoning-format<br />
参数值: none / deepseek / deepseek-legacy / auto<br />
默认: auto<br />
环境变量: LLAMA_ARG_THINK<br />
作用: 思考内容怎么交付：none=不解析留在 content；deepseek=放进 message.reasoning_content；legacy=两处都有<br />
能否被请求级覆盖: <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> body 里 reasoning_format<br />
────────────────────────────────────────<br />
Flag: --chat-template-kwargs<br />
参数值: JSON 对象<br />
默认: {}<br />
环境变量: LLAMA_ARG_CHAT_TEMPLATE_KWARGS<br />
作用: 任意模板 kwargs（你现在进程里的 {"enable_thinking": true, "preserve_thinking": true} 就是它设的）。<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a0.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--warning" style="height:23px;width:auto;vertical-align:middle" title="⚠" alt="⚠" />️ 新版已弃用 enable_thinking<br />
走这里，会打警告让你改用 --reasoning on/off<br />
能否被请求级覆盖: <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> body 里 chat_template_kwargs 与之合并，请求值优先<br />
────────────────────────────────────────<br />
Flag: --no-reasoning-in-content（force_pure_content 相关）<br />
参数值: bool<br />
默认: disabled<br />
环境变量: —<br />
作用: 禁止工具调用/思考混进 content 块<br />
能否被请求级覆盖: （随解析器配置，一般不用动）</p>
]]></description><link>https://lcz.me/post/14402</link><guid isPermaLink="true">https://lcz.me/post/14402</guid><dc:creator><![CDATA[Gary Pan]]></dc:creator><pubDate>Thu, 27 Aug 2026 14:37:11 GMT</pubDate></item></channel></rss>