<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪]]></title><description><![CDATA[<p dir="auto">今天做同一个任务，分别用两个模型跑了一下：</p>
<p dir="auto"><strong>任务内容</strong>：更新Github Repo <code>readme.md</code>，要求架构图、数据流图、类图、流程图各一张（Mermaid），正文中文、图中英文</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>模型</th>
<th>结果</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Qwen3.8:27b</strong></td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/23f1.png?v=60716d54ab2" class="not-responsive emoji emoji-android emoji--stopwatch" style="height:23px;width:auto;vertical-align:middle" title="⏱" alt="⏱" />️ 跑了 ~25min，没出结果（后被我停止）</td>
</tr>
<tr>
<td><strong>Qwen3.6:27b</strong></td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=60716d54ab2" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 13min 完成</td>
</tr>
</tbody>
</table>
<p dir="auto">两个模型参数量都是 27B，差距挺明显的。3.8 在复杂文档生成 + Mermaid 多图表这种任务上似乎不太稳，偶尔“空转/内耗”，3.6 反而更靠谱。</p>
<p dir="auto">两天用下来总体感受<br />
1 复杂长链路，3.8逻辑会逻辑链路断裂，导致"内耗" -- 只费电，不做功<br />
2 速度方面3.6明显更快，不是pp/token generation测试速度，而是总体任务速度，同一个任务，同样的输入，3.6快100%甚至更多。<br />
3 复杂图像识别方面，3.8好很多。 “3.6 + 路由 + 3B的小图像识别模型”也可以解决问题，结果几乎一致<br />
4 有人说3.8代码能力强，可是代码能力基于逻辑能力，逻辑都不理解，代码能力从何谈起？ 问题都找不到，怎么改？<br />
总体来说：local 3.8 qwen 实用性表示怀疑。</p>
<p dir="auto">群里有没有跑过类似对比的？大家感觉 <strong>3.8 相对 3.6 的提升在哪类任务上比较明显</strong>？欢迎点评、讨论</p>
]]></description><link>https://lcz.me/topic/1173/随手记录-qwen3.8-27b-vs-qwen3.6-27b-实测对比</link><generator>RSS for Node</generator><lastBuildDate>Sat, 22 Aug 2026 03:28:48 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1173.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 18 Aug 2026 01:16:30 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Fri, 21 Aug 2026 04:47:36 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/daniel-zhou" aria-label="Profile: daniel-zhou">@<bdi>daniel-zhou</bdi></a> <a href="/post/13043">说</a>:</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/d8d26a01-96b1-4b7d-a592-fdb355e2788a.jpeg" alt="35de943f-8efa-47bf-b274-f61d1e7e7c07-image.jpeg" class=" img-fluid img-markdown" /><br />
唉，不知道怎么搞了</p>
</blockquote>
<p dir="auto">我的结论以下。单卡3090 24g不适合干很重的活，而轻量的不如走api。<br />
老特也推荐，跑图文必须自己搞。llm 买api合适。</p>
]]></description><link>https://lcz.me/post/13225</link><guid isPermaLink="true">https://lcz.me/post/13225</guid><dc:creator><![CDATA[wwcd2016]]></dc:creator><pubDate>Fri, 21 Aug 2026 04:47:36 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Thu, 20 Aug 2026 06:04:24 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/d8d26a01-96b1-4b7d-a592-fdb355e2788a.jpeg" alt="35de943f-8efa-47bf-b274-f61d1e7e7c07-image.jpeg" class=" img-fluid img-markdown" /><br />
唉，不知道怎么搞了</p>
]]></description><link>https://lcz.me/post/13043</link><guid isPermaLink="true">https://lcz.me/post/13043</guid><dc:creator><![CDATA[daniel zhou]]></dc:creator><pubDate>Thu, 20 Aug 2026 06:04:24 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Thu, 20 Aug 2026 00:47:44 GMT]]></title><description><![CDATA[<p dir="auto">更新一下，我目前用签名的模型，双卡单流256K，已经稳定使用了。再也不折腾本地基础设施了。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/7d62ff11-dac2-41bb-bbbc-c5f0e0412c10.jpeg" alt="ac1bc264-2226-458f-9be7-c48351e7f899-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/7fefb174-ae11-49e4-91d2-8b30ea79b835.jpeg" alt="064a8c3a-79af-4aec-8df5-5d21e47198e3-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/12987</link><guid isPermaLink="true">https://lcz.me/post/12987</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Thu, 20 Aug 2026 00:47:44 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Wed, 19 Aug 2026 13:56:10 GMT]]></title><description><![CDATA[<p dir="auto">3.8在编程上体感是大幅领先于3.6，3.6接入Agent根本不能干啥正儿八经的活，3.8则大部分工作都能完成。数据很详细，虽然是AI整理的，但是贴主主导思想，把握的很好。格式公正，易于阅读，置顶。</p>
]]></description><link>https://lcz.me/post/12925</link><guid isPermaLink="true">https://lcz.me/post/12925</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 19 Aug 2026 13:56:10 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Wed, 19 Aug 2026 13:03:16 GMT]]></title><description><![CDATA[<h1>qwen3.8:27b vs qwen3.6:27b — Thinking 模式实测对比</h1>
<blockquote>
<p dir="auto">测试环境：内网 AI 服务器（Ubuntu + AMD 7900 XTX 24G + ROCm 7.x + Ollama 0.32.13）<br />
调用方式：OpenAI 兼容接口，<code>max_tokens=6000</code>，非流式<br />
测试时间：2026-08-19</p>
</blockquote>
<h2>背景</h2>
<p dir="auto">qwen3.8:27b 默认 thinking = max，慢到令人发指（简单问题也要等 3-5 分钟）。给请求注入 <code>reasoning_effort: medium</code> 后明显改善。但有人还在用 qwen3.6:27b，于是做了一组公平对比。</p>
<h2>测试设计</h2>
<p dir="auto">两道长链推理题，统一参数发送：</p>
<ul>
<li><strong>过桥题</strong>：四人过桥（1/2/5/10 分钟），求最短总时间及方案（经典优化题，答案 17 分钟）</li>
<li><strong>逻辑题</strong>：5 人 5 题答对关系，求哪题答错人数最多（多约束推理，答案：第 1 题，B/C/E）</li>
</ul>
<p dir="auto">三组配置：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>编号</th>
<th>模型</th>
<th>thinking 模式</th>
<th>说明</th>
</tr>
</thead>
<tbody>
<tr>
<td>A</td>
<td>qwen3.8:27b</td>
<td>medium</td>
<td>服务端默认注入</td>
</tr>
<tr>
<td>B</td>
<td>qwen3.6:27b</td>
<td>medium</td>
<td>同等级，对比模型质量</td>
</tr>
<tr>
<td>C</td>
<td>qwen3.6:27b</td>
<td>default</td>
<td>不传 reasoning_effort，模型自然 thinking</td>
</tr>
</tbody>
</table>
<h2>结果</h2>
<h3>第一轮（max_tokens=3000）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>配置</th>
<th>过桥题</th>
<th>正文输出</th>
<th>思考长度</th>
<th>耗时</th>
<th>正确？</th>
</tr>
</thead>
<tbody>
<tr>
<td>A: 3.8+medium</td>
<td>推理得 17 分钟</td>
<td>119 字（截断）</td>
<td>3842 字</td>
<td>63s</td>
<td>正文被 max_tokens 截断</td>
</tr>
<tr>
<td>B: 3.6+medium</td>
<td>—</td>
<td><strong>0 字</strong></td>
<td>7867 字</td>
<td>100s</td>
<td>全部 token 被思考吃光</td>
</tr>
<tr>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f604.png?v=60716d54ab2" class="not-responsive emoji emoji-android emoji--smile" style="height:23px;width:auto;vertical-align:middle" title="C:" alt="😄" /> 3.6+default</td>
<td>17 分钟</td>
<td>885 字</td>
<td>5856 字</td>
<td>85s</td>
<td>正确</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong>3.6 + medium 直接坏掉</strong>：3000 token 全被 thinking 消耗，正文输出 0 字符。和 3.8 + max 的问题如出一辙。</p>
</blockquote>
<h3>第二轮（max_tokens=6000，只测两个可用配置）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>配置</th>
<th>任务</th>
<th>耗时</th>
<th>正文</th>
<th>思考</th>
<th>Token</th>
<th>正确？</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>3.8+medium</strong></td>
<td>过桥</td>
<td><strong>61.8s</strong></td>
<td>994 字</td>
<td>5831 字</td>
<td>2646</td>
<td>17 分钟</td>
</tr>
<tr>
<td>3.6+default</td>
<td>过桥</td>
<td>112.9s</td>
<td>1346 字</td>
<td>5897 字</td>
<td>3305</td>
<td>17 分钟</td>
</tr>
<tr>
<td><strong>3.8+medium</strong></td>
<td>逻辑</td>
<td><strong>19.3s</strong></td>
<td>564 字</td>
<td>413 字</td>
<td>777</td>
<td>第1题 BCE</td>
</tr>
<tr>
<td>3.6+default</td>
<td>逻辑</td>
<td>86.5s</td>
<td>634 字</td>
<td>5397 字</td>
<td>2532</td>
<td>第1题 BCE</td>
</tr>
</tbody>
</table>
<h3>汇总</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>3.8 + medium</th>
<th>3.6 + default</th>
<th>倍率</th>
</tr>
</thead>
<tbody>
<tr>
<td>总耗时</td>
<td><strong>81.1s</strong></td>
<td>199.4s</td>
<td>3.8 快 <strong>2.5x</strong></td>
</tr>
<tr>
<td>总 token</td>
<td><strong>3423</strong></td>
<td>5837</td>
<td>3.8 省 <strong>1.7x</strong></td>
</tr>
<tr>
<td>逻辑题思考长度</td>
<td><strong>413 字</strong></td>
<td>5397 字</td>
<td>3.8 少 <strong>13x</strong></td>
</tr>
<tr>
<td>两题正确率</td>
<td>2/2</td>
<td>2/2</td>
<td>持平</td>
</tr>
</tbody>
</table>
<h2>分析</h2>
<h3>1. 速度：3.8 碾压</h3>
<p dir="auto">过桥题 62s vs 113s（快 1.8x），逻辑题 19s vs 87s（快 <strong>4.5x</strong>）。简单任务上差距最大——3.6 不管题目难易都重度思考，3.8 会自适应：难题 5831 字符思考，简单题仅 413 字符。</p>
<h3>2. Token 效率：3.8 省三倍</h3>
<p dir="auto">3.8 总用 3423 token，3.6 用 5837。3.6 在逻辑题上花了 2532 token，其中 5397 字符是英文思考过程——对"统计答错人数"这种题来说严重过度推理。</p>
<h3>3. 质量：打平</h3>
<p dir="auto">两道题两个模型都答对了。3.8 的过桥答案带了下界证明（数学严谨），3.6 的带了两种策略公式对比。都是好答案。</p>
<h3>4. 3.6 + medium 不可用</h3>
<p dir="auto">这是最关键的发现：给 3.6 显式设 <code>reasoning_effort=medium</code> 会导致 thinking 吞掉全部 token 预算，正文输出为零。3.6 只能用 default（不加 effort 字段），但那样又慢又费 token。</p>
<h2>结论</h2>
<blockquote>
<p dir="auto"><strong>主力模型用 <code>qwen3.8:27b</code> + <code>think=medium</code>。</strong></p>
</blockquote>
<ul>
<li>速度快 2.5 倍</li>
<li>Token 省 1.7 倍</li>
<li>答题质量相同</li>
<li>3.8 会根据任务难度自适应思考量，3.6 不会</li>
<li>服务端注入 <code>reasoning_effort: medium</code> 是当前最优配置</li>
</ul>
<p dir="auto">如果还有客户端在用 <code>qwen3.6:27b</code>，改成 <code>qwen3.8:27b</code> 即可——速度翻倍，质量不变。</p>
<hr />
<p dir="auto"><em>补充：Ollama 的 thinking 等级合法值为 <code>low / medium / high / max / true / false</code>。OpenAI 兼容端点的 <code>reasoning_effort</code> 会被自动映射为对应等级；客户端自带的值优先于服务端默认值。</em></p>
]]></description><link>https://lcz.me/post/12915</link><guid isPermaLink="true">https://lcz.me/post/12915</guid><dc:creator><![CDATA[Fan Rex]]></dc:creator><pubDate>Wed, 19 Aug 2026 13:03:16 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Tue, 18 Aug 2026 06:53:40 GMT]]></title><description><![CDATA[<p dir="auto">内部思维链应该是也调整过，总体还是超过预期，本来我以为只是3.6套个 壳换下版本号，没想到还真有东西。</p>
]]></description><link>https://lcz.me/post/12688</link><guid isPermaLink="true">https://lcz.me/post/12688</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Tue, 18 Aug 2026 06:53:40 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Tue, 18 Aug 2026 05:57:54 GMT]]></title><description><![CDATA[<p dir="auto">我觉得应用3.8思维深度的medium档去与3.6去对比，因为medium档是没有额外指令的，其他两个都有额外指令，3.6默认也没有（猜的）。</p>
]]></description><link>https://lcz.me/post/12681</link><guid isPermaLink="true">https://lcz.me/post/12681</guid><dc:creator><![CDATA[neo]]></dc:creator><pubDate>Tue, 18 Aug 2026 05:57:54 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Tue, 18 Aug 2026 04:22:39 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E4%B9%9D%E9%BE%99%E6%9D%A8%E7%94%9F" aria-label="Profile: 九龙杨生">@<bdi>九龙杨生</bdi></a> 看自己的模型和硬件吧，我这KV CACHE池有2X 256K容量，我平时主要被HINDSIGHT卡死的.  如果没测过自己实战里面超过180K上下文的实际表现和速度，再谈循环也是枉然吧？  特别是llama.cpp，基本就是玩具的，这里量化，那里量化，上下文还打折，别指望它能发挥全部实力</p>
]]></description><link>https://lcz.me/post/12665</link><guid isPermaLink="true">https://lcz.me/post/12665</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Tue, 18 Aug 2026 04:22:39 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Tue, 18 Aug 2026 04:16:46 GMT]]></title><description><![CDATA[<p dir="auto">会的，xhigh 死循环/空转不是个例，论坛上有现成证据链：</p>
<ol>
<li>
<p dir="auto">TID:1136 里 terry 实测 3.8 长链推理频繁崩溃（"替代 DeepSeek 不现实"），那还只是默认档；xhigh 是最高思考预算，思考链最长，一旦模型在某个子问题上反复绕圈，就是纯空转——只费电不做功。</p>
</li>
<li>
<p dir="auto">6% 这个数字跟论坛体感对得上，而且概率不是均匀的：任务越难、链路越长越容易触发，hard 任务上比 6% 高不少也正常。</p>
</li>
<li>
<p dir="auto">档位怎么选（包磊 TID:1147 实测）：low/medium 实测效果几乎无差；xhigh 只在一次性高难度题上有意义——stxpnet 说"一次成功率提升"也是真的，但代价是时间和死循环风险。给 agent/长链任务用就锁 low/medium；一次性难题可以开 xhigh，但要能等，超时手动打断重问一次就行。</p>
</li>
<li>
<p dir="auto">想省心可以装 froggeric/Qwen-Fixed-Chat-Templates（TID:1151），v22 模板对思考档位控制更好，还能 enable_thinking=false 开快速档。</p>
</li>
</ol>
<p dir="auto">一句话：xhigh 是"高难度单题"档，不是"日常/长链"档。</p>
]]></description><link>https://lcz.me/post/12664</link><guid isPermaLink="true">https://lcz.me/post/12664</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 18 Aug 2026 04:16:46 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Tue, 18 Aug 2026 03:52:25 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stxpnet" aria-label="Profile: stxpnet">@<bdi>stxpnet</bdi></a> 不是说开了思考Xhigh比较容易陷入死循环吗？我看他们的截图有超过6%可能性；</p>
]]></description><link>https://lcz.me/post/12655</link><guid isPermaLink="true">https://lcz.me/post/12655</guid><dc:creator><![CDATA[九龙杨生]]></dc:creator><pubDate>Tue, 18 Aug 2026 03:52:25 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Tue, 18 Aug 2026 03:22:34 GMT]]></title><description><![CDATA[<p dir="auto">只要你显卡的算力和 k v cache给够了，正确使用3.8新的模板，里面有不同档位的，效率将会起飞： 后面等大神更新更智能一些的模板吧：</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/f13053b0-d6eb-480b-8685-48e6822f968e.jpeg" alt="1d71a463-cc6a-4318-bb4c-7cee0134cc33-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">新写代码，或者是人类判断比较难的任务，直接点xhigh,一次成功率大大提升：<br />
<img src="https://upload.lcz.me/uploads/8af98ccc-fe63-4a60-a92b-0f9bf69de4af.jpeg" alt="e9898dea-f8e6-40fb-9c44-802f2b871e37-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/12645</link><guid isPermaLink="true">https://lcz.me/post/12645</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Tue, 18 Aug 2026 03:22:34 GMT</pubDate></item><item><title><![CDATA[Reply to 随手记录：Qwen3.8:27b vs Qwen3.6:27b 实测对比 🧪 on Tue, 18 Aug 2026 03:06:45 GMT]]></title><description><![CDATA[<p dir="auto">Qwen3.8:27b 剛剛為了解一個高難度題目 跑了33分鐘, 我20分鐘時還以為當機了, 插話問它“現狀如何？” , 他回我後繼續計算, 過一會兒33分鐘時我回去看 發現居然吐出了答案 而且答案正確</p>
<p dir="auto">我測試過Qwen3.8:27b Reasoning On and Off<br />
有些題目不開推理(Off), LLM直接手推 快很多 而且答案也正確</p>
]]></description><link>https://lcz.me/post/12643</link><guid isPermaLink="true">https://lcz.me/post/12643</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Tue, 18 Aug 2026 03:06:45 GMT</pubDate></item></channel></rss>