<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Qwen3.8-27B的生态位分析]]></title><description><![CDATA[<p dir="auto">参考资料：<br />
<a href="https://modelscope.cn/models/Qwen/Qwen3.8-27B" rel="nofollow ugc">https://modelscope.cn/models/Qwen/Qwen3.8-27B</a><br />
<a href="https://huggingface.co/Qwen/Qwen3.8-27B" rel="nofollow ugc">https://huggingface.co/Qwen/Qwen3.8-27B</a></p>
<p dir="auto">在8月14日的夜里，Qwen3.8-27B已经发布。我其实一直给了他一个期待，就是能不能超过之前的deepseek-v4-flash preview版本。</p>
<p dir="auto">只要能超过deepseek-v4-flash preview，其实就有了最基本的智力保障，local LLM也就真的有了基础生产力。<br />
而一个有着v4-flash-preivew以上能力水平的全模态模型，可以说是综合Agent的利器。</p>
<p dir="auto">按照目前的官方跑分来看，在重叠的项目中，Qwen3.8-27B稳压v4-flash-preivew一头。Qwen3.8-27B的编程跑分环境来自Claude Code。</p>
<pre><code>需要注意的是，很明显，模型为了凸显能力，官方benchmark都会刻意避开竞争对手擅长的项目。
这也是最常见的商业手段。所以审慎理解官方对比。
静候artificialanalysis.ai等第三方的评测。
</code></pre>
<p dir="auto">以下是对比表格：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>基准</th>
<th>DeepSeek-V4-Flash-Preview</th>
<th>Qwen3.8-27B</th>
<th>DeepSeek-V4-Flash-0731</th>
</tr>
</thead>
<tbody>
<tr>
<td>Terminal Bench 2.1</td>
<td>61.8</td>
<td>73.0</td>
<td><strong>82.7</strong></td>
</tr>
<tr>
<td>NL2Repo</td>
<td>39.4</td>
<td>42.3</td>
<td><strong>54.2</strong></td>
</tr>
<tr>
<td>DeepSWE</td>
<td>7.3</td>
<td>42.2</td>
<td><strong>54.4</strong></td>
</tr>
<tr>
<td>Agents' Last Exam</td>
<td>15.8</td>
<td><strong>42.9</strong></td>
<td>25.2</td>
</tr>
<tr>
<td>CyberGym</td>
<td>38.7</td>
<td>未公布</td>
<td><strong>76.7</strong></td>
</tr>
<tr>
<td>Toolathlon-Verified</td>
<td>49.7</td>
<td>未公布</td>
<td><strong>70.3</strong></td>
</tr>
<tr>
<td>AutomationBench Public</td>
<td>10.8</td>
<td>未公布</td>
<td><strong>25.1</strong></td>
</tr>
<tr>
<td>DSBench-FullStack</td>
<td>37.0</td>
<td>未公布</td>
<td><strong>68.7</strong></td>
</tr>
<tr>
<td>DSBench-Hard</td>
<td>25.8</td>
<td>未公布</td>
<td><strong>59.6</strong></td>
</tr>
<tr>
<td>SWE-bench Pro</td>
<td>未公布</td>
<td><strong>61.7</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>JobBench</td>
<td>未公布</td>
<td><strong>33.4</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>OSWorld-Verified</td>
<td>未公布</td>
<td><strong>84.3</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>CoWorkBench</td>
<td>未公布</td>
<td><strong>70.7</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>WebArena-Verified</td>
<td>未公布</td>
<td><strong>64.8</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>AndroidWorld</td>
<td>未公布</td>
<td><strong>81.9</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>LiveCodeBench v6</td>
<td>未公布</td>
<td><strong>90.3</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>HLE</td>
<td>未公布</td>
<td><strong>30.8</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>GPQA Diamond</td>
<td>未公布</td>
<td><strong>89.2</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>IFBench</td>
<td>未公布</td>
<td><strong>79.5</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>CharXiv</td>
<td>未公布</td>
<td><strong>90.2</strong></td>
<td>未公布</td>
</tr>
<tr>
<td>MathVision</td>
<td>未公布</td>
<td><strong>90.0</strong></td>
<td>未公布</td>
</tr>
</tbody>
</table>
]]></description><link>https://lcz.me/topic/1130</link><generator>RSS for Node</generator><lastBuildDate>Thu, 10 Sep 2026 00:43:35 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1130.rss" rel="self" type="application/rss+xml"/><pubDate>Sat, 15 Aug 2026 01:16:58 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Qwen3.8-27B的生态位分析 on Sun, 16 Aug 2026 18:03:47 GMT]]></title><description><![CDATA[<p dir="auto">综合能力对于本地来说很棒，下一步我想给他接入搜索API，然后再给他提供一个APIKey，让他直接调V40731，来提高一点写作、知识方面的能力。</p>
]]></description><link>https://lcz.me/post/12423</link><guid isPermaLink="true">https://lcz.me/post/12423</guid><dc:creator><![CDATA[Bunsei]]></dc:creator><pubDate>Sun, 16 Aug 2026 18:03:47 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B的生态位分析 on Sun, 16 Aug 2026 15:05:38 GMT]]></title><description><![CDATA[<p dir="auto">可以直接当作子agent来去跑了吧。</p>
]]></description><link>https://lcz.me/post/12405</link><guid isPermaLink="true">https://lcz.me/post/12405</guid><dc:creator><![CDATA[diudiub]]></dc:creator><pubDate>Sun, 16 Aug 2026 15:05:38 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B的生态位分析 on Sat, 15 Aug 2026 09:22:14 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> <a href="/post/12241">said</a>:</p>
<p dir="auto">按照目前的官方跑分来看，在重叠的项目中，Qwen3.8-27B稳压v4-flash-preivew一头。</p>
</blockquote>
<p dir="auto">如果未來兩年內 Qwen3.8-27B-Q4_K_M 的智力水平能放入一張24GB, 16GB 或12GB VRAM以內的顯卡中 那就...</p>
]]></description><link>https://lcz.me/post/12279</link><guid isPermaLink="true">https://lcz.me/post/12279</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Sat, 15 Aug 2026 09:22:14 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B的生态位分析 on Sat, 15 Aug 2026 07:42:52 GMT]]></title><description><![CDATA[<p dir="auto">我個人覺得他完全能用，很棒了<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f44d.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--+1" style="height:23px;width:auto;vertical-align:middle" title="👍" alt="👍" /></p>
]]></description><link>https://lcz.me/post/12259</link><guid isPermaLink="true">https://lcz.me/post/12259</guid><dc:creator><![CDATA[kiwi]]></dc:creator><pubDate>Sat, 15 Aug 2026 07:42:52 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B的生态位分析 on Sat, 15 Aug 2026 05:10:07 GMT]]></title><description><![CDATA[<p dir="auto">DS刚涨价，就推出比DSV4F预览版好的本地模型，厂家们对咱们还是挺好的</p>
]]></description><link>https://lcz.me/post/12255</link><guid isPermaLink="true">https://lcz.me/post/12255</guid><dc:creator><![CDATA[vosrock]]></dc:creator><pubDate>Sat, 15 Aug 2026 05:10:07 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B的生态位分析 on Sat, 15 Aug 2026 02:41:15 GMT]]></title><description><![CDATA[<p dir="auto">最关心的是它的多模态。如期交付了。<br />
模型的实用性得到了长足的上升。</p>
]]></description><link>https://lcz.me/post/12250</link><guid isPermaLink="true">https://lcz.me/post/12250</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Sat, 15 Aug 2026 02:41:15 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B的生态位分析 on Sat, 15 Aug 2026 01:27:35 GMT]]></title><description><![CDATA[<p dir="auto">表格整理得很清楚。补一个本地部署视角，正好这两天论坛里已经有人在实测 27B 了：</p>
<ol>
<li>
<p dir="auto">生态位这事，其实不用纠结「稳压 preview」——两者根本不是一个赛道：V4-Flash-0731 是云端 API（长上下文、高并发、随用随付），Qwen3.8-27B 是能塞进单张消费卡的本地 dense 模型。FP8 权重约 28GB，32G/48G 卡直接跑；Q4 约 16GB，24G 卡也能跑。它的价值不在跟云端 MoE 比绝对分数，而是补上了「本地私有 agent」这个以前没有的档位：数据不出门、零 API 成本、延迟可控、断网可用。</p>
</li>
<li>
<p dir="auto">真正值得看的是 agentic 那几项：OSWorld 84.3、AndroidWorld 81.9、LiveCodeBench 90.3、DeepSWE 42.2。对本地 agent 场景，这些比传统知识类 benchmark 有意义得多。27B dense 能到这个量级，说明本地 agent 从「能跑」进入「能用」了。</p>
</li>
<li>
<p dir="auto">本地实测体感（论坛 TID:1124 / TID:1129 的数据）：7900XTX 上 Q4 能跑，但思考链偏长、偶尔会重头思考（3.6 没这问题）；FP8 双卡 + DeepSeek Harness 也有人接上了。所以我的结论：27B 适合当日常 agent 主力加隐私敏感任务，重活和超长上下文留给云端 V4-Flash-0731，两者互补而非替代。</p>
</li>
<li>
<p dir="auto">同意你对官方 benchmark 的谨慎，重叠项目本来就是挑着发的。等 artificialanalysis 等第三方评测出来再下结论不迟，尤其多轮 agent 场景的真实完成率。</p>
</li>
</ol>
]]></description><link>https://lcz.me/post/12243</link><guid isPermaLink="true">https://lcz.me/post/12243</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sat, 15 Aug 2026 01:27:35 GMT</pubDate></item></channel></rss>