<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Qwen3.6 27b dense vs Gemma4 31b dense]]></title><description><![CDATA[<p dir="auto">有没有大佬测试过</p>
<p dir="auto">我自己试了几次<br />
感觉gemma4 很笨</p>
<p dir="auto">同一个任务 qwen 省心很多<br />
读pdf ，还要教他怎样读<br />
用mermaid cli, skill 已经有了 但是他说他不会</p>
<p dir="auto">我都怀疑是不是我的设置错了</p>
]]></description><link>https://lcz.me/topic/608/qwen3.6-27b-dense-vs-gemma4-31b-dense</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 23:29:53 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/608.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 18 Jun 2026 06:36:56 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Qwen3.6 27b dense vs Gemma4 31b dense on Thu, 18 Jun 2026 11:24:39 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a></p>
<p dir="auto">我找了一下<br />
说是比较偏向coding<br />
但是不能调用工具查资料 coding 也废了一半<br />
我的coding 比较多是商业逻辑</p>
<p dir="auto">是有看到gemma比较差 但是没想到我完全用不了</p>
<p dir="auto">我感觉QWEN3.6 27b 的智力有以前MINIMAX 2.7 的智力<br />
MINIMAX3.0 没用过所以不知道</p>
<p dir="auto">看了你的那个链接 感觉我上面的体验是对的</p>
]]></description><link>https://lcz.me/post/7330</link><guid isPermaLink="true">https://lcz.me/post/7330</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Thu, 18 Jun 2026 11:24:39 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.6 27b dense vs Gemma4 31b dense on Thu, 18 Jun 2026 11:18:23 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a><br />
除了 Tool-call parser<br />
其他的我觉得都没什么关系</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>Setting</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Image</td>
<td><code>vllm/vllm-openai:v0.22.0</code></td>
</tr>
<tr>
<td>Container</td>
<td><code>vllm-gemma-bf16</code></td>
</tr>
<tr>
<td>Restart</td>
<td><code>"no"</code> (managed by <code>vllm.service</code>)</td>
</tr>
<tr>
<td>Model</td>
<td><code>Intel/gemma-4-31B-it-int4-AutoRound</code> (mounted as <code>gemma-4-31b-autoround-int4</code>)</td>
</tr>
<tr>
<td>Draft (MTP)</td>
<td><code>google/gemma-4-31B-it-assistant</code> (0.5B drafter)</td>
</tr>
<tr>
<td>Served name</td>
<td><code>gemma-4-31b-bf16</code></td>
</tr>
<tr>
<td>Tensor parallel</td>
<td>2 (dual 3090)</td>
</tr>
<tr>
<td>Max model len</td>
<td>131,072</td>
</tr>
<tr>
<td>GPU mem util</td>
<td>0.95</td>
</tr>
<tr>
<td>Max num seqs</td>
<td>4 (concurrent)</td>
</tr>
<tr>
<td>Max batched tokens</td>
<td>4,096 (must fit vision tower's 2,496 mm tokens)</td>
</tr>
<tr>
<td>KV cache</td>
<td>default (BF16) — explicitly NOT fp8 (Ampere can't)</td>
</tr>
<tr>
<td>Speculative decoding</td>
<td>MTP n=4 (Google's official assistant)</td>
</tr>
<tr>
<td>Reasoning parser</td>
<td><code>gemma4</code> (separates <code>&lt;channel&gt;thought</code> from content)</td>
</tr>
<tr>
<td>Tool-call parser</td>
<td><code>gemma4</code> (+ <code>--enable-auto-tool-choice</code>)</td>
</tr>
<tr>
<td>Chat template</td>
<td><code>tool_chat_template_gemma4.jinja</code> (from vLLM image)</td>
</tr>
<tr>
<td>Sampling override</td>
<td>T=1.0, top_p=0.95, top_k=64, min_p=0, no rep penalty</td>
</tr>
<tr>
<td>Overlays</td>
<td>1 only — PR #42006 (streaming multi-tool-call fix)</td>
</tr>
</tbody>
</table>
]]></description><link>https://lcz.me/post/7329</link><guid isPermaLink="true">https://lcz.me/post/7329</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Thu, 18 Jun 2026 11:18:23 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.6 27b dense vs Gemma4 31b dense on Thu, 18 Jun 2026 09:07:22 GMT]]></title><description><![CDATA[<p dir="auto">各种benchmark也都支持你的结论。比如：<a href="https://artificialanalysis.ai/models/gemma-4-31b" rel="nofollow ugc">https://artificialanalysis.ai/models/gemma-4-31b</a></p>
<p dir="auto">而且他的benchmark很偏科，有的跑分可以跑到和qwen3.6-27B旗鼓相当，有的只在qwen3.6-35B-A3B的一半……</p>
<p dir="auto">有一种说法是gemma4不针对benchmark优化<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f602.png?v=9786174bac0" class="not-responsive emoji emoji-android emoji--joy" style="height:23px;width:auto;vertical-align:middle" title=":joy:" alt="😂" /> ，但是这种东西谁知道呢……</p>
]]></description><link>https://lcz.me/post/7304</link><guid isPermaLink="true">https://lcz.me/post/7304</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Thu, 18 Jun 2026 09:07:22 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.6 27b dense vs Gemma4 31b dense on Thu, 18 Jun 2026 08:17:24 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/applejuice" aria-label="Profile: applejuice">@<bdi>applejuice</bdi></a> 展示你的配置。我们猜测不了你的配置。gamma4 我用了。智力尚可。</p>
]]></description><link>https://lcz.me/post/7291</link><guid isPermaLink="true">https://lcz.me/post/7291</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Thu, 18 Jun 2026 08:17:24 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.6 27b dense vs Gemma4 31b dense on Thu, 18 Jun 2026 07:03:26 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/applejuice" aria-label="Profile: applejuice">@<bdi>applejuice</bdi></a> 不是你的设置问题，Gemma4 在指令跟随和工具调用方面确实比 Qwen 3.6 弱不少。我测试下来几个关键区别：</p>
<ol>
<li>
<p dir="auto">训练数据差异：Gemma4 的预训练数据偏学术/代码，SFT 阶段的 tool-use 数据量不如 Qwen 3.6 丰富，所以"教它用 mermaid CLI 也不会"这个现象很常见。</p>
</li>
<li>
<p dir="auto">指令跟随风格：Gemma 系列对 system prompt 的敏感度比 Qwen 低，同样的 skill 描述在 Qwen 上能严格执行，在 Gemma 上可能就理解成"参考一下"。</p>
</li>
<li>
<p dir="auto">PDF 阅读：Qwen 3.6 的 long-context 能力更强（131K tokens 实测稳定），Gemma4 虽然支持 256K 但中间层 attention 容易丢失细节，尤其是多页 PDF 的结构化信息。</p>
</li>
</ol>
<p dir="auto">建议：如果你主要做 PDF 处理 + tool use，继续用 Qwen 3.6 27b 作为主力。Gemma4 可以留着做辅助判断（比如分类任务、简单 QA），它在这个场景下表现还不错。</p>
]]></description><link>https://lcz.me/post/7283</link><guid isPermaLink="true">https://lcz.me/post/7283</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 18 Jun 2026 07:03:26 GMT</pubDate></item></channel></rss>