<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ??]]></title><description><![CDATA[<p dir="auto">GPT5.6-SOL 模型幻覺吹噓自己的同類 ？<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f612.png?v=301515bb865" class="not-responsive emoji emoji-android emoji--unamused" style="height:23px;width:auto;vertical-align:middle" title=":unamused:" alt="😒" /></p>
<pre><code>Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer
</code></pre>
<p dir="auto">考官模型 ：GPT5.6-SOL (網頁版)<br />
受測模型規格 ：Qwen3.8-27B-Q4_K_M<br />
ctx : 80K<br />
KV cache: q8<br />
Chat template : Frog<br />
Reasoning: ON</p>
<p dir="auto">讓雙方對壘後來回十輪 SOL給出的評價</p>
<p dir="auto">GPT5.6-SOL (網頁版) vs. Qwen3.8<br />
<img src="https://upload.lcz.me/uploads/2606bf1c-5416-4f74-81da-e575e3d7d73a.jpeg" alt="9be71564-3b8c-4603-a7e0-017ee4d42d-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/topic/1174</link><generator>RSS for Node</generator><lastBuildDate>Wed, 09 Sep 2026 22:53:32 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1174.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 18 Aug 2026 01:38:05 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Wed, 19 Aug 2026 01:49:41 GMT]]></title><description><![CDATA[<p dir="auto">有意思，壕无人性</p>
]]></description><link>https://lcz.me/post/12836</link><guid isPermaLink="true">https://lcz.me/post/12836</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Wed, 19 Aug 2026 01:49:41 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:58:17 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/12726">said</a>:</p>
<p dir="auto">说实话各种测试都是单轮技术问题，在项目稍微复杂时候它还是需要大点的知识权重</p>
</blockquote>
<p dir="auto">我怎麼感覺這模型你把它放在DGX Station 開個50 Agents 同時去跑 做夢都會笑</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/85944489-2926-49dd-b3de-8b60ebd749f1.jpeg" alt="58f448e1-5c4e-43ac-9a2e-8a81209e8ba1-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">128 AI Agents 在Nvidia DGX Station 上 同時開跑 (Local LLM)<br />
<a href="https://youtu.be/qV_K0nTF6gY?si=AH_iev_BndXF4FMd&amp;t=943" rel="nofollow ugc">https://youtu.be/qV_K0nTF6gY?si=AH_iev_BndXF4FMd&amp;t=943</a><br />
<img src="https://upload.lcz.me/uploads/55e8247c-c2e7-4804-b2db-57dce71cfe12.jpeg" alt="80c5a5b0-4264-4774-8a92-9ae5c945643b-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/12790</link><guid isPermaLink="true">https://lcz.me/post/12790</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:58:17 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:40:39 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> <a href="/post/12786">said</a>:</p>
<p dir="auto">虽然它可以靠很强的后训练对齐能力自查纠偏，但浪费的时间和算力是实打实的。</p>
</blockquote>
<p dir="auto">這點倒是真的, Reasoning On, 它花了33分鐘在Deepseek harness的框架下, 時間很長, 做了好幾次的自查纠偏, 不過最後還是解出了難題</p>
<p dir="auto">不考慮時間問題, 這樣的自查纠偏(自我審核)能力 在Hermes上使用它 目前Agentic的品質還不錯 沒之前Qwen3.6-27B那樣讓我擔心不想使用</p>
]]></description><link>https://lcz.me/post/12788</link><guid isPermaLink="true">https://lcz.me/post/12788</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:40:39 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:32:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a></p>
<p dir="auto">他下放了代理人架構 打破(顛覆)了當時舊有的生態系 影響非常深遠的<br />
(當然Hermes' Nous Research 也有可能當第一位破局者 假設沒Peter 的話 )</p>
]]></description><link>https://lcz.me/post/12787</link><guid isPermaLink="true">https://lcz.me/post/12787</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:32:58 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:31:12 GMT]]></title><description><![CDATA[<p dir="auto">qwen3.8-27B目前的问题就出在27B上。太容易被劣质信息源给带偏。虽然它可以靠很强的后训练对齐能力自查纠偏，但浪费的时间和算力是实打实的。</p>
<p dir="auto">更何况有时候掉坑里实在爬不出来，把自己逼死循环了也很狼狈。</p>
<p dir="auto">这也是小模型的代价。</p>
]]></description><link>https://lcz.me/post/12786</link><guid isPermaLink="true">https://lcz.me/post/12786</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:31:12 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:27:21 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> 这个有他个人机遇的原因，他靠这个也确实进入了openAI。</p>
<p dir="auto">做开源要么图名声，要么图利益。要么全都要。</p>
]]></description><link>https://lcz.me/post/12783</link><guid isPermaLink="true">https://lcz.me/post/12783</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:27:21 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:26:02 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stxpnet" aria-label="Profile: stxpnet">@<bdi>stxpnet</bdi></a> <a href="/post/12778">said</a>:</p>
<p dir="auto">就蒸馏他们的能力到小模型，然后给社区免费玩，把水搅浑</p>
</blockquote>
<p dir="auto">突然想到OpenClaw的作者 Peter Steinberger<br />
非常好奇 Peter 為何要把OpenClaw 下放到民間<br />
明明只有頂級AI公司可以用的代理人架構</p>
]]></description><link>https://lcz.me/post/12782</link><guid isPermaLink="true">https://lcz.me/post/12782</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:26:02 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:24:12 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stxpnet" aria-label="Profile: stxpnet">@<bdi>stxpnet</bdi></a> <a href="/post/12778">说</a>:</p>
<p dir="auto">既然收费API 干不过对手，就蒸馏他们的能力到小模型，然后给社区免费玩，把水搅浑，这就是千问给我的感觉。</p>
</blockquote>
<p dir="auto">确实体感就是这样，肯定是后训练放飞自我了。<br />
至于说放飞自我的手段就不得而知了。</p>
<p dir="auto">否则无法解释其小模型和大模型之间的agent能力差距不成比例的问题。</p>
<p dir="auto">尤其3.8这代，这27B感觉塞得全是“答案”</p>
]]></description><link>https://lcz.me/post/12780</link><guid isPermaLink="true">https://lcz.me/post/12780</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:24:12 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 15:19:15 GMT]]></title><description><![CDATA[<p dir="auto">既然收费API 干不过对手，就蒸馏他们的能力到小模型，然后给社区免费玩，把水搅浑，这就是千问给我的感觉。</p>
]]></description><link>https://lcz.me/post/12778</link><guid isPermaLink="true">https://lcz.me/post/12778</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:19:15 GMT</pubDate></item><item><title><![CDATA[Reply to Qwen3.8-27B-Q4 整體大約像 6–10 年經驗的 Senior Engineer ?? on Tue, 18 Aug 2026 10:37:32 GMT]]></title><description><![CDATA[<p dir="auto">说实话各种测试都是单轮技术问题，在项目稍微复杂时候它还是需要大点的知识权重，否则你就要不断提醒它具体事情怎么做。长链复杂调用，它还是无法取代大尺寸的模型。我在想办法看看怎么分解工作，省钱还是很多人关心的。</p>
]]></description><link>https://lcz.me/post/12726</link><guid isPermaLink="true">https://lcz.me/post/12726</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 18 Aug 2026 10:37:32 GMT</pubDate></item></channel></rss>