<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？]]></title><description><![CDATA[<p dir="auto">如题！<br />
问了一下AI，大概就是8G以下的大小，能达到Qwen3.6 27B Q4左右的效果？！</p>
<p dir="auto">模型链接：<a href="https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf" rel="nofollow ugc">https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf</a></p>
]]></description><link>https://lcz.me/topic/856/大伙儿对最近reddit-localllm社区中很火的ternary-bonsai-27b怎么看</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 22:16:54 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/856.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 15 Jul 2026 07:23:10 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Thu, 16 Jul 2026 23:11:15 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/wwcd2016" aria-label="Profile: wwcd2016">@<bdi>wwcd2016</bdi></a><br />
27B的话可以考虑这个版本，实测和Q4几乎没区别，速度更快，上下文可以开的更大而且还可以开视觉。<br />
IQ 量化非常强，我感觉只要你做的任务不是特别偏门IQ3XXS版本完全可以替代 Q4，但是注意，只有这个 XXS 版本比较好，我试过它的 XS 版本和 S 版本在某些时候下会输出会蹦还不如这个 XXS 的稳定。我这种魔改版的都是第 3 方的这些人。 每天都在那量化好多模型，他们不可能一个一个测试的，所以如果用模改版的模型的话，还是要自己多测试几个版本，找一个最适合的。</p>
<p dir="auto"><a href="https://huggingface.co/mradermacher/Huihui-Qwen3.6-27B-abliterated-i1-GGUF?show_file_info=Huihui-Qwen3.6-27B-abliterated.i1-IQ3_XXS.gguf" rel="nofollow ugc">https://huggingface.co/mradermacher/Huihui-Qwen3.6-27B-abliterated-i1-GGUF?show_file_info=Huihui-Qwen3.6-27B-abliterated.i1-IQ3_XXS.gguf</a></p>
]]></description><link>https://lcz.me/post/10002</link><guid isPermaLink="true">https://lcz.me/post/10002</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Thu, 16 Jul 2026 23:11:15 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Thu, 16 Jul 2026 22:44:52 GMT]]></title><description><![CDATA[<p dir="auto">对了，这是单R9700的数据，7900XTX的话测试速度应该可以在PP2700 TP85左右（因为7900XTX不支持WMMA但是带宽要高30%左右，分别限制预处理和解码）。<br />
所以你要玩双卡的话应该是双7900XTX的组合更好，48G干啥都够了，只不过速度不能完全叠加能有个1.4-1.7倍就到头了，<strong>云的，没测试过。</strong></p>
]]></description><link>https://lcz.me/post/10000</link><guid isPermaLink="true">https://lcz.me/post/10000</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Thu, 16 Jul 2026 22:44:52 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Thu, 16 Jul 2026 22:40:46 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bunsei" aria-label="Profile: Bunsei">@<bdi>Bunsei</bdi></a> 除非时间不敏感或者需要质量（其实用v4 flash更好）的时候在开启27B就行。<br />
我发现这个35B的Ornith Q5好强，对于我的用途几乎都能替代flash了。跑分的PP3800/TP57,实际抓日志看Opencode和Zcode还有Hermes的调用大概在PP2900，TP单发50，双发合计60左右，这个体验就很爽了。</p>
<p dir="auto">本地的模型知识不够多和不够新（其实任何参数的模型都一样），所以这个时间配个好用的搜索引擎就哼重要了，付费的EXA或者Searngx（我本地配的质量能到exa的80%左右）。<br />
一套下来体验很好。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/20a141d8-dafa-41d1-b57e-c855af231e8a.jpeg" alt="aab76594-307d-4abe-b3c7-109f60d82d31-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/9999</link><guid isPermaLink="true">https://lcz.me/post/9999</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Thu, 16 Jul 2026 22:40:46 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Thu, 16 Jul 2026 09:07:55 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> 我还在考虑要不要买第二块R9700呢，一块R9700长上下文prefill的速度太要命了。</p>
]]></description><link>https://lcz.me/post/9980</link><guid isPermaLink="true">https://lcz.me/post/9980</guid><dc:creator><![CDATA[Bunsei]]></dc:creator><pubDate>Thu, 16 Jul 2026 09:07:55 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Thu, 16 Jul 2026 09:02:35 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/wwcd2016" aria-label="Profile: wwcd2016">@<bdi>wwcd2016</bdi></a></p>
<p dir="auto">核心的参数如下，我是成功开启WMMA以后速度才狂飙起来的（以前也就2200左右，R9700），PP跑分能到3800tps，tp57左右，好爽。实际用Agent调用大概在2800tps，tp单槽49，双槽并行总共大概60tps。<br />
记得要用对聊天模板，Agent场景这个非常重要。</p>
<p dir="auto">模型与模板：<br />
Bash<br />
<br />
<br />
--model    /models/Ornith-1.0-35B-Q5_K_M.gguf<br />
--mmproj   /models/mmproj-35B-A3B-Q8_0.gguf<br />
--alias    Ornith-35B-WMMA<br />
--chat-template-file  /app/chat_template_ornith.jinja</p>
<p dir="auto">服务端：<br />
copy<br />
<br />
<br />
--host    0.0.0.0<br />
--port    8080</p>
<p dir="auto">KV Cache：<br />
copy<br />
<br />
<br />
--cache-type-k  q8_0<br />
--cache-type-v  q8_0<br />
--ctx-size      524288     ← 512K</p>
<p dir="auto">批处理：<br />
copy<br />
<br />
<br />
--batch-size    1024<br />
--ubatch-size   1024<br />
--flash-attn    on<br />
--n-gpu-layers  99</p>
<p dir="auto">并行槽位：<br />
copy<br />
<br />
<br />
--parallel       2<br />
--slot-prompt-similarity  0.50<br />
--cache-reuse    1</p>
<p dir="auto">采样参数：<br />
copy<br />
<br />
<br />
--temp             0.6<br />
--top-p            0.95<br />
--top-k            20<br />
--repeat-penalty   1.1<br />
--frequency-penalty  0.1<br />
--predict          8192</p>
<p dir="auto">视觉：<br />
copy<br />
<br />
<br />
--image-min-tokens  1536</p>
<p dir="auto">性能 Tuning：<br />
copy<br />
<br />
<br />
--poll        100<br />
--poll-batch  1<br />
--prio-batch  3</p>
]]></description><link>https://lcz.me/post/9979</link><guid isPermaLink="true">https://lcz.me/post/9979</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Thu, 16 Jul 2026 09:02:35 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Thu, 16 Jul 2026 08:50:22 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> 求你的具体模型名，启动参数。麻烦你了。我是觉得35b速度快，但是编程全是乱的。我主要是用来做qmt量化。</p>
]]></description><link>https://lcz.me/post/9978</link><guid isPermaLink="true">https://lcz.me/post/9978</guid><dc:creator><![CDATA[wwcd2016]]></dc:creator><pubDate>Thu, 16 Jul 2026 08:50:22 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 23:57:42 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/wwcd2016" aria-label="Profile: wwcd2016">@<bdi>wwcd2016</bdi></a><br />
你可以试试 K V 量化开到 4，然后，如果开了 MTP 的话，MTP 头那一部分的计算量化也可以开到 Q8这样子就又可以省一下显存了。24G 显存的卡确实 27B 更强，但是如果你的卡显存更大的话，35B 用高粱话的版本其实是更好的选择，速度能快个三倍多。35B 这个 MOE 处理系统提示词之类的超级快，我自己折腾了好久，最后实测PP能到3800 tps，比硬磕 27B 舒服太多了</p>
]]></description><link>https://lcz.me/post/9962</link><guid isPermaLink="true">https://lcz.me/post/9962</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Wed, 15 Jul 2026 23:57:42 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 22:55:59 GMT]]></title><description><![CDATA[<p dir="auto">qwen 3.6 27b 稳定性好。速度还行。<br />
跟着论坛用了3090  长上下文的。148k 稳定驱动 hermes 。没有问题。<br />
但vison 模式 只有50k上下文，目前还没有解。</p>
]]></description><link>https://lcz.me/post/9960</link><guid isPermaLink="true">https://lcz.me/post/9960</guid><dc:creator><![CDATA[wwcd2016]]></dc:creator><pubDate>Wed, 15 Jul 2026 22:55:59 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 14:50:30 GMT]]></title><description><![CDATA[<p dir="auto">27b更强，无论哪个领域，就是慢点</p>
]]></description><link>https://lcz.me/post/9956</link><guid isPermaLink="true">https://lcz.me/post/9956</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 15 Jul 2026 14:50:30 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 14:20:44 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/9939">说</a>:</p>
<p dir="auto">本地老老实实Qwen3.6 27b</p>
</blockquote>
<p dir="auto">本地應該考慮 ,Qwen3.6 27b 還是 Qwen3.6-35B-A3B呢? 兩者都是24GB VRAM能吞下的。 多模態來說， 是密集裹是專家模型更強? 問了AI， 它也說不準</p>
]]></description><link>https://lcz.me/post/9954</link><guid isPermaLink="true">https://lcz.me/post/9954</guid><dc:creator><![CDATA[exe127]]></dc:creator><pubDate>Wed, 15 Jul 2026 14:20:44 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 10:37:11 GMT]]></title><description><![CDATA[<p dir="auto">说实话，本地老老实实Qwen3.6 27b，跟着版本走就对了，未来很长一段时间，Qwen就是本地天花板代名词，小作坊是不可能干的过阿里这样的大企业的。在线就是Deepseek V4 Flash，高端的就上A O两家的模型，过分折腾耽误做事。</p>
]]></description><link>https://lcz.me/post/9939</link><guid isPermaLink="true">https://lcz.me/post/9939</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 15 Jul 2026 10:37:11 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 10:06:37 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bunsei" aria-label="Profile: Bunsei">@<bdi>Bunsei</bdi></a></p>
<p dir="auto">雖然運用Hybrid Attention分別壓縮重要跟不重要的Layer到1 BPW是很厲害沒錯啦, 不過個人覺得有點喧嘩取眾</p>
<p dir="auto">模型的BPW對於輸出跟Tool Call是很重要的, 這也是為什麼BPW低的27B更容易陷入Thinking Loop跟Tool Call失敗</p>
<p dir="auto"><a href="https://ai-coding.wiselychen.com/bonsai-27b-qwen36-compression-local-inference/" rel="nofollow ugc">根據這個評測</a>, 單是Tool Call上多20%的失敗就足夠判死刑</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/fd365603-ac9c-47b3-a3e6-958f0edaf8c7.jpeg" alt="4e3ffbd3-1bbf-4d63-97d9-43e49ef5b48b-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/9935</link><guid isPermaLink="true">https://lcz.me/post/9935</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Wed, 15 Jul 2026 10:06:37 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 07:50:50 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> 也就是说实际效果可能更差，可能只有部分任务的情况下才能达到宣称的效果？可惜运行需要自己编译llama.cpp，再加上我的卡是R9700 ROCm，之前一阵折腾，才搞好。跑这个实际测试一遍对我来说还是有难度的。 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f915.png?v=1376b21da6c" class="not-responsive emoji emoji-android emoji--face_with_head_bandage" style="height:23px;width:auto;vertical-align:middle" title=":face_with_head_bandage:" alt="🤕" /></p>
]]></description><link>https://lcz.me/post/9926</link><guid isPermaLink="true">https://lcz.me/post/9926</guid><dc:creator><![CDATA[Bunsei]]></dc:creator><pubDate>Wed, 15 Jul 2026 07:50:50 GMT</pubDate></item><item><title><![CDATA[Reply to 大伙儿对最近Reddit LocalLLM社区中很火的Ternary Bonsai 27B怎么看？ on Wed, 15 Jul 2026 07:46:03 GMT]]></title><description><![CDATA[<p dir="auto">看了下，作者自述的性能如下：</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/2ff9b82d-ed75-4914-a051-634dcdacd7cd.jpeg" alt="2705356b-ba45-47c3-ba01-50b1389d0939-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">但他这个性能描述有很大的不合理之处。至少是benchmark的选型不合理。<br />
他自述Gemma-4-31B FP16的能力是Qwen3.6-27B FP16的 99.4%。但任何benchmark网站都不支持此结论。<br />
比如：<a href="https://artificialanalysis.ai/models/comparisons/qwen3-6-27b-vs-gemma-4-31b" rel="nofollow ugc">https://artificialanalysis.ai/models/comparisons/qwen3-6-27b-vs-gemma-4-31b</a></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/bbfd42ef-dcda-43ac-a3cd-404dc2ba71da.jpeg" alt="d5fb2e3f-5af0-4273-b50d-be50222a89f9-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">基于此，作者对于其27B模型是Qwen3.6-27B FP16的94.6%的结论至少是值得探讨的。</p>
]]></description><link>https://lcz.me/post/9925</link><guid isPermaLink="true">https://lcz.me/post/9925</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Wed, 15 Jul 2026 07:46:03 GMT</pubDate></item></channel></rss>