<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/b47acd21-8749-491a-acc5-bf0dd7426bb5.jpeg" alt="2c1c0705-5d39-45da-880c-8b396146864c-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">这个模型支持多模态，所以日常被我拿来用于工作，速度挺快的，也能持续干活，<br />
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ4_NL.gguf<br />
启动参数</p>
<pre><code>./llama-cli -m /path/to/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ4_NL.gguf \
  --mmproj /path/to/mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf \
  -ngl 99 \
  -c 131072 \
  -fa \
  --batch-size 512 \
  -t 8 \
  --temp 0.7 \
  --repeat-penalty 1.1
</code></pre>
]]></description><link>https://lcz.me/topic/792/实测-7900xtx使用qwen3.6-35b-a3b速度稳定在80-t-s</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 20:02:15 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/792.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 07 Jul 2026 02:27:08 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Sun, 26 Jul 2026 13:07:17 GMT]]></title><description><![CDATA[<p dir="auto">不信的只有自己实践了才能信，35B,A3B，大概就是6B的速度，质量嘛，大概15B。 其实有个简单的判断方法，你给它配好harnness工具，然后让它写一些复杂点的小游戏。然后观察nvtop曲线。 就算你调好参数，让它能满载跑，它写程序也是写100行，过一会儿删80行， 电力和时间就这样被浪费了。  它只适合做一些简单的任务编排，但是这样的任务,deepseek flash就能做了，也便宜。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/5e14f9c2-0fbf-4d55-93c8-360fc1a4924e.jpeg" alt="50e83ee5-0d5b-4cb2-887e-4769b475a92d-image.jpeg" class=" img-fluid img-markdown" /><br />
这个模型唯一优点的就是刚开始像打了鸡血一样快。140T/S，中后期会掉到80，比较难受。 （我一直用0.6温度， 不然后期智力下降厉害） 。</p>
]]></description><link>https://lcz.me/post/10604</link><guid isPermaLink="true">https://lcz.me/post/10604</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Sun, 26 Jul 2026 13:07:17 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Sat, 25 Jul 2026 17:10:15 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/rex" aria-label="Profile: Rex">@<bdi>Rex</bdi></a> 强调了很多次，35B A3B 干不了什么正儿八经的活，小玩玩可以。但大家似乎都不太相信，不知道原因是什么。</p>
]]></description><link>https://lcz.me/post/10546</link><guid isPermaLink="true">https://lcz.me/post/10546</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sat, 25 Jul 2026 17:10:15 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Sat, 25 Jul 2026 11:07:26 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stormaround" aria-label="Profile: stormaround">@<bdi>stormaround</bdi></a> 同感，我用35B-A3A写一个代码项目，他干了5个小时，修了上百行代码。 claude code评价：5个小时干的活没有正向做功，因为没有找到根本问题，在现象层面修了东边，坏了西边，反过来修西边，又坏了东边。直到手动停止它</p>
]]></description><link>https://lcz.me/post/10509</link><guid isPermaLink="true">https://lcz.me/post/10509</guid><dc:creator><![CDATA[Rex]]></dc:creator><pubDate>Sat, 25 Jul 2026 11:07:26 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Wed, 22 Jul 2026 08:27:48 GMT]]></title><description><![CDATA[<p dir="auto">轻度聊天非常舒服的，我的单卡3090速度有80-100token/s,用完就关了，比在线的省点token,比本地qwen 3.6 27b ttft快4-5倍。  楼主这个参数没指定，似乎就F16的K V CACHE，有点牛，多轮可能会爆显存</p>
]]></description><link>https://lcz.me/post/9893</link><guid isPermaLink="true">https://lcz.me/post/9893</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Wed, 22 Jul 2026 08:27:48 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Sat, 11 Jul 2026 16:19:02 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stormaround" aria-label="Profile: stormaround">@<bdi>stormaround</bdi></a><br />
我也发现这个问题了，试了不下10个这个基座模型的变种，有的量化精度低一些的在Hermes调用几步以后就开始陷入死循环，良化进度高一点的（新入了r9700）会好一些，但是他的思维会陷入一个死循环，就是不断的尝试，但是每一轮尝试下来都是一样的结果，根本解决不了问题。<br />
后来我把跟他的交互记录让deepseek Pro看了一下，Deepseek 2分钟就发现了问题，根本的原因是Hermes斯的对话管理机制，因为它的定位是做一个轻量的助手，所以他的对话几乎都不能持久，然后你过一会不用的话，那个对话就给你关掉了，你在聊天的时候又是一个新的对话，所以说它不是靠对话来管理上下文的，它实际上是靠搜索它的那个对话的数据库来搞的，这样的话，尤其像这种本地小模型智商没那么高的情况下，他有可能抓到的内容是碎片化不相关的。所以导致使用体验非常糟糕，然后我就把同样的咖啡UI的自动化视频流程让他总结了一遍，迁移到了opencode里面。中间有一个断点在opencode里面调用一个魔改版的这个35b一次跑通！因为open code的一个对话就是一个项目，这种开发类的agent框架的上下文比较干净。<br />
工具和模型同样重要，多么痛的领悟，我已经卡了三四天了了。</p>
]]></description><link>https://lcz.me/post/9754</link><guid isPermaLink="true">https://lcz.me/post/9754</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Sat, 11 Jul 2026 16:19:02 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Thu, 09 Jul 2026 16:50:54 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stormaround" aria-label="Profile: stormaround">@<bdi>stormaround</bdi></a> 开个新的对话就可以了，上下文太长了就得开，claude对于新对话支持的比较好</p>
]]></description><link>https://lcz.me/post/9570</link><guid isPermaLink="true">https://lcz.me/post/9570</guid><dc:creator><![CDATA[koala]]></dc:creator><pubDate>Thu, 09 Jul 2026 16:50:54 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Thu, 09 Jul 2026 16:50:05 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/densha" aria-label="Profile: densha">@<bdi>densha</bdi></a> 在claude里面跑起来效果不错，有时候27B出错的时候我就切换他来跑，</p>
]]></description><link>https://lcz.me/post/9569</link><guid isPermaLink="true">https://lcz.me/post/9569</guid><dc:creator><![CDATA[koala]]></dc:creator><pubDate>Thu, 09 Jul 2026 16:50:05 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Tue, 07 Jul 2026 12:56:18 GMT]]></title><description><![CDATA[<p dir="auto">这个模型跑的很快，用m5max 6bit 能跑将近100toks，但是用hermes调用时候经常会死循环，一直重复几句话，干不了大活</p>
]]></description><link>https://lcz.me/post/9381</link><guid isPermaLink="true">https://lcz.me/post/9381</guid><dc:creator><![CDATA[stormaround]]></dc:creator><pubDate>Tue, 07 Jul 2026 12:56:18 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Tue, 07 Jul 2026 10:36:56 GMT]]></title><description><![CDATA[<p dir="auto">大佬的ctx開 132072這麼大@@，我的顯卡跟您一樣，照抄參數，跑起來很喘20t/s左右，而且系統內存32G也被佔滿了...</p>
]]></description><link>https://lcz.me/post/9370</link><guid isPermaLink="true">https://lcz.me/post/9370</guid><dc:creator><![CDATA[densha]]></dc:creator><pubDate>Tue, 07 Jul 2026 10:36:56 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Tue, 07 Jul 2026 06:37:07 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 做总项目调度，调试软件、工作流，等等。</p>
]]></description><link>https://lcz.me/post/9340</link><guid isPermaLink="true">https://lcz.me/post/9340</guid><dc:creator><![CDATA[koala]]></dc:creator><pubDate>Tue, 07 Jul 2026 06:37:07 GMT</pubDate></item><item><title><![CDATA[Reply to 【实测】7900xtx使用Qwen3.6-35B-A3B速度稳定在80+t/s on Tue, 07 Jul 2026 05:01:20 GMT]]></title><description><![CDATA[<p dir="auto">你用它来干什么呢？多模态是它的最大优势，也正好踩在本地显卡的算力负担甜点上。长上下文怕是显存不太够。</p>
]]></description><link>https://lcz.me/post/9330</link><guid isPermaLink="true">https://lcz.me/post/9330</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 07 Jul 2026 05:01:20 GMT</pubDate></item></channel></rss>