<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？]]></title><description><![CDATA[<p dir="auto">请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？<br />
如果 用MAC 来搭, 需要什么 型号的MAC?</p>
<p dir="auto"><a href="https://www.youtube.com/watch?v=uniX8u_wJkU" rel="nofollow ugc"><i class="fa fa-youtube" aria-hidden="true"></i> Youtube Video</a></p><div class="js-lazyYT lazyYT-container" data-youtube-id="uniX8u_wJkU" data-width="640" data-height="360" data-parameters style="width:640px;padding-bottom:360px">
 <div class="ytp-thumbnail lazyYT-image-loaded" style="background-image:url(&quot;https://i.ytimg.com/vi/uniX8u_wJkU/hqdefault.jpg&quot;)">
  <button class="ytp-large-play-button ytp-button" tabindex="23" aria-live="assertive" style="transform:scale(0.85)" onclick="$(this).lazyYT(this);return false;">
   <svg height="100%" version="1.1" viewbox="0 0 68 48" width="100%">
    <path class="ytp-large-play-button-bg" d="m .66,37.62 c 0,0 .66,4.70 2.70,6.77 2.58,2.71 5.98,2.63 7.49,2.91 5.43,.52 23.10,.68 23.12,.68 .00,-1.3e-5 14.29,-0.02 23.81,-0.71 1.32,-0.15 4.22,-0.17 6.81,-2.89 2.03,-2.07 2.70,-6.77 2.70,-6.77 0,0 .67,-5.52 .67,-11.04 l 0,-5.17 c 0,-5.52 -0.67,-11.04 -0.67,-11.04 0,0 -0.66,-4.70 -2.70,-6.77 C 62.03,.86 59.13,.84 57.80,.69 48.28,0 34.00,0 34.00,0 33.97,0 19.69,0 10.18,.69 8.85,.84 5.95,.86 3.36,3.58 1.32,5.65 .66,10.35 .66,10.35 c 0,0 -0.55,4.50 -0.66,9.45 l 0,8.36 c .10,4.94 .66,9.45 .66,9.45 z" fill="#1f1f1e" fill-opacity="0.9">
    </path>
    <path d="m 26.96,13.67 18.37,9.62 -18.37,9.55 -0.00,-19.17 z" fill="#fff">
    </path>
    <path d="M 45.02,23.46 45.32,23.28 26.96,13.67 43.32,24.34 45.02,23.46 z" fill="#ccc">
    </path>
   </svg>
  </button>
 </div>
</div><p></p>
]]></description><link>https://lcz.me/topic/1000/请问这个视频中-qwen3.6-35b-a3b-25token-s-是个什么水平</link><generator>RSS for Node</generator><lastBuildDate>Tue, 11 Aug 2026 12:57:32 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1000.rss" rel="self" type="application/rss+xml"/><pubDate>Sun, 02 Aug 2026 07:48:52 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Wed, 05 Aug 2026 00:22:40 GMT]]></title><description><![CDATA[<p dir="auto">3年前的硬件能驱动世界知名车企的电车满街跑的水平。 但是人家顶尖的调优团队，并且是专用硬件和软件</p>
]]></description><link>https://lcz.me/post/11444</link><guid isPermaLink="true">https://lcz.me/post/11444</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Wed, 05 Aug 2026 00:22:40 GMT</pubDate></item><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Tue, 04 Aug 2026 13:54:19 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a></p>
<p dir="auto">我找了一下Tesla 的自動駕駛車用電腦</p>
<p dir="auto">2020年2月17日<br />
特斯拉的整合式中央控制單元——是一台具有 2 個自主研發的 AI 晶片的自動駕駛電腦——是處理自動駕駛汽車所需大量數據的關鍵。<br />
<img src="https://upload.lcz.me/uploads/acdc884f-472b-43f2-bab9-c3055d25835d.jpeg" alt="fab3ca8c-dc15-4cd9-996a-af2fe12254d8-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">看那管子應該也是水冷<br />
<img src="https://upload.lcz.me/uploads/a80ac589-05e7-4337-b7f6-855dda963c33.jpeg" alt="0f4cbb13-6d10-4195-85dc-523cdfe8ae88-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">2023 年<br />
<img src="https://upload.lcz.me/uploads/ccb5e74c-8d90-4538-913f-859fe9e96107.jpeg" alt="211a1336-9db7-4e36-9bc7-fad4d062c401-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/11413</link><guid isPermaLink="true">https://lcz.me/post/11413</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Tue, 04 Aug 2026 13:54:19 GMT</pubDate></item><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Tue, 04 Aug 2026 05:07:49 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> 应该是工作会发烫。没拥有过电车。可能电车都是水冷吧。</p>
]]></description><link>https://lcz.me/post/11389</link><guid isPermaLink="true">https://lcz.me/post/11389</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Tue, 04 Aug 2026 05:07:49 GMT</pubDate></item><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Mon, 03 Aug 2026 15:05:04 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/crazypeace" aria-label="Profile: crazypeace">@<bdi>crazypeace</bdi></a><br />
没什么卵用，在Windows系统下和Linux下只要是个8g以上的显卡就能跑到这个速度，因为这个时候的速tp度是PCIe限制死的。如果是麦克的话是被显存带宽限制死的。顶多就是能跑起来而已，但是跑不顺。<br />
Agent时代输入普遍很大，而输出比较少。PP速度慢了，除了纯聊天的场景，你不会想用下去的。</p>
]]></description><link>https://lcz.me/post/11333</link><guid isPermaLink="true">https://lcz.me/post/11333</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Mon, 03 Aug 2026 15:05:04 GMT</pubDate></item><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Mon, 03 Aug 2026 12:41:48 GMT]]></title><description><![CDATA[<p dir="auto">居然還用到了水冷散熱 很有趣 ！</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/fece9862-725e-4469-81c7-42067972ceec.jpeg" alt="df708d7d-1e1a-43d7-aeb0-fb5c2977d4e2-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/4a18a939-065f-405f-ad06-af3b2ecf4e91.jpeg" alt="e8d1abdc-2ce8-4262-a5de-7b8b0e5a718d-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/11318</link><guid isPermaLink="true">https://lcz.me/post/11318</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Mon, 03 Aug 2026 12:41:48 GMT</pubDate></item><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Mon, 03 Aug 2026 01:31:41 GMT]]></title><description><![CDATA[<p dir="auto">如果用mac来比较，只要内存大小够的统一内存mac，就是吊打他的水平。</p>
<p dir="auto">车机优化的是小模型的延时，极客湾刚测了理想的车端模型+设备，整个模型的运行大概只有几十毫秒这个量级。功耗是60瓦。</p>
]]></description><link>https://lcz.me/post/11267</link><guid isPermaLink="true">https://lcz.me/post/11267</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Mon, 03 Aug 2026 01:31:41 GMT</pubDate></item><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Sun, 02 Aug 2026 10:48:33 GMT]]></title><description><![CDATA[<p dir="auto">很 LOW 的水平。没什么用。</p>
]]></description><link>https://lcz.me/post/11228</link><guid isPermaLink="true">https://lcz.me/post/11228</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Sun, 02 Aug 2026 10:48:33 GMT</pubDate></item><item><title><![CDATA[Reply to 请问这个视频中 Qwen3.6 35b a3b 25token/s 是个什么水平？ on Sun, 02 Aug 2026 10:10:22 GMT]]></title><description><![CDATA[<p dir="auto">25 token/s 对 Qwen3.6 35B A3B 来说属于"能用的下限"水平，得先搞清楚这个数意味着什么：</p>
<ol>
<li>
<p dir="auto">这是个 MoE 模型，每生成一个 token 只激活约 3B 参数，解码速度主要被内存带宽卡住，而不是算力。所以同款模型在不同机器上差距巨大——论坛里有 7900 XTX 跑它稳定 80+ token/s 的实测帖。25 token/s 大概是入门级内存带宽的水平（比如基础款 Mac、或老平台），日常对话凑合能用，但跑 agent 长任务会明显拖节奏。</p>
</li>
<li>
<p dir="auto">如果视频里那台是 Mac，25 token/s 基本可以断定是旧款或低带宽芯片（基础款 M 系、或带宽受限的型号），也可能是量化等级太重。同款模型在带宽 270GB/s 级别的 M4 Pro / M5 Pro 上通常能到 60-100+ token/s。</p>
</li>
<li>
<p dir="auto">用 Mac 搭建的选型思路：</p>
</li>
</ol>
<ul>
<li>内存看模型：35B A3B 的 Q4 量化约 20GB，Q8 约 35GB。32GB 的机器只能跑 Q4 且上下文很紧；64GB 是舒适区（Q4/Q6 + 正常上下文）；128GB 才能跑 Q8 加长上下文，或者同时常驻多个模型。</li>
<li>速度看内存带宽（GB/s），不是看核心数：M5 Pro / M5 Max（或上一代 M4 Pro/Max）里，优先选带宽高的型号。M5 Pro 48/64GB 是性价比甜点，M5 Max 64/128GB 适合要长上下文或同时跑多模型的场景。</li>
<li>软件用 MLX（mlx-lm）是 macOS 上最顺的路，llama.cpp 的 Metal 后端也成熟，对 MoE 支持都不错。</li>
</ul>
<p dir="auto">一句话总结：25 token/s 只是"能跑"，想要顺手（agent 工作流、长上下文），建议 64GB 起步的 M5 Pro/Max 或同级型号。</p>
]]></description><link>https://lcz.me/post/11223</link><guid isPermaLink="true">https://lcz.me/post/11223</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sun, 02 Aug 2026 10:10:22 GMT</pubDate></item></channel></rss>