<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？]]></title><description><![CDATA[<p dir="auto">电脑配置是MacBook pro M5 MAX 128G，因为不懂编程，一直使用deepseek v4 flash通过hermes做一些金融小模型的训练和量化编程，截止目前都很不错。</p>
<p dir="auto">介于ds开始涨价，而自己平常使用的token确实有点儿多，加之对LLM的专业能力范围不大，就想使用Qwen3.8-27B-MLX-4bit本地化部署，通过oMLX让hermes调用。结果发现在oMLX对话页面的速度是很好的，但是一连接上hermes就降速厉害，半天没个反应，速度严重拉胯。</p>
<p dir="auto">真心请教各位，该如何解决hermes调用本地模型严重降速的问题，另外，oMLX如何优化配置Qwen3.8-27B这个模型呢？我对上下文的长度有刚性需求，也不知道如何调整，先谢谢各位了！</p>
]]></description><link>https://lcz.me/topic/1175/真心求教hermes-agent运行本地omlx的qwen3.8-27b-mlx-4bit如何提速</link><generator>RSS for Node</generator><lastBuildDate>Sat, 22 Aug 2026 03:26:47 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1175.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 18 Aug 2026 03:05:02 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Wed, 19 Aug 2026 08:06:02 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/stxpnet" aria-label="Profile: stxpnet">@<bdi>stxpnet</bdi></a> <a href="/post/12869">说</a>:</p>
<p dir="auto">我的建议是，让AI上谷姐和REDDIT去帮你找 大本营，总有大神会和你用同款机器，然后他会精调一个专门针对这个机型的框架和模型,然后你接个deepseek v4,设计 好提示词，让梁文峰直接帮你抄作业就完事了。 我等了几天已经等到了。  以下给你一个HERMES可用的提示词：</p>
<p dir="auto">去谷歌和reddit 帮我调研一下，最近7天内的，关于qwen 3.8 27b的模型比如（<a href="http://huggingface.co/zhnagchenchne/Qwen3.8-27B-GPTQ-W4A16" rel="nofollow ugc">huggingface.co/zhnagchenchne/Qwen3.8-27B-GPTQ-W4A16</a> ），看看有哪些比较适合MacBook pro M5 MAX 128G 的，列最适合的10个，并用表格说明各自优劣势 ，我目前用的这个是 <a href="http://huggingface.co/twolven/Qwen3.8-27B-abliterated-AWQ-MTP" rel="nofollow ugc">huggingface.co/twolven/Qwen3.8-27B-abliterated-AWQ-MTP</a> ，感觉还行，但是离完美还有一些距离. 我比较喜欢【什么样的模型，用途主要是XXXX 】的。 如果reddit上找不到相关内容的话，可以从 <a href="http://huggingface.co/models?p=1&amp;sort=created&amp;search=qwen+3.8+27b" rel="nofollow ugc">huggingface.co/models?p=1&amp;sort=created&amp;search=qwen+3.8+27b</a> 这个作为起点进去找找。 本会话可能要做为每日常规任务。主要目标是发掘一个适合 我本地 MacBook pro M5 MAX 128G 硬件的vllm或者sglang模型及框架 .</p>
</blockquote>
<p dir="auto"><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f44d.png?v=60716d54ab2" class="not-responsive emoji emoji-android emoji--+1" style="height:23px;width:auto;vertical-align:middle" title=":+1:" alt="👍" />  非常感谢，这个也很有帮助。。。马上办！</p>
]]></description><link>https://lcz.me/post/12883</link><guid isPermaLink="true">https://lcz.me/post/12883</guid><dc:creator><![CDATA[Bin Mao]]></dc:creator><pubDate>Wed, 19 Aug 2026 08:06:02 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Wed, 19 Aug 2026 06:22:48 GMT]]></title><description><![CDATA[<p dir="auto">我的建议是，让AI上谷姐和REDDIT去帮你找 大本营，总有大神会和你用同款机器，然后他会精调一个专门针对这个机型的框架和模型,然后你接个deepseek v4,设计 好提示词，让梁文峰直接帮你抄作业就完事了。 我等了几天已经等到了。  以下给你一个HERMES可用的提示词：</p>
<p dir="auto">去谷歌和reddit 帮我调研一下，最近7天内的，关于qwen 3.8 27b的模型比如（<a href="http://huggingface.co/zhnagchenchne/Qwen3.8-27B-GPTQ-W4A16" rel="nofollow ugc">huggingface.co/zhnagchenchne/Qwen3.8-27B-GPTQ-W4A16</a> ），看看有哪些比较适合MacBook pro M5 MAX 128G 的，列最适合的10个，并用表格说明各自优劣势 ，我目前用的这个是 <a href="http://huggingface.co/twolven/Qwen3.8-27B-abliterated-AWQ-MTP" rel="nofollow ugc">huggingface.co/twolven/Qwen3.8-27B-abliterated-AWQ-MTP</a> ，感觉还行，但是离完美还有一些距离. 我比较喜欢【什么样的模型，用途主要是XXXX 】的。 如果reddit上找不到相关内容的话，可以从 <a href="http://huggingface.co/models?p=1&amp;sort=created&amp;search=qwen+3.8+27b" rel="nofollow ugc">huggingface.co/models?p=1&amp;sort=created&amp;search=qwen+3.8+27b</a> 这个作为起点进去找找。 本会话可能要做为每日常规任务。主要目标是发掘一个适合 我本地 MacBook pro M5 MAX 128G 硬件的vllm或者sglang模型及框架 .</p>
]]></description><link>https://lcz.me/post/12869</link><guid isPermaLink="true">https://lcz.me/post/12869</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Wed, 19 Aug 2026 06:22:48 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Wed, 19 Aug 2026 06:17:10 GMT]]></title><description><![CDATA[<p dir="auto">只能等 苹果官方搞定了。本身就是个闭源系统。</p>
]]></description><link>https://lcz.me/post/12867</link><guid isPermaLink="true">https://lcz.me/post/12867</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Wed, 19 Aug 2026 06:17:10 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Wed, 19 Aug 2026 05:44:58 GMT]]></title><description><![CDATA[<p dir="auto">谢谢各位的解答！按照xiaote的模式，我暂时应用起来了，慢点就慢点，熬着试试再说，实在不行继续用ds v4 flash外加agnes 2.5 flash，再次感谢！</p>
]]></description><link>https://lcz.me/post/12863</link><guid isPermaLink="true">https://lcz.me/post/12863</guid><dc:creator><![CDATA[Bin Mao]]></dc:creator><pubDate>Wed, 19 Aug 2026 05:44:58 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Wed, 19 Aug 2026 04:17:23 GMT]]></title><description><![CDATA[<p dir="auto">小特说的“关思考”这步可以好好试一下，另外 Mac 本来就是个偏科生，这么低的功耗下，以纯学习为主，兼顾娱乐和影音， tok/s 能到 30 以上就不错了<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f601.png?v=60716d54ab2" class="not-responsive emoji emoji-android emoji--grin" style="height:23px;width:auto;vertical-align:middle" title="😁" alt="😁" />别指望太多</p>
]]></description><link>https://lcz.me/post/12850</link><guid isPermaLink="true">https://lcz.me/post/12850</guid><dc:creator><![CDATA[老茶]]></dc:creator><pubDate>Wed, 19 Aug 2026 04:17:23 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Wed, 19 Aug 2026 03:10:32 GMT]]></title><description><![CDATA[<p dir="auto">现在才醒觉，会不会太晚了呢？当时用DS 4 flash的时候就没想过它会涨价吗？你觉得它那么便宜是合理的吗？能一直持续下去吗？现在能用得爽的卡 已经很贵很贵了，都买不起了</p>
]]></description><link>https://lcz.me/post/12844</link><guid isPermaLink="true">https://lcz.me/post/12844</guid><dc:creator><![CDATA[Mediali Li]]></dc:creator><pubDate>Wed, 19 Aug 2026 03:10:32 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Tue, 18 Aug 2026 04:16:15 GMT]]></title><description><![CDATA[<p dir="auto">kop wang 说的根因是对的：统一内存带宽决定 prefill 速度，Mac 上跑 27B 的 prefill 就是慢。但"无解"太绝对了，有几个真实可用的杠杆，按性价比排序：</p>
<ol>
<li>
<p dir="auto">先分清慢在哪：Hermes 每轮请求都带一套固定系统提示词（约 8-10K token），新会话第一轮必须整体 prefill 一遍；agent 用法下工具调用结果不断拼进上下文，前缀一变又得重新 prefill——这才是"半天没反应"的真相。oMLX 对话页快，是因为单轮短输入、前缀稳定。</p>
</li>
<li>
<p dir="auto">上下文配置：你有刚性需求——oMLX/Ollama 侧把 num_ctx 设到 65536（64K），Hermes 侧 context_length 也配 65536，两边对齐。Hermes 的 64K 是启动硬门槛，不能低于它；但也别开 128K+，KV cache 越大 prefill 越贵。32-64K 是本地 agent 的甜点区。</p>
</li>
<li>
<p dir="auto">关思考（提速最狠的一步）：Qwen3.8 默认开 thinking 会疯狂刷思考链，token 量直接翻几倍。装 froggeric/Qwen-Fixed-Chat-Templates 的新模板（论坛 TID:1151 有分享），思考档位锁 low/medium，简单任务直接 enable_thinking=false。agent 场景千万别用 xhigh。</p>
</li>
<li>
<p dir="auto">模型保活：把 keep_alive 设长（或常驻），避免每轮重新加载——MLX 加载 27B Q4 也要十几秒，频繁加载等于白等。</p>
</li>
<li>
<p dir="auto">务实路线：本地 27B 在 M5 MAX 上解码大概 25-35 t/s（带宽 ÷ 权重大小，量级跑不掉），日常问答、短任务完全够用；但"金融模型训练/量化编程"这种 agent 长任务，上下文反复重排，Mac 带宽就是硬伤。建议混合用：日常短对话走本地，复杂 agent 任务继续 DeepSeek API——涨价后的 Flash 在总成本上依然干不过官方（kop wang 引过 opencode Go 的原话：总成本没人能干过原厂）。别二选一，各干各擅长的。</p>
</li>
</ol>
]]></description><link>https://lcz.me/post/12662</link><guid isPermaLink="true">https://lcz.me/post/12662</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 18 Aug 2026 04:16:15 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Tue, 18 Aug 2026 03:46:51 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> <a href="/post/12646">说</a>:</p>
<p dir="auto">“该如何解决hermes调用本地模型严重降速的问题”</p>
<p dir="auto">这就是统一内存的内存带宽太小导致prefill太慢导致的，无解。<br />
hermes的冷启动提示词大概在10k token左右，1000t/s的prefill能力，就需要罚站10秒。更何况M5 MAX 彪不到1000t/s<br />
而且prefill会随着上下文的占用线性变慢。</p>
</blockquote>
<p dir="auto">那这岂不是只能继续用云端模型了？！哎，天了噜！！！</p>
]]></description><link>https://lcz.me/post/12654</link><guid isPermaLink="true">https://lcz.me/post/12654</guid><dc:creator><![CDATA[Bin Mao]]></dc:creator><pubDate>Tue, 18 Aug 2026 03:46:51 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Tue, 18 Aug 2026 03:29:08 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> 看来苹果玩ai是真不行</p>
]]></description><link>https://lcz.me/post/12649</link><guid isPermaLink="true">https://lcz.me/post/12649</guid><dc:creator><![CDATA[dong wei]]></dc:creator><pubDate>Tue, 18 Aug 2026 03:29:08 GMT</pubDate></item><item><title><![CDATA[Reply to 真心求教Hermes agent运行本地oMLX的Qwen3.8-27B-MLX-4bit如何提速？ on Tue, 18 Aug 2026 03:26:39 GMT]]></title><description><![CDATA[<p dir="auto">“该如何解决hermes调用本地模型严重降速的问题”</p>
<p dir="auto">这就是统一内存的内存带宽太小导致prefill太慢导致的，无解。<br />
hermes的冷启动提示词大概在10k token左右，1000t/s的prefill能力，就需要罚站10秒。更何况M5 MAX 彪不到1000t/s<br />
而且prefill会随着上下文的占用线性变慢。</p>
]]></description><link>https://lcz.me/post/12646</link><guid isPermaLink="true">https://lcz.me/post/12646</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Tue, 18 Aug 2026 03:26:39 GMT</pubDate></item></channel></rss>