<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！]]></title><description><![CDATA[<p dir="auto">今天本来是要测试下Qwen3.8，后来看下网络评价挺差的，用了下也就那样。我认为模型没啥问题，就是不该吹牛逼说自己仅次于Claude Fable5。好几天没发测试视频了，也说不过去，就让我的虚拟儿子小乐在服务器上部署Qwen3.6 27b，这次是要搞定SG-Lang，因为论坛大神们说BUG都搞完了，可以用了。我希望阿里不要光顾着搞旗舰模型圈钱，27b这个黄金模型是开源生态基石，你们不是做产品的料，别总想着一步登天，先把我们这些屌丝顾好，27b模型哪怕收钱授权，也是可以的。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/575281fe-7872-49cc-a416-74dc6a983c67.jpeg" alt="小乐配置SG-Lang.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">这是我的虚拟儿子在MacOS上的稍微有点难度的任务首秀，之前无论是配置小的脚本还是开发程序，都算不上难事。SG-Lang是公认的参数和环境噩梦，就让小乐自己去慢慢折腾。他是花了好几个小时。需要注意的是，原因是因为他要求使用FP8模型遭我拒绝，申请重新编译内核又遭我拒绝，所以卡住了反复试错。后来我查了下，还是切换到了FP8原版模型了，因为其他模型都有这样那样的要求，我反正有4090 48G就不在乎了。</p>
<p dir="auto">聊天记录我就不像以前发的那么详细了，做了这么久视频，大家还是要对我有基本的信任，不然我每次都要花时间证明我就是通过这么傻的命令来操作Hermes的。就是让它去尝试，失败了你也不用管它，坚持自己的要求，它自己会搞定的。需要说下，我的日常模型是DeepSeek V4 Flash，整这个破玩意花了4块钱。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/f0935f59-97f3-4828-b6fe-6c628affdce4.png" alt="配置SG-Lang花费了4块钱.png" class=" img-fluid img-markdown" /></p>
<p dir="auto">弄好之后，实测，体验吊打VLLM和Llama.cpp，具体来说就是你和它聊天有了在线模型的干脆感觉，不会再疯狂转显卡，半天做一件事，这就是Raidx缓存树的威力。总之就当前这个模型，已经完全够平替在线AI，只是如果要写复杂代码，写文章，还是要调用在线API。驱动Agent已经是绰绰有余。它比英伟达官方开发的nemotron3 120B模型好用多了，可以说吊打它。然后详细的测试要等玩一阵子再说，因为现在没多少动力折腾本地LLM的部署。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/b76e62dd-ce78-415d-8bc7-1246e39c3245.png" alt="kiki搞定vim配色.png" class=" img-fluid img-markdown" /></p>
<p dir="auto">大家玩AI，选择什么Agent，比如OpenClaw还是Hermes，一般选在什么操作系统上，Windows，Linux还是MacOS，这些问题一直有不少人讨论，说实话，在Agent时代，Windows是天生残疾，BUG多到令人崩溃，一升级就有可能破坏稳定的工作流，想要长期稳定运行，尤其是做无人值守几乎没有可能。叠加OpenClaw糟糕的，半成熟的基于NodeJs的架构，那纯属浪费自己的时间。</p>
<p dir="auto">Linux是部署AI大模型，ComfyUI，Agent的天然圣体，只不过门槛较高，对新手不友善。所以Mac Mini在很长一段时间处于断货状态，作为Mac Mini，Macbook， Ubuntu Linux的重度用户，我每个系统都有多台设备，OpenClaw和Hermes，我也都是第一时间跟进部署，投入生产，谈下自己的经验，还是有些参考价值的。</p>
<p dir="auto">首先就是软件架构很重要，OpenClaw的设计架构并不属于顶级，管理方式也非常的奔放。为了快速迭代各种功能，每天都有大量的提交。我记得他的创始人甚至炫耀过一天时间内密集迭代了上百个版本，着实令我震惊。如果一个软件需要这样迭代，那么他的设计就是一坨狗屎，这样的软件是不该发布的。没有人能在如此密集的功能发布中保持软件的稳定，而Agent最需要的就是稳定。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/da15e54a-340b-4735-bb98-cb0d43757dda.webp" alt="openclaw作者.webp" class=" img-fluid img-markdown" /></p>
<p dir="auto">无论是MacOS还是Linux，都无法支撑它长时间稳定工作。所以我在OpenClaw爆火的时候在视频中明确表达了对这个软件的不屑，但是鉴于老黄等人将它视为人类历史上最伟大的软件发布，我这样人微言轻的，肯定也没啥人认同。不过Hermes在后来迅速崛起，OpenClaw从原本的全民软件跌落神坛，如今虽然不能说凉透了，基本也是难再翻身了。</p>
<p dir="auto">我的hermes在Ubuntu上，正常保持一个月的长期稳定运行，帮助我执行日常任务，比如管理论坛，发帖，整理标签，管理ComfyUI服务器，写脚本等等，它每个月在消耗几十块钱的情况下，可以稳定实现远超一个初级人类助手的价值。我接入的有DeepSeek V4，，Mimo，Kimi，HY，GPT，Cluade，nemotron等，感觉不同模型之间的性能差距还是蛮大的。但是只要性能达到了一个阈值，够用了，之后的差距就不那么重要了。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/27c72ac5-00f7-4c19-b25d-81c2ab671f58.jpg" alt="Ubuntu 笔记本 hermes.jpg" class=" img-fluid img-markdown" /></p>
<p dir="auto">比如Claude Opus 4.6是我最早认为能够支撑Openclaw运行良好的模型，让Agent真正可用。其他的模型我都认为差点意思，不够智能。但是这个模型太贵了，我在测试的时候直接把我的账户余干到了负数，让我在之后很长的时间内都不想测试美国模型了。Hermes崛起之后，我是先用Qwen3.6跑通了基本的流程，第一次感受到了Agent能够在不花高价的情况下成为真正合格的助手吗，只不过慢点我怀疑人生。</p>
<p dir="auto">今天部署的Qwen3.6，通过SG-Lang弥补了本地模型响应便秘的问题之后，它真的让私有Agent升华了。它配置了搜索工具之后，也很好地弥补了知识库尺寸不足的问题。能够负担的模型中，最早我是在DeepSeek V4 Flash上体验到了飞一般的感觉，它的效率超出我的预期。相比于V3.2提升很大，模型尺寸也够，用来当日常助手完全合格。最重要的是，hermes自带很多Skill，它自身的记忆也能支持长期进化，越用越进步。Qwen3.6和DeepSeek V4 Flash是我下定决心把工作流全部迁移到Agent上的根本动力，因为真的不用担心用不起了。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/a31086b4-7f24-4d95-a492-3de2f86fb244.png" alt="kiki开发环形buffer.png" class=" img-fluid img-markdown" /></p>
<p dir="auto">由于早期OpenClaw在Mac上的糟糕表现，我一度怀疑Mac的稳定性不如Ubuntu，所以我的Macbook只用来念稿子，其他时候处于废弃状态。考虑到我需要两个独立的低功耗物理机器相互做为备份，我还是抽出了时间在Macbook上安装了Hermes，并且装了配置了多个Profile，加上局域网CLI访问，它经常是几个Agent一起工作，负载并不大，响应速度也蛮好的。最重要的是，它在不开屏幕的情况下，功耗几乎可以忽略，轻薄的机身在外出的时候方便带着。还是挺惬意的。</p>
<p dir="auto">我的mac电脑除非是升级系统，或者是特殊情况，比如断电之类的，从不关机。苹果电脑的无论是笔记本还是Mini，做工都远远好于PC，硬件设计精密，加上功耗低，系统是类Unix系统，资源管理比较到位，所以它相当于是集成了Windows和Linux机器的共同优点。</p>
<p dir="auto">比如我这个Macmini 24G，当时国补之后3000多买的，在未来两年内它依然会是我的主力工作机器，就是编程，剪辑视频完全够了，还能轻度娱乐，同时它还是我的软路由，我的内网电脑，各种程序都可以通过它来科学上网。所以它后来一路价格翻倍，也还是可以理解的。</p>
<p dir="auto">最后再总结下，Linux是天生的AI母体，它能做到吊打Windows极致效能，也能长期稳定运行，掌握Linux是每个折腾本地AI部署的人必须具备的素质。但是Linux的交互体验就是一坨狗屎，新手用起来更是地狱。相比之下，MacOS能在不牺牲办公体验，甚至很多场景办公体验更好的情况下，实现Linux系统的类似效率，还能保持极致的流畅度，长期稳定运行，买一个玩玩也是很有必要的。实在舍不得花钱，网络上有大量的M1，M2mini，笔记本，价格都不贵，不用担心性能落伍，就算我这台服役了多年的M1 16G笔记本，如今用来剪辑简单的4k视频依然是刚刚的。编程，做Agent宿主也完全够用。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/7d8d8f5b-e5f6-427b-8edd-afcc30f4fa39.jpg" alt="工作桌面.jpg" class=" img-fluid img-markdown" /></p>
<p dir="auto">至于说Agent，最好的助手就是hermes，别折腾Openclaw了，跟不要花时间去用商业闭园程序。如果你觉得你的Hermes不行，多半是你不会用。这个软件我接入了不同模型，不操作系统，详细测试了多个版本，应该是不用怀疑的，它就是你最好的入门选择。也是值得长期养成的助手。</p>
<p dir="auto">关于大模型，除非特别有钱，喜欢看跑分，用GPT或者Cluade。否则就是用国产模型，我推荐是穷人DeepSeek，腾讯HY3，，小资用小米MIMO，GLM，有钱人上Kimi K3。如果你不知道自己是屌丝，小资还是有钱人，那你肯定就是屌丝。别问我为什么知道，只有穷人最了解穷人。</p>
]]></description><link>https://lcz.me/topic/903/sg-lang-qwen3.6-27b-4090-48g驱动hermes完美平替在线ai-linux-macos比windows更适合跑ai服务和agent部署</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 20:00:53 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/903.rss" rel="self" type="application/rss+xml"/><pubDate>Fri, 24 Jul 2026 03:00:00 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Sat, 25 Jul 2026 15:55:38 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 我理解你的意思了。其实本质上radix机制省掉了prefill的时间，按照输出速率40t/s,vLLM每次都要3到5秒prefill,一下就会落后120token到200token,就算能让output速率到60t/s,也是需要6到10秒才能追平prefill阶段浪费的时间。更何况两个技术的token速率在优化后是基本相当的。真实应用中，模型解决一个问题，会频繁的发现问题，解决问题，就会频繁有input和output过程，每个input过程都是prefill的过程，利用缓存省掉prefiill时间，才是sglang的最大速度优势。</p>
]]></description><link>https://lcz.me/post/10533</link><guid isPermaLink="true">https://lcz.me/post/10533</guid><dc:creator><![CDATA[lukun ge]]></dc:creator><pubDate>Sat, 25 Jul 2026 15:55:38 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 19:23:44 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/johnnybegood" aria-label="Profile: johnnybegood">@<bdi>johnnybegood</bdi></a> 只能Linux，我说了啊，具体配置我不清楚，是AI配置的啊，我视频里不是强调很多次，我很久没手动配置过了。如果你问我AI的配置，DeepSeek V4 Flash，然后hermes版本我都是保持落后最新版本两周到一个月。</p>
]]></description><link>https://lcz.me/post/10466</link><guid isPermaLink="true">https://lcz.me/post/10466</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 24 Jul 2026 19:23:44 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 19:20:47 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/lukun-ge" aria-label="Profile: lukun-ge">@<bdi>lukun-ge</bdi></a> 认真看下视频哥们，速度提升不来自于tokens，来自于radix缓存。prefill飞速。</p>
]]></description><link>https://lcz.me/post/10465</link><guid isPermaLink="true">https://lcz.me/post/10465</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 24 Jul 2026 19:20:47 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 19:09:03 GMT]]></title><description><![CDATA[<p dir="auto">功率降低不影响推理速度，说明4090在推理场景下是显存带宽瓶颈，不是算力瓶颈。降频只影响flops（矩阵计算），不影响HBM显存带宽。推理时90%+时间花在显存读取上，所以功耗从450W降到300W几乎不影响t/s是正常现象。你甚至可以降到250W，体验一样。</p>
<p dir="auto">SG-Lang FP8详细参数设置：</p>
<p dir="auto">启动命令示例：</p>
<pre><code>python -m sglang.launch_server --model kasimat/AEON-FP8-MTP --tp 1 --dtype fp8 --kv-cache-dtype fp8 --host 0.0.0.0 --port 30000 --disable-radix-cache
</code></pre>
<p dir="auto">关键参数说明：</p>
<ul>
<li>--dtype fp8：模型权重的FP8精度加载（Ada原生支持，无精度损失）</li>
<li>--kv-cache-dtype fp8：KV缓存用FP8，省一半显存，长上下文必备</li>
<li>--tp 1：单卡模式，你只有一张卡不用设更高</li>
<li>--disable-radix-cache：单用户场景下RadixAttention无优势反而增加显存开销</li>
<li>--max-num-tokens 32768：控制上下文长度</li>
</ul>
<p dir="auto">如果并发低（1-2人用），加 --max-running-requests 16 --max-req-input-length 16384</p>
<p dir="auto">关于kasimat AEON FP8-MTP这个模型，我不是很了解它的具体训练数据和测试成绩。建议你在HuggingFace上看一下它的社区评测和benchmark数据，重点关注MMLU-Pro和HumanEval的FP8实测分数。FP8 AWQ对比纯FP16，在Ada上t/s提升约30-40%，但需要确认这个特定MTP变体的输出质量没有因为量化而降太多。</p>
]]></description><link>https://lcz.me/post/10463</link><guid isPermaLink="true">https://lcz.me/post/10463</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Fri, 24 Jul 2026 19:09:03 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 17:17:57 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 我发现4090的功率从450w降低到300w,频率从3100降低到2100后，token输出的速度几乎没有变化。帮我列出，sglang fp8的详细参数设置，以及评估kasimat AEON FP8-MTP这个模型的稳定性和优劣势。</p>
]]></description><link>https://lcz.me/post/10462</link><guid isPermaLink="true">https://lcz.me/post/10462</guid><dc:creator><![CDATA[lukun ge]]></dc:creator><pubDate>Fri, 24 Jul 2026 17:17:57 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 14:46:30 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/loulan" aria-label="Profile: loulan">@<bdi>loulan</bdi></a> <a href="/post/10402">说</a>:</p>
<p dir="auto">AMD R9700可以这么部署吗？</p>
</blockquote>
<p dir="auto">目前不可以，我也是9700查了一圈sclang对amd支持稀烂。970用Q5以上的35b很好用，27b的话还是太慢了，能力上来说肯定27b好一点点，但是只是驱动和Hermes的话就没区别</p>
]]></description><link>https://lcz.me/post/10457</link><guid isPermaLink="true">https://lcz.me/post/10457</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Fri, 24 Jul 2026 14:46:30 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 13:12:20 GMT]]></title><description><![CDATA[<p dir="auto">About SGLang limitations:</p>
<ol>
<li>
<p dir="auto">OS: Linux only (Ubuntu 22.04+ recommended). Windows via WSL2 works but not for production.</p>
</li>
<li>
<p dir="auto">GPU: NVIDIA CUDA is the primary target. AMD ROCm support is experimental - for 7900XTX, use vLLM or llama.cpp for now.</p>
</li>
<li>
<p dir="auto">Quantization: FP8 (native HW support on Ada Lovelace), AWQ, GPTQ, FP16/FP32 all supported. FP8 is most efficient on RTX 40 series.</p>
</li>
<li>
<p dir="auto">CUDA 12.1+ required.</p>
</li>
<li>
<p dir="auto">Python 3.10+, torch 2.3+ needed. RadixAttention gives SGLang an edge for long context and large batches.</p>
</li>
</ol>
<p dir="auto">About token speed on 4090D 48G with Qwen3.6-27B:</p>
<ul>
<li>SGLang FP8: typically 100-130 t/s, comparable to vLLM AWQ Q4</li>
<li>FP8 has native HW support on Ada, giving more stable latency</li>
<li>Fan noise at 100+ t/s is normal - 4090 draws 250-350W under load</li>
</ul>
<p dir="auto">To reduce fan noise:</p>
<ul>
<li>Power limit to 250W (10-15% t/s drop, much quieter)</li>
<li>Or use vLLM with --kv-cache-dtype fp8_e5m2</li>
</ul>
<p dir="auto">4090D 48G doing 130 t/s with vLLM AWQ is already excellent.</p>
]]></description><link>https://lcz.me/post/10452</link><guid isPermaLink="true">https://lcz.me/post/10452</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Fri, 24 Jul 2026 13:12:20 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 13:06:27 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 请问sglang下token输出能达到多少呀？我使用vLLM,Q4下，越狱模型shawnw3i/Huihui-Qwen3.6-27B-abliterated-AWQ-MTP，能否达到130t/s.但是确实风扇狂转是真的，我也是4090魔改的显卡。</p>
]]></description><link>https://lcz.me/post/10449</link><guid isPermaLink="true">https://lcz.me/post/10449</guid><dc:creator><![CDATA[lukun ge]]></dc:creator><pubDate>Fri, 24 Jul 2026 13:06:27 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 11:22:05 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 看过视频了， 但是还是没有 GET 到 SG-LANG具体对软件环境、硬件和模型有什么限制。  比如说， 只能在 linux下？ 只能在某某内核版本下？ 只能在 nvidia下？ 只能 FP8?  具体的限制是什么呢</p>
]]></description><link>https://lcz.me/post/10445</link><guid isPermaLink="true">https://lcz.me/post/10445</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Fri, 24 Jul 2026 11:22:05 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 11:20:23 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 赞， 这个对于本地跑AI来说很重要啊</p>
]]></description><link>https://lcz.me/post/10444</link><guid isPermaLink="true">https://lcz.me/post/10444</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Fri, 24 Jul 2026 11:20:23 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 09:44:53 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/10395">说</a>:</p>
<p dir="auto">SG-Lang</p>
</blockquote>
<p dir="auto">7900xtx 能用上吗?</p>
]]></description><link>https://lcz.me/post/10438</link><guid isPermaLink="true">https://lcz.me/post/10438</guid><dc:creator><![CDATA[laobenxiong]]></dc:creator><pubDate>Fri, 24 Jul 2026 09:44:53 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 08:28:18 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E4%BF%A0%E5%AE%A2%E9%A2%A8" aria-label="Profile: 俠客風">@<bdi>俠客風</bdi></a> 我不是大师，论坛不是有超凡大师吗，有问题@他们，我积分高是因为我是站长，发帖多。</p>
]]></description><link>https://lcz.me/post/10426</link><guid isPermaLink="true">https://lcz.me/post/10426</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 24 Jul 2026 08:28:18 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 08:26:56 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/zhou-jian" aria-label="Profile: zhou-jian">@<bdi>zhou-jian</bdi></a> 大哥你看了视频吗？你好歹要看下视频啊，说的很清楚，就是交给DeepSeek+Hermes，让它们配置就好了啊。</p>
]]></description><link>https://lcz.me/post/10425</link><guid isPermaLink="true">https://lcz.me/post/10425</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 24 Jul 2026 08:26:56 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 08:26:11 GMT]]></title><description><![CDATA[<p dir="auto">大師您好，想加您好友，但不得其門而入！</p>
]]></description><link>https://lcz.me/post/10424</link><guid isPermaLink="true">https://lcz.me/post/10424</guid><dc:creator><![CDATA[俠客風]]></dc:creator><pubDate>Fri, 24 Jul 2026 08:26:11 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 05:34:18 GMT]]></title><description><![CDATA[<p dir="auto">楼主，能给出你的部署脚本吗？我也想抄作业。</p>
]]></description><link>https://lcz.me/post/10406</link><guid isPermaLink="true">https://lcz.me/post/10406</guid><dc:creator><![CDATA[zhou jian]]></dc:creator><pubDate>Fri, 24 Jul 2026 05:34:18 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 04:29:31 GMT]]></title><description><![CDATA[<p dir="auto">AMD R9700可以这么部署吗？</p>
]]></description><link>https://lcz.me/post/10402</link><guid isPermaLink="true">https://lcz.me/post/10402</guid><dc:creator><![CDATA[loulan]]></dc:creator><pubDate>Fri, 24 Jul 2026 04:29:31 GMT</pubDate></item><item><title><![CDATA[Reply to SG-Lang + Qwen3.6 27B + 4090 48G驱动Hermes完美平替在线AI，Linux/MacOS比Windows更适合跑AI服务和Agent部署！ on Fri, 24 Jul 2026 03:00:00 GMT]]></title><description><![CDATA[<p dir="auto">Youtube视频地址：<a href="https://youtu.be/7d3bxccoX9g" rel="nofollow ugc">https://youtu.be/7d3bxccoX9g</a></p>
<p dir="auto">有AWQ调试通了的，效率不错的欢迎留言，另外就是FP8模型我在服务器端和客户端设置No Thinking，Hermes请求还是会推理，虽然也不太影响效率，但是不喜欢思考模式，视频做完了不想继续调试，太忙了。有搞定的可以分享下，我准备直接抄作业，谢谢！</p>
<p dir="auto">就是一定要顺畅跑起来Hermes上下文拉满的啊，千万不要半成品发作业让我抄。我这套稳定工作了很久，没啥问题。</p>
]]></description><link>https://lcz.me/post/10396</link><guid isPermaLink="true">https://lcz.me/post/10396</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 24 Jul 2026 03:00:00 GMT</pubDate></item></channel></rss>