<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测]]></title><description><![CDATA[<h2>我的配置</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>配置</th>
</tr>
</thead>
<tbody>
<tr>
<td>系统</td>
<td>Ubuntu 26.04 LTS</td>
</tr>
<tr>
<td>CPU</td>
<td>i5-12600KF</td>
</tr>
<tr>
<td>内存</td>
<td>32 GB</td>
</tr>
<tr>
<td>显卡</td>
<td><strong>AMD RX 7900 XTX 24G</strong></td>
</tr>
<tr>
<td>推理后端</td>
<td>llama.cpp</td>
</tr>
<tr>
<td>模型</td>
<td><strong>Qwen 3.8 27B，Q4_K_M 量化</strong></td>
</tr>
<tr>
<td>Agent 框架</td>
<td>DeepSeek Harness（<code>@deepseek-ai/dsh</code> v0.1.0-rc.6）</td>
</tr>
</tbody>
</table>
<p dir="auto">模型文件约 16G，24G 显存能<strong>全部塞进 GPU</strong></p>
<h2>速度实测</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th>输入长度</th>
<th>输出 token</th>
<th>首字延迟</th>
<th>生成速度</th>
</tr>
</thead>
<tbody>
<tr>
<td>简单问答</td>
<td>67 tok</td>
<td>128</td>
<td>0.29s</td>
<td><strong>33.4 tok/s</strong></td>
</tr>
<tr>
<td>写代码</td>
<td>127 tok</td>
<td>512</td>
<td>0.15~0.31s</td>
<td><strong>33.0 tok/s</strong></td>
</tr>
<tr>
<td>长文档问答（9K 上下文）</td>
<td>9,392 tok</td>
<td>160</td>
<td>0.17~0.22s</td>
<td><strong>29 tok/s</strong></td>
</tr>
<tr>
<td>冷启动预填充（2.8K 新提示）</td>
<td>2,815 tok</td>
<td>—</td>
<td>4.25s</td>
<td><strong>~663 tok/s</strong></td>
</tr>
</tbody>
</table>
<h2>具体让它干了什么</h2>
<p dir="auto">手上有一个电机工程的技能包（一个 <code>.skill</code> 文件，里面是算法脚本），功能是把电机 MAP 长表转成四象限二维 Lookup Table，给 MATLAB/Simulink 用。</p>
<p dir="auto">任务是：<strong>把它工程化成一个符合 MCP 规范的 Server</strong>，让 Claude Desktop、Cursor、Kimi 这些支持 MCP 的客户端都能直接调用。它自己用了快20分钟，生成了交付物，我测了一下，效果确实是我想要的效果，看起来实现的非常不错。它自己开发，测试脚本自己写自己测，整个流程都是它自己在优化，全程没有在指导过什么。<br />
总结：非常好用，能力很强的稠密模型再加上harness工程，确实是一把干活的利器。</p>
]]></description><link>https://lcz.me/topic/1124</link><generator>RSS for Node</generator><lastBuildDate>Thu, 10 Sep 2026 00:43:27 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1124.rss" rel="self" type="application/rss+xml"/><pubDate>Fri, 14 Aug 2026 19:09:12 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测 on Sun, 16 Aug 2026 13:42:41 GMT]]></title><description><![CDATA[<p dir="auto">实测了，35B还是不行，老老实实用27B了</p>
]]></description><link>https://lcz.me/post/12396</link><guid isPermaLink="true">https://lcz.me/post/12396</guid><dc:creator><![CDATA[vosrock]]></dc:creator><pubDate>Sun, 16 Aug 2026 13:42:41 GMT</pubDate></item><item><title><![CDATA[Reply to 7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测 on Sat, 15 Aug 2026 08:58:30 GMT]]></title><description><![CDATA[<p dir="auto">最近35B感觉又行了，之前测试用了Q4量化，最近发现实际上用Q8量化也是可以跑的，就是速度慢点，而且KV量化也可以用Q80，甚至不做KV量化，这样比原来测试的时候精度提高不少</p>
]]></description><link>https://lcz.me/post/12269</link><guid isPermaLink="true">https://lcz.me/post/12269</guid><dc:creator><![CDATA[vosrock]]></dc:creator><pubDate>Sat, 15 Aug 2026 08:58:30 GMT</pubDate></item><item><title><![CDATA[Reply to 7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测 on Sat, 15 Aug 2026 08:04:48 GMT]]></title><description><![CDATA[<p dir="auto">rocm看来和cuda还是有差距</p>
]]></description><link>https://lcz.me/post/12261</link><guid isPermaLink="true">https://lcz.me/post/12261</guid><dc:creator><![CDATA[ezios]]></dc:creator><pubDate>Sat, 15 Aug 2026 08:04:48 GMT</pubDate></item><item><title><![CDATA[Reply to 7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测 on Sat, 15 Aug 2026 01:41:37 GMT]]></title><description><![CDATA[<p dir="auto">日常改个程序，驱动hermes  ，还是qwen35b的那种更丝滑。</p>
]]></description><link>https://lcz.me/post/12245</link><guid isPermaLink="true">https://lcz.me/post/12245</guid><dc:creator><![CDATA[wwcd2016]]></dc:creator><pubDate>Sat, 15 Aug 2026 01:41:37 GMT</pubDate></item><item><title><![CDATA[Reply to 7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测 on Sat, 15 Aug 2026 01:29:43 GMT]]></title><description><![CDATA[<p dir="auto">这玩意直接下载。然后把模型名字一改。跑起来 和原来的3.6 没有感觉有变化。</p>
]]></description><link>https://lcz.me/post/12244</link><guid isPermaLink="true">https://lcz.me/post/12244</guid><dc:creator><![CDATA[wwcd2016]]></dc:creator><pubDate>Sat, 15 Aug 2026 01:29:43 GMT</pubDate></item><item><title><![CDATA[Reply to 7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测 on Fri, 14 Aug 2026 19:23:06 GMT]]></title><description><![CDATA[<p dir="auto">这个截图到没有，晚上快快的部署了，用了用，qwen3.8 27B的思考链太长了，有时候还会重头思考，3.6 27B好像没有出现过这个问题。上面那个任务它跑了20多分钟。花在思考上的时间非常多，明天准备把思考链关了试试，看看输出有没有太大区别。</p>
]]></description><link>https://lcz.me/post/12229</link><guid isPermaLink="true">https://lcz.me/post/12229</guid><dc:creator><![CDATA[Phuong Ngo]]></dc:creator><pubDate>Fri, 14 Aug 2026 19:23:06 GMT</pubDate></item><item><title><![CDATA[Reply to 7900XTX+Qwen 3.8 27B Q4 + DeepSeek Harness 实测 on Fri, 14 Aug 2026 19:16:22 GMT]]></title><description><![CDATA[<p dir="auto">有实测截图吗，挺有意思的，速度还挺快的</p>
]]></description><link>https://lcz.me/post/12228</link><guid isPermaLink="true">https://lcz.me/post/12228</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 14 Aug 2026 19:16:22 GMT</pubDate></item></channel></rss>