<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Deepseek-Harness 单卡（7900XTX）运行Qwen3.8-27B 完整文档-3]]></title><description><![CDATA[<h1>6. 性能测试</h1>
<h3>6.1 测试脚本</h3>
<p dir="auto">存放到 <code>~/llama-bin/bench.py</code>：</p>
<pre><code class="language-python">#!/usr/bin/env python3
import time, json, urllib.request, concurrent.futures, statistics

API = "http://localhost:8090/v1/chat/completions"
MODEL = "qwen3.8-27b"

def chat(prompt, max_tokens, temp=0):
    payload = {"model": MODEL, "messages": [{"role": "user", "content": prompt}],
               "max_tokens": max_tokens, "temperature": temp, "stream": True}
    req = urllib.request.Request(API, data=json.dumps(payload).encode(),
                                  headers={"Content-Type": "application/json"})
    start = time.time(); first = None; tok = 0
    try:
        with urllib.request.urlopen(req) as r:
            for line in r:
                line = line.decode().strip()
                if line.startswith("data: "):
                    d = line[6:]
                    if d == "[DONE]": break
                    try:
                        o = json.loads(d)
                        c = o.get("choices", [{}])[0].get("delta", {}).get("content")
                        if c:
                            if first is None: first = time.time()
                            tok += 1
                    except: pass
    except: pass
    total = time.time() - start
    ttft = first - start if first else 0
    return ttft, total, tok

def sec(t):
    print(f"\n{chr(61)*64}\n  {t}\n{chr(61)*64}")

# ① 吞吐量
sec("① 吞吐量测试")
tr = []
for label, prompt, mx in [
    ("短输出 100", "用一句话介绍人工智能。", 100),
    ("中输出 500", "请写一段关于春天的短文，约200字。", 500),
    ("长输出 1000", "请写一篇关于AI发展历程的文章，尽量详细。", 1000),
    ("超长 2000", "请详细解释量子计算原理，写3000字文章。", 2000),
]:
    ttft, total, tok = chat(prompt, mx)
    gt = total - ttft
    tps = tok / gt if gt &gt; 0 else 0
    tr.append({"l": label, "ttft": ttft, "tps": tps})
    print(f"  {label:&lt;12} TTFT:{ttft*1000:&gt;6.0f}ms {gt:&gt;6.2f}s {tok:&gt;5}tok {tps:&gt;6.1f}tok/s")
print(f"  平均: {statistics.mean(r['tps'] for r in tr):.1f} tok/s")

# ② 并发
sec("② 并发测试")
for n in [2, 4, 8]:
    def w(i): return chat(f"{i}+{i}等于多少？", 50)
    s = time.time()
    with concurrent.futures.ThreadPoolExecutor(max_workers=n) as p:
        results = [f.result() for f in [p.submit(w, i) for i in range(n)]]
    wall = time.time() - s
    total_tok = sum(r[2] for r in results)
    print(f"  并发={n} 耗时:{wall:.2f}s tokens:{total_tok} 吞吐:{total_tok/wall:.1f}tok/s")

# ③ 稳定性
sec("⑥ 稳定性测试 (5次)")
st = []
for i in range(5):
    ttft, total, tok = chat("用一句话介绍机器学习。", 100)
    gt = total - ttft
    tps = tok / gt if gt &gt; 0 else 0
    st.append({"ttft": ttft, "tps": tps})
    print(f"  第{i+1}次: TTFT:{ttft*1000:.0f}ms {tok}tok {tps:.1f}tok/s")
print(f"  TTFT: min={min(r['ttft'] for r in st)*1000:.0f} max={max(r['ttft'] for r in st)*1000:.0f}")
print(f"  速度: min={min(r['tps'] for r in st):.1f} max={max(r['tps'] for r in st):.1f} avg={statistics.mean(r['tps'] for r in st):.1f}")
</code></pre>
<h3>6.2 实测结果</h3>
<p dir="auto"><strong>吞吐量</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>输出长度</th>
<th>TTFT</th>
<th>生成耗时</th>
<th>Tokens</th>
<th>速度</th>
</tr>
</thead>
<tbody>
<tr>
<td>短 (100)</td>
<td>5132ms*</td>
<td>0.58s</td>
<td>32</td>
<td>55.2 tok/s</td>
</tr>
<tr>
<td>中 (500)</td>
<td>619ms</td>
<td>2.62s</td>
<td>147</td>
<td>56.2 tok/s</td>
</tr>
<tr>
<td>长 (1000)</td>
<td>623ms</td>
<td>17.35s</td>
<td>1000</td>
<td>57.7 tok/s</td>
</tr>
<tr>
<td>超长 (2000)</td>
<td>660ms</td>
<td>34.68s</td>
<td>1839</td>
<td>53.0 tok/s</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">*首次含冷启动加载，后续 TTFT 稳定 ~600ms。平均 <strong>55.5 tok/s</strong></p>
</blockquote>
<p dir="auto"><strong>并发</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>并发数</th>
<th>总耗时</th>
<th>总Tokens</th>
<th>吞吐</th>
</tr>
</thead>
<tbody>
<tr>
<td>2</td>
<td>1.49s</td>
<td>7</td>
<td>4.7 tok/s</td>
</tr>
<tr>
<td>4</td>
<td>2.64s</td>
<td>19</td>
<td>7.2 tok/s</td>
</tr>
<tr>
<td>8</td>
<td>4.39s</td>
<td>46</td>
<td>10.5 tok/s</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">并发下 TTFT 线性增长（串行处理，--parallel 1）</p>
</blockquote>
<p dir="auto"><strong>温度影响</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>温度</th>
<th>TTFT</th>
<th>速度</th>
</tr>
</thead>
<tbody>
<tr>
<td>0</td>
<td>333ms</td>
<td>53.4 tok/s</td>
</tr>
<tr>
<td>0.5</td>
<td>451ms</td>
<td>43.0 tok/s</td>
</tr>
<tr>
<td>1.0</td>
<td>450ms</td>
<td>46.2 tok/s</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>长上下文</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>上下文</th>
<th>TTFT</th>
<th>速度</th>
</tr>
</thead>
<tbody>
<tr>
<td>短 (~200字)</td>
<td>622ms</td>
<td>46.4 tok/s</td>
</tr>
<tr>
<td>长 (~1600字)</td>
<td>1439ms</td>
<td>55.1 tok/s</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>稳定性 (5次相同请求)</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>最小</th>
<th>最大</th>
<th>平均</th>
</tr>
</thead>
<tbody>
<tr>
<td>TTFT</td>
<td>153ms</td>
<td>624ms</td>
<td>305ms</td>
</tr>
<tr>
<td>速度</td>
<td>48.7</td>
<td>52.7</td>
<td>51.8 tok/s</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">速度波动 &lt;8%，非常稳定</p>
</blockquote>
<blockquote>
<p dir="auto"><strong>口径说明</strong>：上表测于<strong>未启用 k4v</strong> 时期，且为纯文本负载。启用 k4v 后的解码表现见 §6.3。<br />
两组数字不可直接比较——负载类型不同（散文 vs 逐字重现）。</p>
</blockquote>
<hr />
<h3>6.3 投机解码专项（k4v）</h3>
<p dir="auto">llama.cpp 的投机解码分支支持<strong>多个投机器同时启用</strong>，<code>--spec-type</code> 吃逗号清单。<br />
在原有 <code>draft-mtp</code>（模型内建 MTP 头自投机）之上叠加 <code>ngram-map-k4v</code>（从已生成上下文做<br />
n-gram 查表，命中一次吐最多 48 个 draft token），二者互补：<strong>n-gram 命中时是"猜一大段"，<br />
没命中就退回 MTP，所以不会拖慢其他内容。</strong></p>
<pre><code class="language-bash">--spec-type ngram-map-k4v,draft-mtp \
--spec-draft-n-max 3 \
--spec-ngram-map-k4v-size-n 32 \
--spec-ngram-map-k4v-size-m 48 \
--spec-ngram-map-k4v-min-hits 1
</code></pre>
<p dir="auto"><strong>零 VRAM 成本、零重编</strong>，已在 <code>llama.sh</code> 中默认开启（<code>K4V=0</code> 可关）。</p>
<h4>实测结果</h4>
<p dir="auto">方法：temp=0、每项 2 轮取中位数、读 <code>timings.predicted_per_second</code>。<br />
<strong>两配置 <code>n_gen</code> 完全一致</strong>（407/365/500/271/630），故无「输出长度不同造成的假信号」。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>工作负载</th>
<th>仅 draft-mtp</th>
<th>+k4v (n=32)</th>
<th>变化</th>
<th>接受率变化</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>code_refactor</code></td>
<td>86.1</td>
<td><strong>112.6</strong></td>
<td><strong>+30.7%</strong></td>
<td>1.00 → 0.96</td>
</tr>
<tr>
<td><code>repeat_synth</code></td>
<td>86.0</td>
<td><strong>122.4</strong></td>
<td><strong>+42.3%</strong></td>
<td>1.00 → 0.99</td>
</tr>
<tr>
<td><code>template_batch</code></td>
<td>47.2</td>
<td>46.7</td>
<td>−1.1%</td>
<td>0.37 → 0.37</td>
</tr>
<tr>
<td><code>reasoning</code></td>
<td>72.9</td>
<td>72.1</td>
<td>−1.1%</td>
<td>0.79 → 0.79</td>
</tr>
<tr>
<td><code>prose</code></td>
<td>47.2</td>
<td>46.7</td>
<td>−1.1%</td>
<td>0.38 → 0.38</td>
</tr>
</tbody>
</table>
<p dir="auto">负项均落在 ±1.1%，与同配置重跑的噪音水平一致。经 systemd 重启后复验：112.3 / 122.7。</p>
<h4>两个反直觉点</h4>
<ol>
<li><strong>接受率下降不是退化。</strong> <code>code_refactor</code> 接受率 1.00 → 0.96 看着像变差，实则<br />
草稿 token 从 306 涨到 400、总吞吐 86 → 112。<strong>接受率必须与 <code>drafted</code> 一起看</strong>，<br />
孤立看会误判（参见 §9.6 记录的同类教训）。</li>
<li><strong>收益来自「逐字重现」，不是「格式相似」。</strong></li>
</ol>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>你的任务长这样</th>
<th>k4v 收益</th>
</tr>
</thead>
<tbody>
<tr>
<td>让模型把整份文件/文章原样吐回、只改几行（<strong>改代码、翻译校对、逐条批注</strong>）</td>
<td><strong>+30~60%，最值钱</strong></td>
</tr>
<tr>
<td>大量重复的结构化输出（JSON/YAML/SRT/表格批次）</td>
<td>中等</td>
</tr>
<tr>
<td>逐条照抄原文再加判定（规格检查、RAG 引述）</td>
<td>小幅或持平</td>
</tr>
<tr>
<td>格式固定但内容全新（套模板产新数据）</td>
<td>持平（<code>size-n</code> 取 12 时倒扣）</td>
</tr>
<tr>
<td>纯创作、思考链、对话</td>
<td>无感（也不会变慢）</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">对 dsh 而言第一行才是主场：agent 改文件时经常要把整个文件回吐一遍，这就是 +30.7% 那档。</p>
</blockquote>
<h4>为什么 <code>size-n</code> 必须是 32</h4>
<p dir="auto"><code>size-n</code> 决定「要比对多长的前文才算命中」。拉长 = 更严格 = 少开火但更准。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>size-n</th>
<th>code_refactor</th>
<th>template_batch</th>
<th>字幕重排类</th>
<th>合成重复</th>
</tr>
</thead>
<tbody>
<tr>
<td>关闭</td>
<td>86.1</td>
<td>47.2</td>
<td>106.1</td>
<td>86.0</td>
</tr>
<tr>
<td>12</td>
<td>166.4</td>
<td>68.4 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=4368ee61982" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> −5%</td>
<td>95.8 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=4368ee61982" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> −10%</td>
<td>178.6</td>
</tr>
<tr>
<td>24</td>
<td>169.1</td>
<td>70.2</td>
<td>100.9 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=4368ee61982" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /></td>
<td>—</td>
</tr>
<tr>
<td><strong>32</strong></td>
<td><strong>157.0</strong></td>
<td>71.1</td>
<td>105.4</td>
<td>237.0</td>
</tr>
<tr>
<td>40</td>
<td>136.0</td>
<td>71.9</td>
<td>104.8</td>
<td>236.9</td>
</tr>
</tbody>
</table>
<p dir="auto">（上表为来源帖在 RTX 4080S / 7900 XTX 上的扫描值，用于说明趋势，本机仅实测 n=32 一档。）</p>
<p dir="auto"><code>size-n</code> 太小 → n-gram 一直开火、一直猜错，<strong>白做的 draft 计算比不开启还多</strong>。<br />
典型反例是"文字照抄、时间轴改掉"的字幕重排：n-gram 看到前文就把旧时间戳一起预测出来，<br />
每到时间轴就撞墙，draft 992 个只接受 510。<strong>格式重复但内容要换的任务，n-gram 是负资产。</strong></p>
<h4>不要动 <code>min-hits</code></h4>
<p dir="auto">来源帖试过 <code>--spec-ngram-map-k4v-min-hits 2</code>（要求 n-gram 出现两次才采用），<br />
结果各项 draft 计数与完全关闭 k4v <strong>逐项相同</strong>——等于直接关掉功能。保持 1。</p>
<h4>不要设 <code>p_min</code>，<code>n_max</code> 保持 3</h4>
<p dir="auto"><code>n_max=5</code> + <code>p_min=0.4</code> 在来源帖实测中「重复内容看起来更快，但推理 −22%、散文 −37%」：<br />
draft 猜长了接受率上不去（散文只有 0.39），每步都在做白工。<code>n_max</code> 越大，<br />
可预测内容越快、不可预测内容越慢——纯创作偏多用 3，结构化输出偏多可试 4~5。</p>
<p dir="auto">来源：<code>https://lcz.me/topic/1398</code>（7900 XTX / Vulkan 实测报告，附 RTX 交叉验证）</p>
<hr />
<h2>7. 日常运维</h2>
<h3>7.1 常用命令</h3>
<pre><code class="language-bash"># 启动模型（自动带 shim；有 systemd 单元且无覆盖变量时自动走 systemctl）
~/llama-bin/llama.sh 38

# systemd 托管时的启动/重启（推荐；走 llama.sh 38 --fg，会继承脚本内全部参数）
systemctl --user start  qwen38
systemctl --user restart qwen38
systemctl --user stop   qwen38     # 手动停，Restart=on-failure 不会拉回

# 查看状态
~/llama-bin/llama.sh status

# 停止
~/llama-bin/llama.sh stop

# 切换模型
~/llama-bin/llama.sh 36 --use

# 临时覆盖参数（单次生效）
K4V=0 ~/llama-bin/llama.sh 38          # 关掉 ngram-k4v，只留 MTP
NMAX=2 CRAM=0 ~/llama-bin/llama.sh 38   # 关掉 KV 驻内存
CTX=65536 ~/llama-bin/llama.sh 38       # 加大窗口前先读 §9.8：逼近上限会掉进 GTT

# shim 单独管理
~/llama-bin/shim.sh status
~/llama-bin/shim.sh restart

# 查看日志
tail -f /tmp/qwen38-vulkan.log    # llama.cpp
tail -f /tmp/dsh-llm-shim.log     # shim
journalctl --user -u qwen38 -f     # systemd 下 llama.cpp 的日志
</code></pre>
<blockquote>
<p dir="auto"><strong>改 <code>llama.sh</code> 后不需要动 systemd 单元</strong> —— 单元的 <code>ExecStart</code> 是 <code>llama.sh 38 --fg</code>，<br />
会自动继承脚本里的新参数。<code>systemctl --user restart qwen38</code> 即可生效。</p>
</blockquote>
<h3>7.2 健康检查</h3>
<pre><code class="language-bash"># llama.cpp
curl -s http://127.0.0.1:8080/health | python3 -m json.tool

# shim
curl -s http://127.0.0.1:8090/health | python3 -m json.tool

# 模型列表
curl -s http://127.0.0.1:8090/v1/models | python3 -m json.tool

# 快速测试
curl -s http://127.0.0.1:8090/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"你好"}],"max_tokens":50}'
</code></pre>
<hr />
<h2>8. 故障排查</h2>
<h3>8.1 常见问题</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>症状</th>
<th>原因</th>
<th>解决</th>
</tr>
</thead>
<tbody>
<tr>
<td>shim: UPSTREAM_UNREACHABLE</td>
<td>llama.cpp 未启动</td>
<td>~/llama-bin/llama.sh 38</td>
</tr>
<tr>
<td>CONTEXT_OVERFLOW</td>
<td>上下文超限</td>
<td>/compact 压缩或新开会话</td>
</tr>
<tr>
<td>工具调用参数错乱</td>
<td>绕过 shim 直连 llama.cpp</td>
<td>确保 baseURL 指向 :8090</td>
</tr>
<tr>
<td>模型跑飞 (32k token)</td>
<td>llama.cpp 解析器 bug</td>
<td>必须经 shim，不能直连</td>
</tr>
<tr>
<td>端口冲突</td>
<td>8080/8081/8090 被占用</td>
<td>改 P38_PORT/SHIM_PORT 环境变量</td>
</tr>
<tr>
<td>显存不足 OOM</td>
<td>两模型同时运行</td>
<td>只能跑一个，先 stop 再启另一个</td>
</tr>
<tr>
<td>MTP 头被忽略</td>
<td>--spec-type 未生效</td>
<td>检查日志 grep "nextn.*ignoring"</td>
</tr>
</tbody>
</table>
<h3>8.2 日志位置</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>服务</th>
<th>日志文件</th>
</tr>
</thead>
<tbody>
<tr>
<td>llama.cpp 3.8</td>
<td>/tmp/qwen38-vulkan.log</td>
</tr>
<tr>
<td>llama.cpp 3.6</td>
<td>/tmp/qwen36-vulkan.log</td>
</tr>
<tr>
<td>shim</td>
<td>/tmp/dsh-llm-shim.log</td>
</tr>
<tr>
<td>DSH</td>
<td>~/.dsh/logs/</td>
</tr>
</tbody>
</table>
<h3>8.3 完全重置</h3>
<pre><code class="language-bash"># 停止所有服务
~/llama-bin/llama.sh stop
~/llama-bin/shim.sh stop

# 验证端口已释放
ss -tlnp | grep -E "8080|8081|8090"

# 重新启动
~/llama-bin/llama.sh 38
</code></pre>
<hr />
<h2>9. 诊断与实战记录</h2>
<h3>9.1 直连失败的具体表现与根因</h3>
<p dir="auto">直连 llama.cpp 的 OpenAI 端点时：</p>
<pre><code>assistant/message: content: []   stopReason: "length"
usage: output 32768 token   耗时 7.7 分钟   → 0 可用输出
</code></pre>
<p dir="auto">根因不在模型，而在 <strong>llama.cpp b11223 自带的 tool-call 解析器</strong>（该 build 报 <code>chat_format: peg-native</code>）。用 <code>/apply-template</code> + <code>/completion</code> 绕过解析器取模型原始输出：</p>
<pre><code>&lt;tool_call&gt; &lt;function=bash&gt; &lt;parameter=command&gt; ls -1 /home/hnz/桌面 | wc -l &lt;/parameter&gt;
&lt;parameter=description&gt; Count entries in Desktop directory &lt;/parameter&gt; &lt;/function&gt; &lt;/tool_call&gt;
</code></pre>
<p dir="auto">49 token、格式完全正确、正常 EOS。但同一个请求走 <code>/v1/chat/completions</code>（带 <code>tools</code>）时，解析结果是：</p>
<pre><code class="language-json">{"command": "ls -1 /home/hnz/桌面 | wc -l &lt;/parameter&gt; &lt;parameter=description&gt; ... &lt;/tool_call&gt;\n&lt;tool_call&gt;..."}
</code></pre>
<p dir="auto">即第二个参数起的整段被吞进第一个参数的值里；而且<strong>这种失败会诱发模型把同一个 <code>&lt;tool_call&gt;</code> 重复几百遍</strong>直到打满 <code>max_tokens</code>。模型自带模板明确允许参数值跨行（"that can span multiple lines"），解析器却处理不了。</p>
<p dir="auto">对照实验（同一份真实 DSH 请求，24 个工具）：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>路径</th>
<th>结果</th>
</tr>
</thead>
<tbody>
<tr>
<td>直连 <code>:8081</code></td>
<td><strong>0/10</strong> 成功，每次跑飞约 30s / 满 <code>max_tokens</code></td>
</tr>
<tr>
<td>经 shim</td>
<td><strong>10/10</strong> 成功，每次约 1s / 49~51 token</td>
</tr>
</tbody>
</table>
<p dir="auto">采样侧只能缓解（<code>--dry-multiplier 0.8</code> 能让它不再无限重复，但参数照样被解析坏），换 build 可能有效但会丢掉这套 MTP 调优（69 tok/s），所以选了中间层方案。</p>
<blockquote>
<p dir="auto"><strong><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a0.png?v=4368ee61982" class="not-responsive emoji emoji-android emoji--warning" style="height:23px;width:auto;vertical-align:middle" title="⚠" alt="⚠" /> 后续复核推翻了这组对照（2026-09-29，见 §1.2 与 §9.7）</strong>：当前配置下直连可正常完成<br />
工具调用，0/10 复现不出来。下文保留为<strong>历史记录</strong>，不代表当前实际行为。</p>
<p dir="auto">该故障的<strong>现场记录</strong>已归档：本目录 <code>local_model.sh.txt</code><br />
（模型输出的原始抓取，含 4058 个零宽字符 U+200B 残留的重复 <code>&lt;tool_call&gt;</code> 标签，<br />
<code>bash -n</code> 直接语法错误 —— 正是「重复到打满 <code>max_tokens</code>」的产物）。</p>
</blockquote>
<p dir="auto">对照另一个模型：<code>qwen3.8-27b</code>（2026-09-29 实测，当时 <code>n_ctx</code> 还是 262144；现已改为 131072，见 §9.8）经 shim 的干净工具调用 <strong>5/5</strong>、每次 1~2 秒；纯文本回复在<strong>流式与非流式</strong>下都与直连逐字节一致（<code>temperature 0</code> + 固定 seed 对齐）。3.6/3.8 两个实例的严格模式提示双向验证通过。</p>
<h3>9.2 验证记录</h3>
<ol>
<li><strong>解析单测</strong> <code>node ~/llama-bin/shim-parser-test.mjs</code> → 7/7（含 <code>&lt;bash&gt;</code> 漏写 <code>function=</code> 的容错）。</li>
<li><strong>真实请求重放</strong> 10/10 干净 <code>tool_calls</code>，每次 1s。</li>
<li><strong>headless 端到端</strong>（<code>dsh --profile headless --patch dsh-shim-overlay.yml "…"</code>）：
<ul>
<li>单步任务：<code>tool-call bash → "3" → 最终回答 → turn/end: completed</code></li>
<li>多步任务：一轮里并行发出 3 个 bash + 1 个 grep，再补一次 bash，最终答案<br />
（9 个文件 / 1 个子目录 / <code>SHIM_PORT</code> 默认 8090）全部正确。</li>
<li>缓存命中正常（cacheRead 4928 / 5402 token）。</li>
</ul>
</li>
<li><strong>失败路径</strong>：上游不可用时返回 <code>HTTP 502</code> + 明确错误信息；<code>/health</code> 只反映 shim 自身。</li>
<li><strong>k4v 启用后的复验</strong>（2026-09-29）：经 systemd <code>llama.sh 38 --fg</code> 启动后，<br />
投机解码基准复现手工启动的数值（112.3 / 122.7），dsh 经 shim 端到端冒烟正常。</li>
<li><strong>默认模型切换</strong>：<code>dsh-shim-overlay.yml</code> 的默认模型已由 <code>qwen3.6-27b</code> 改为 <code>qwen3.8-27b</code>，<br />
与 <code>llama.sh</code> 默认保持一致。改前若 3.6 未启动，shim 会返回<br />
<code>502 UPSTREAM_UNREACHABLE</code> 并提示当前实际在跑哪个实例。</li>
</ol>
<h3>9.3 修复记录（都是 shim / 脚本侧的问题，DSH 走流式不受影响）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>症状</th>
<th>原因</th>
<th>修法</th>
</tr>
</thead>
<tbody>
<tr>
<td>非流式纯文本回复只剩最后 ~11 个字符</td>
<td>防标签跨 chunk 的 <code>sent</code> 指针在非流式路径也前进，但那条路并不发送</td>
<td>只在真正流式时推进 <code>sent</code></td>
</tr>
<tr>
<td>工具调用之前的正文被丢掉</td>
<td>流式分支没把保留尾巴里的正文发出去</td>
<td>发 <code>tool_calls</code> 前先补发 <code>prefixText</code></td>
</tr>
<tr>
<td><code>shim.sh stop</code> 误杀无关进程（连调用它的 shell 都被杀过）</td>
<td><code>pkill -f "dsh-llm-shim.mjs"</code> 按文件名裸匹配，<code>vim 该文件</code> 也会中招</td>
<td>启动写 PID 文件，停止时按 PID + 只认 <code>argv[0]</code> 含 <code>node</code> 的进程</td>
</tr>
<tr>
<td><code>llama.sh status</code> 偶尔显示「停止」而模型其实在跑</td>
<td>pgrep 在某些环境下看不到别的进程</td>
<td>状态行加端口健康检查兜底</td>
</tr>
</tbody>
</table>
<h3>9.4 上下文限制与超限处理</h3>
<ul>
<li><strong>本地模型的上下文比云端小得多，切模型前要看一眼历史长度</strong>。<code>qwen3.8-27b</code> 的<br />
<code>n_ctx</code> 现在是 <strong>131072</strong>（不再是 262144，见 §9.8），<code>maxTokens: 32768</code>，<br />
所以<strong>留给 prompt 的预算约 9.8 万 token</strong>——本地会话比想象的短得多。<br />
把一个已经很长的云端会话直接切到本地，llama.cpp 会回 <code>400 exceed_context_size_error</code>。正确做法：
<ul>
<li><strong>另开一个新会话</strong>用本地模型（最稳）；或</li>
<li>先在当前会话 <code>/compact</code> 压缩，再 <code>/model</code> 切到本地；</li>
<li>确实需要更长窗口，见 §9.8 关于 <code>CTX</code> 与显存的取舍，不要盲目调大。<br />
现在这种情况会明确报 <code>CONTEXT_OVERFLOW</code> 并给出 token 数，不会再静默无响应。<br />
实测 DSH 侧呈现为 <code>dsh: INVALID_REQUEST: 400: 本地模型上下文不足：请求 N token，而模型 ctx 只有 131072。请在 DSH 里 /compact 压缩会话，或另开一个新会话。</code>，<br />
并且<strong>不再重试 5 次</strong>（<code>INVALID_REQUEST</code> 不在 <code>llm-retry</code> 的重试类别里，立刻失败并给出原因）。</li>
</ul>
</li>
<li><strong><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a0.png?v=4368ee61982" class="not-responsive emoji emoji-android emoji--warning" style="height:23px;width:auto;vertical-align:middle" title="⚠" alt="⚠" /> 超限之后的"压缩死循环"（2026-09-29 实测）</strong>：会话一旦超过本地窗口，DSH 会自动<br />
触发压缩，但<strong>压缩请求本身也要把整段历史发给同一个模型</strong>，于是照样超限：<pre><code>13:08:09  provider=local model=qwen3.8-27b ctx=262144   ← 当时的配置，现为 131072
13:08:11  compaction/start
13:08:13  compaction/end  error: 400 请求 330277 token &gt; ctx 262144
13:08:15  turn/end  error: 400 CONTEXT_OVERFLOW（请求 329927 token）
</code></pre>
而 <code>/compact</code> <strong>不接受参数、用当前选中的模型</strong>，所以在本地模型上手动压缩同样会失败。<br />
正确顺序（二选一）：
<ol>
<li><strong>另开新会话再选本地模型</strong>（已验证：provider=local/qwen3.8-27b，2 次工具调用、<br />
<code>turn/end: completed</code>，prompt 仅约 5.5k token）；</li>
<li>想保留本会话：先 <code>/model</code> 切回 <strong>deepseek-flash → 发一条消息让切换生效 → <code>/compact</code></strong><br />
（云端 100 万窗口才压得动）→ 确认压缩成功后，再 <code>/model</code> 切回本地。<br />
注意切换生效前触发的压缩仍会用旧模型（实测踩过），所以要先发一条消息。</li>
</ol>
</li>
</ul>
<h3>9.5 日常使用注意</h3>
<ul>
<li><strong>一条命令就够</strong>：<code>llama.sh 36|38</code> 在模型就绪后会自动 <code>shim.sh start</code>（幂等，已在跑就跳过）；<br />
<code>llama.sh stop</code> 与 <code>llama.sh 36 --use</code> 会连带停掉 shim。<code>llama.sh status</code> 也会打印 shim 状态行。<br />
（改动前已备份为 <code>~/llama-bin/llama.sh.bak-*</code>、<code>shim.sh.bak-*</code>。）</li>
<li>单独控制 shim：<code>~/llama-bin/shim.sh start|stop|restart|status</code>。<br />
升级过 <code>dsh-llm-shim.mjs</code> 之后要 <code>shim.sh restart</code> 才会生效（下次 <code>llama.sh 36|38</code> 起模型也会自动重启它）。</li>
<li>端口占用兜底：若 :8090 已被别的实例服务（例如别的终端拉起的），<code>shim.sh start</code> 会跳过启动<br />
而不是抢端口；<code>llama.sh status</code> / <code>shim.sh status</code> 也会把这种情况算作“运行”。</li>
<li><strong>模型要对上</strong>：选 3.6 就得跑 <code>llama.sh 36</code>，选 3.8 就得跑 <code>llama.sh 38</code>（两个实例互斥）。<br />
对不上时 shim 会直接报错并说明现在跑的是哪个、该用哪条命令，不会拿另一个模型顶替。<br />
<code>llama.sh 36 --use</code> 切换模型时 shim 不用重启（它按模型名转发）。</li>
<li>shim 没启动时本地路由会 502；此时 GUI 里 <code>/model</code> 切回 <code>deepseek-flash</code> 即可。</li>
<li>本会话里曾临时拉起过一个 shim 实例；<code>llama.sh stop</code> / <code>shim.sh stop</code> 都能把它停掉，<br />
之后由 <code>llama.sh</code> 正常接管。</li>
<li>日志：<code>/tmp/dsh-llm-shim.log</code>（<code>SHIM_VERBOSE=1</code> 时记录每次解析结果与重试）。</li>
</ul>
<h3>9.6 治本方向（上游修复）</h3>
<ol>
<li>换/升级 llama.cpp build —— 该 build 的 PEG 工具解析器是根因；上游修好后 shim 可以直接摘掉。</li>
<li>若换 build：本方案与 MTP 调优无关，shim 不依赖 <code>--spec-type draft-mtp</code>。</li>
<li>上报点：模板文档明示参数值可跨行，而 PEG 解析器处理不了跨行值，且失败会诱发重复生成。</li>
<li>另外值得反馈 DSH 上游：从大窗口模型（deepseek 100 万）切到小窗口模型（本地 26.2 万）时，<br />
自动压缩会用<strong>切换前</strong>的模型调度、且压缩请求本身可能超过目标窗口 —— 于是"越超限越压不动"。<br />
建议改成按目标模型窗口判断、并允许指定压缩用的模型。</li>
</ol>
<h3>9.7 补充复核：shim A/B 与多轮稳定性</h3>
<p dir="auto">2026-09-29 在当时配置（Qwen3.8-27B @262144、k4v 开启）下重做 A/B，<br />
两份 overlay 仅 <code>baseURL</code> 端口不同。</p>
<h4>A/B 结果</h4>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>路径</th>
<th>结果</th>
</tr>
</thead>
<tbody>
<tr>
<td>经 shim <code>:8090</code></td>
<td>66 个文件 / <code>ubuntu-ai</code> ✓</td>
</tr>
<tr>
<td>直连 <code>:8080</code></td>
<td>66 个文件 / <code>ubuntu-ai</code> ✓（并多解释了 <code>find -type f</code> 的口径）</td>
</tr>
</tbody>
</table>
<h4><code>coerce()</code> 类型强转专项</h4>
<p dir="auto">shim 的 <code>coerceCall()</code> 只在模型把数字/布尔<strong>发成带引号字符串</strong>时才触发修补<br />
（<code>dsh-llm-shim.mjs</code> 中 <code>coerce()</code>：<code>number</code>→<code>Number()</code>、布尔映射 <code>true/yes/1</code><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2194.png?v=4368ee61982" class="not-responsive emoji emoji-android emoji--left_right_arrow" style="height:23px;width:auto;vertical-align:middle" title="↔" alt="↔" /><code>false/no/0</code>、<br />
array/object→<code>JSON.parse()</code>，且仅处理 <code>typeof v === "string"</code>）。</p>
<p dir="auto">为最大化触发机会，构造 7 字段<strong>全部 required</strong> 的工具<br />
（string / integer / integer / boolean / number / array / object），逼模型逐个输出：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>路径</th>
<th>样本</th>
<th>类型全对</th>
<th>字符串化</th>
</tr>
</thead>
<tbody>
<tr>
<td>直连 <code>:8080</code></td>
<td>12</td>
<td><strong>12/12</strong></td>
<td><strong>0</strong></td>
</tr>
<tr>
<td>经 shim <code>:8090</code></td>
<td>6</td>
<td>6/6</td>
<td>0</td>
</tr>
</tbody>
</table>
<p dir="auto">直连返回 <code>{"task_name":"backup","delay_ms":1500,"repeat":3,"enabled":true,"threshold":0.75,"tags":["system","disk"],"opts":{"mode":"full"}}</code><br />
—— 全是原生 JSON 类型，<strong>18 次采样零次需要 <code>coerce</code> 修补</strong>。</p>
<blockquote>
<p dir="auto"><strong>这不能证明 <code>coerce</code> 永远无用。</strong> 它是"模型偶尔把数字发成字符串"的兜底，<br />
而本次采样温度固定 0.3、提示固定，输出高度一致，未触发该边缘情况。<br />
真实长会话的输出分布更杂。<strong>建议保留 shim、切直连观察几天，再决定是否下线。</strong></p>
</blockquote>
<h4>多轮对话不存在「越聊越慢」</h4>
<p dir="auto">社区方案建议加稳定性四件套<br />
<code>--no-kv-unified --parallel 1 --ctx-checkpoints 2 --cache-ram 4096</code> 来修多轮掉速。<br />
本机 4 轮实测（<code>cache_prompt:true</code>）显示<strong>该问题不存在</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>轮次</th>
<th><code>cache_n</code>（复用前缀）</th>
<th><code>prompt_n</code>（本轮新算）</th>
<th>decode t/s</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>0</td>
<td>67</td>
<td>64.1</td>
</tr>
<tr>
<td>2</td>
<td>366</td>
<td>60</td>
<td>66.4</td>
</tr>
<tr>
<td>3</td>
<td>725</td>
<td>49</td>
<td>70.9</td>
</tr>
<tr>
<td>4</td>
<td>1073</td>
<td>47</td>
<td>64.8</td>
</tr>
</tbody>
</table>
<p dir="auto">第 2 轮起每轮只重算 47~67 个 token，decode 无衰减（甚至微升）。<br />
<strong>原因：本机一直是 <code>--parallel 1</code></strong>，单槽下不存在 <code>kv_unified</code> 跨槽共享导致的状态失效。<br />
那套参数「免费」的前提是解决一个本机没有的问题，<strong>故不加</strong>。</p>
<h4>复现命令</h4>
<pre><code class="language-bash">DSH=~/.npm/_npx/1e7f6d9597241db0/node_modules/.bin/dsh
# 两份 overlay 仅 baseURL 端口不同（:8080 / :8090），其余完全一致
$DSH --profile headless --patch ~/llama-bin/dsh-shim-overlay.yml \
     headless "统计 ~/llama-bin 下的文件数并读出 /etc/hostname"

# 类型强转压测脚本原为 /tmp/opencode/coerce-test.py（临时目录，已随重启丢失）
# 需要复现时按 §9.7 的 7 字段 schema 描述重写即可
</code></pre>
<h3>9.8 CTX 从 256K 降到 128K 的根因（Vulkan GTT 回落）</h3>
<p dir="auto">2026-09-30 发现 <code>llama.sh 38</code> 慢到不可用（decode 0.49 tok/s），根因不是 shim、不是投机解码，<br />
而是**<code>-c 262144</code> 把权重从显存挤进了 GTT**。</p>
<p dir="auto">机制：模型权重 16.4GB + 256K 的 KV cache ≈ 26.3GB，超出 25.75GB 显存在 DEVICE_LOCAL 上<br />
能分配到的额度，<code>vkAllocateMemory</code> 失败后 Vulkan 后端<strong>静默改用 GTT（系统内存）映射</strong>。<br />
服务照常起来、<code>/health</code> 报 ok、<code>n_ctx</code> 也如实报 262144 —— 但每个 token 都要走 PCIe。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th><code>-c 262144</code></th>
<th><code>-c 131072</code></th>
</tr>
</thead>
<tbody>
<tr>
<td><code>vis_vram_used</code></td>
<td>567 MiB</td>
<td>权重完整驻留</td>
</tr>
<tr>
<td><code>gtt_used</code></td>
<td>23490 MiB</td>
<td>~2.5 GB（KV）</td>
</tr>
<tr>
<td>decode</td>
<td><strong>0.49 tok/s</strong></td>
<td>正常（数十 tok/s）</td>
</tr>
<tr>
<td>prefill</td>
<td>30 t/s</td>
<td>400+ t/s</td>
</tr>
</tbody>
</table>
<p dir="auto">日志里的特征行：<code>print_timing: ... 0.49 tokens per second</code>（decode 阶段）、<br />
<code>prompt processing ... 30.08 tokens per second</code>（prefill 阶段）。</p>
<p dir="auto"><strong>"多分配的 KV 不花钱"是错的。</strong> 只有在 KV 仍能装进 DEVICE_LOCAL 时这句话才成立；<br />
一旦溢出到 GTT，代价是每个 token 都走系统内存带宽。换句话说，<strong>判断依据是"装不装得下"，<br />
不是"填不填得满"</strong>。</p>
<p dir="auto">显存扫描（<code>-cram -1</code> 下，KV 在系统内存，只算权重+运行时开销）：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>ctx</th>
<th>显存占用</th>
<th>备注</th>
</tr>
</thead>
<tbody>
<tr>
<td>32K</td>
<td>~17.4 GB</td>
<td>安全</td>
</tr>
<tr>
<td>128K</td>
<td>~21.2 GB</td>
<td><strong>当前默认，留 ~4.5 GB 余量</strong></td>
</tr>
<tr>
<td>256K</td>
<td>~26.3 GB</td>
<td>爆，触发 GTT 回落</td>
</tr>
</tbody>
</table>
<p dir="auto">下界 64K：Hermes Agent 硬性要求 <code>context &gt;= 64000</code>，低于此直接拒绝连接。</p>
<p dir="auto">顺带加了两个参数：</p>
<ul>
<li><code>-cram -1</code>：KV cache 显式放系统内存，不与权重抢 DEVICE_LOCAL。<br />
对混合架构（本模型仅 16 层真注意力）代价很小，却能给权重留出余量。</li>
<li><code>P38_CTX</code> 默认 131072：<code>CTX=65536 ./llama.sh 38</code> 可临时覆盖，<br />
但<strong>每次加大都必须重测 <code>llama.sh status</code> 的显存占用</strong>，逼近 ~24GB 时立刻退回。</li>
</ul>
<p dir="auto"><code>P36_CTX</code> 仍是 196608，因为 Qwen3.6 的实测配置一直如此；若 3.6 也出现 0.49 tok/s，<br />
按同样逻辑降级即可。</p>
<blockquote>
<p dir="auto"><strong>教训</strong>：llama.cpp 的性能参数（<code>-c</code> / <code>-ngl</code> / <code>-ub</code>）互相耦合，改任何一个都要<br />
看最终<strong>显存落点</strong>（<code>vis_vram_used</code> / <code>gtt_used</code>），不能只看单个参数的账面数字。<br />
相关实测记录见 <code>~/llama-bin/OPTIMIZATION.md</code>。</p>
</blockquote>
<hr />
<h2>10. 附录：文件清单</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>文件</th>
<th>用途</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>~/llama-bin/llama.sh</code></td>
<td>模型服务统一入口（36/38 启停、切换、状态；k4v 参数见 §6.3）</td>
</tr>
<tr>
<td><code>~/llama-bin/OPTIMIZATION.md</code></td>
<td><strong>推理性能优化完整记录</strong>（显存账、ubatch、prefill 曲线、decode 溯源、各外部方案复核）</td>
</tr>
<tr>
<td><code>~/llama-bin/dsh-llm-shim.mjs</code></td>
<td>shim 本体</td>
</tr>
<tr>
<td><code>~/llama-bin/shim.sh</code></td>
<td><code>start</code> / <code>stop</code> / <code>restart</code> / <code>status</code>（风格与 <code>llama.sh</code> 一致）</td>
</tr>
<tr>
<td><code>~/llama-bin/shim-parser-test.mjs</code></td>
<td>解析逻辑离线单测（7 项，改动后跑一下）</td>
</tr>
<tr>
<td><code>~/llama-bin/dsh-shim-overlay.yml</code></td>
<td>headless 端到端验证用的 <code>--patch</code> 覆盖层（默认 <code>qwen3.8-27b</code>）</td>
</tr>
<tr>
<td><code>~/llama-bin/dsh-direct-overlay.yml</code></td>
<td>同上的直连版（<code>:8080</code>，无 shim），供 A/B 用</td>
</tr>
<tr>
<td><code>~/llama-bin/bench.py</code></td>
<td>性能测试脚本（第 6 章）</td>
</tr>
<tr>
<td><code>~/.config/qwen38/api.env</code></td>
<td>服务端密钥（独立文件，不进脚本、不进 shell 环境）</td>
</tr>
<tr>
<td><code>~/.config/systemd/user/qwen38.service</code></td>
<td>systemd 单元，<code>ExecStart=llama.sh 38 --fg</code>（改 <a href="http://llama.sh" rel="nofollow ugc">llama.sh</a> 会自动继承）</td>
</tr>
<tr>
<td><code>~/llama-bin/llama.sh.bak-*</code> / <code>shim.sh.bak-*</code></td>
<td>改动前的脚本备份</td>
</tr>
<tr>
<td><code>~/文档/7900_dsh/README.md</code></td>
<td>本文档目录的维护说明与归档清单</td>
</tr>
<tr>
<td><code>~/文档/7900_dsh/local_model.sh.txt</code></td>
<td>§9.1「无限重复 tool_call」故障的现场抓取（<strong>非脚本，不可执行</strong>）</td>
</tr>
</tbody>
<tbody>
<tr>
<td>目录</td>
<td>内容</td>
</tr>
<tr>
<td>---</td>
<td>---</td>
</tr>
<tr>
<td><code>~/models/</code></td>
<td><code>Qwen3.8-27B-UD-Q4_K_M.gguf</code>、<code>Qwen3.6-27B-Q4_K_M-mtp.gguf</code>、<code>mmproj-model-bf16.gguf</code></td>
</tr>
<tr>
<td><code>~/.dsh/profiles/web/cordis.patch.yml</code></td>
<td>DSH provider 路由（<code>local</code> → <code>:8090/v1</code>）</td>
</tr>
<tr>
<td><code>/tmp/qwen38-vulkan.log</code> / <code>/tmp/qwen36-vulkan.log</code></td>
<td>llama.cpp 日志</td>
</tr>
<tr>
<td><code>/tmp/dsh-llm-shim.log</code></td>
<td>shim 日志（<code>SHIM_VERBOSE=1</code> 时记录每次解析结果与重试）</td>
</tr>
<tr>
<td><code>~/.dsh/logs/</code></td>
<td>DSH 日志</td>
</tr>
</tbody>
</table>
]]></description><link>https://lcz.me/topic/2015</link><generator>RSS for Node</generator><lastBuildDate>Fri, 02 Oct 2026 17:06:57 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/2015.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 30 Sep 2026 16:07:12 GMT</pubDate><ttl>60</ttl></channel></rss>