<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Proxmox VE 9.2+ Intel 核显 SR-IOV + AMD Radeon AI PRO R9700 直通+Ubuntu24.04部署-全方位评测（终章）]]></title><description><![CDATA[<h1>PVE 9.2 物理主机 + Ubuntu 24.04 虚拟机 llama.cpp 服务+AMD Radeon AI PRO R9700 本地推理平台全方位评测验收</h1>
<hr />
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>内容</th>
</tr>
</thead>
<tbody>
<tr>
<td>评测执行端</td>
<td>macOS 工作站 <code>192.168.67.203</code>（2.5GbE）</td>
</tr>
<tr>
<td>PVE 物理主机</td>
<td><code>192.168.67.251</code>（节点名 <code>R9700</code>，Proxmox VE 9.2.18，内核 <code>7.0.14-16-pve</code>）</td>
</tr>
<tr>
<td>推理虚拟机</td>
<td><code>192.168.67.202</code>（VMID 102，Ubuntu 24.04.5，内核 <code>7.0.0-31-generic</code>）</td>
</tr>
<tr>
<td>加速卡</td>
<td>AMD Radeon AI PRO R9700（<code>gfx1201</code>，Navi 48，32 GB GDDR6，PCIe 5.0 x16）</td>
</tr>
<tr>
<td>推理服务</td>
<td><code>qwen38.service</code> → llama.cpp <code>llama-server</code> 0.4.0-dev，监听 <code>0.0.0.0:8080</code></td>
</tr>
<tr>
<td>模型</td>
<td><code>Qwen3.8-27B-Uncensored-Q5_K_M.gguf</code>（27.32 B 参数，Q5_K_M，128K 上下文）</td>
</tr>
<tr>
<td>评测维度</td>
<td>健康度 / 稳定性 / 功耗与散热 / 局域网可用性与性能 / 安全性 / 可运维性</td>
</tr>
<tr>
<td>评测方式</td>
<td>SSH 只读采集 + 局域网端到端压测 + 1 Hz 功耗温度采样</td>
</tr>
</tbody>
</table>
<p dir="auto">本报告所有数据均通过局域网络凭据只读采集。</p>
<hr />
<h1>一、总览</h1>
<hr />
<h2>1.1 结论</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>维度</th>
<th>实测结论</th>
<th>评级</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>PVE 主机健康度</strong></td>
<td>0 个失败单元、0 次 MCE/硬件错误、0 次 CPU 热节流、0 次 PCIe AER，运行 1 小时 51 分无异常</td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><strong>虚拟机健康度</strong></td>
<td>0 个失败单元、0 次 OOM、0 次 swap 换入换出、0 次 GPU 重置，内存压力指标≈0</td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><strong>GPU 健康度</strong></td>
<td>MEM ECC 已启用，UMC RAS 计数 <strong>CE=0 / UE=0 / DE=0</strong>，无坏页；满载 5 分钟无任何错误</td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><strong>推理服务稳定性</strong></td>
<td><code>NRestarts=0</code>，近 7 天日志 <strong>0 条 error/fail</strong>；两次持续满载（252 s + 304 s）<strong>0 错误、吞吐零衰减</strong></td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><strong>推理性能</strong></td>
<td>单流生成 <strong>52.7 token/s</strong>（σ=0.01）；预填充峰值 <strong>949 token/s</strong>；首 token 中位 <strong>145 ms</strong></td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><strong>局域网时延/吞吐</strong></td>
<td>RTT 中位 <strong>0.69 ms</strong>、丢包 0%；双向吞吐 <strong>293 MB/s（2.34 Gbit/s）</strong>，达 2.5GbE 线速的 94%</td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><strong>功耗控制</strong></td>
<td>空闲 GPU 12 W / CPU 21 W；满载 GPU 299 W（功耗墙 300 W）/ CPU 47 W</td>
<td><strong>良</strong></td>
</tr>
<tr>
<td><strong>散热</strong></td>
<td>GPU 结温稳态 <strong>89 °C</strong>（峰值 97 °C）、<strong>显存温度 92 °C</strong>（峰值 94 °C）</td>
<td><strong>需关注</strong></td>
</tr>
<tr>
<td><strong>多客户端并发能力</strong></td>
<td><code>-np 1</code>：请求被<strong>严格串行化</strong>，聚合吞吐恒定约 49 token/s，长请求会阻塞全部其他客户端 （家庭独占）</td>
<td><strong>差</strong></td>
</tr>
</tbody>
</table>
<hr />
<h2>1.2 七个最关键的事实</h2>
<ol>
<li>
<p dir="auto"><strong>性能非常扎实，且完全稳定。</strong> 单流生成稳定在 <strong>52.4–52.7 token/s</strong>（多次测量标准差仅 0.01–0.17）；5 分钟持续满载 78 个请求、0 错误、吞吐零衰减。对比服务日志中长上下文场景下的 30–43 token/s，可见<strong>短上下文性能有明显优势，长上下文会有衰减</strong>（见 <a href="#63-%E9%95%BF%E4%B8%8A%E4%B8%8B%E6%96%87%E9%A2%84%E5%A1%AB%E5%85%85%E8%83%BD%E5%8A%9B%E5%85%B3%E9%94%AE%E7%9F%AD%E6%9D%BF">6.3</a>）。</p>
</li>
<li>
<p dir="auto"><strong>最大的架构短板是 <code>-np 1</code>（单 slot），不是 GPU。</strong> 并发 1→16 时聚合吞吐几乎恒定在 <strong>47.5→49.0 token/s</strong>，说明请求被严格串行排队。实测一个 10 token 的小请求，在 16,200 token 长请求之后排队时，耗时从 <strong>0.40 s 涨到 19.78 s（49 倍）</strong>。多人/多 agent 同时使用同一服务时体验会显著恶化。</p>
</li>
<li>
<p dir="auto"><strong>长上下文预填充是"分钟级"的。</strong> 98,304 token 的输入需要 <strong>261 秒</strong>；按实测速率外推，占满 128K 上下文约需 <strong>5.8 分钟</strong>。在此期间 <code>/health</code> 仍返回 200，但推理接口完全不可用——<strong>健康检查不能代表服务可用性</strong>。</p>
</li>
<li>
<p dir="auto"><strong>GPU 功耗正常但温度偏高。</strong> R9700 空闲 12 W、满载 <strong>299 W（正好打满 300 W 功耗墙，峰值 304 W，曾观测到 354 W 瞬时尖峰）</strong>；结温稳态 89 °C、峰值 <strong>97 °C</strong>，<strong>显存温度稳态 92 °C、峰值 94 °C</strong>——显存温度是本机散热的最热点。</p>
</li>
<li>
<p dir="auto"><strong>数据安全存在真实风险：</strong>  <strong><code>AI服务器整块 1 TB 物理 NVMe 裸盘直通 （方案：停机备份整块盘）</code>。</strong></p>
</li>
<li>
<p dir="auto"><strong>访问控制真实风险：</strong>  <strong><code>局域网内任何设备都可以直接占用这块 300 W 的 GPU</code></strong></p>
</li>
<li>
<p dir="auto"><strong>主机与虚拟机的硬件健康度都很好。</strong> 0 次 MCE、0 次 AER、0 次 CPU 节流、内存压力≈0；GPU ECC 零错误。平台本身的可靠性没有问题。</p>
</li>
</ol>
<hr />
<h1>二、评测环境与拓扑</h1>
<hr />
<h2>2.1 物理拓扑</h2>
<pre><code>        ┌──────────────────────────────────────────────────────────────┐
        │  评测客户端  macOS 工作站   192.168.67.203                    │
        │  网卡 en1：2500Base-T 全双工   （icmp RTT 中位 0.37 ms）        │
        └───────────────────────────┬──────────────────────────────────┘
                                    │  2.5GbE 交换网络（同一广播域）
        ┌───────────────────────────▼──────────────────────────────────┐
        │  PVE 9.2.18 物理主机  192.168.67.251   节点名 R9700            │
        │  Intel Core Ultra 7 265K / 20 核 / 62 GiB DDR5-6400（非 ECC）  │
        │  ASUS TUF GAMING Z890-PRO WIFI  BIOS 3211                     │
        │  启动盘 Samsung 512GB NVMe → LVM(pve-root 467.94G + swap 8G)  │
        │  数据盘 Kingston 1TB NVMe →【整块裸盘直通】                    │
        │  上行 enp135s0 2500Mb/s → vmbr0（无 BMC/IPMI，防火墙已禁用）    │
        └───────────────┬──────────────────────────┬───────────────────┘
                        │ KVM + VFIO 直通          │
        ┌───────────────▼──────────────┐  ┌────────▼─────────┐
        │ VM 102  Ubuntu24.04-desktop  │  │ VM 100 Ikuai8    │
        │ 192.168.67.202               │  │ VM 101 OpenWrt   │
        │ 10 vCPU(host) / 48 GiB       │  │      停止待用     │
        │ 根盘 = Kingston 整块直通      │  └──────────────────┘
        │ GPU 04:00.0 R9700 直通        │
        │ qwen38.service :8080          │
        └──────────────────────────────┘
</code></pre>
<h2>2.2 PVE 物理主机</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>实测值</th>
</tr>
</thead>
<tbody>
<tr>
<td>主机名 / 系统</td>
<td><code>R9700</code> / Debian GNU/Linux 13.6 (trixie)</td>
</tr>
<tr>
<td>PVE 版本</td>
<td>pve-manager <strong>9.2.18</strong>，proxmox-ve 9.2.0，qemu-server 9.2.7，pve-qemu-kvm 11.0.3-3</td>
</tr>
<tr>
<td>内核</td>
<td><code>7.0.14-16-pve</code>（另留存 6.14.11-9、6.14.8-2 备用内核）</td>
</tr>
<tr>
<td>主板 / BIOS</td>
<td>ASUSTeK <strong>TUF GAMING Z890-PRO WIFI</strong> Rev 1.xx / BIOS <strong>3211</strong>（2026-07-28）</td>
</tr>
<tr>
<td>CPU</td>
<td>Intel Core Ultra 7 <strong>265K</strong>，20 核 / 20 线程，800–5500 MHz，L3 30 MiB</td>
</tr>
<tr>
<td>内存</td>
<td><strong>62 GiB</strong>（2 × 32 GB <strong>Kingston KF564C32-32</strong> DDR5-<strong>6400</strong> MT/s @1.4 V），<strong>非 ECC</strong>，4 槽仅用 2 槽</td>
</tr>
<tr>
<td>启动存储</td>
<td>Samsung <strong>MZAMX512HCLV-00BL2</strong> 512 GB → LVM：<code>pve-root</code> 467.94 G(ext4) + <code>pve-swap</code> 8 G</td>
</tr>
<tr>
<td>数据存储</td>
<td>Kingston <strong>SNV2S1000G</strong> 1 TB（<strong>整块直通给 VM 102</strong>，<strong>AI推理服务器</strong>）</td>
</tr>
<tr>
<td>存储方案</td>
<td>LVM + ext4，未使用 ZFS；根分区已用 12 G / 461 G（3%）</td>
</tr>
<tr>
<td>上行链路</td>
<td><code>enp135s0</code>（1c:86:0b:2f:1a:16）<strong>2500 Mb/s 全双工</strong>，桥接至 <code>vmbr0</code></td>
</tr>
<tr>
<td>未用网口</td>
<td><code>eno1</code> / <code>enp136s0</code> / <code>enp137s0</code> / <code>enp138s0</code>（NO-CARRIER）、<code>wlp131s0</code>（WiFi 未启用）</td>
</tr>
<tr>
<td>内核启动参数</td>
<td><code>intel_iommu=on i915.enable_guc=3 i915.max_vfs=4 module_blacklist=xe</code></td>
</tr>
<tr>
<td>调频策略</td>
<td><code>intel_pstate=active</code>，governor=<strong>performance</strong>，EPP=default，<code>no_turbo=0</code>（睿频启用），max_perf_pct=100</td>
</tr>
<tr>
<td>深度 C 态</td>
<td><code>intel_idle</code> + <code>menu</code>，C3 驻留占运行时长约 <strong>83%</strong></td>
</tr>
<tr>
<td>热节流计数</td>
<td>core = <strong>0</strong>，package = <strong>0</strong>（从未节流）</td>
</tr>
<tr>
<td>传感器</td>
<td><code>coretemp</code>（封装 39–40 °C 空闲）、2× NVMe、<code>asus</code> hwmon（<strong>无风扇转速通道</strong>）、<code>acpitz</code>；<strong>无 <code>turbostat</code></strong></td>
</tr>
<tr>
<td>虚拟化</td>
<td>KVM 运行中，仅 VM 102 在跑；KSM 已启用（run=1）</td>
</tr>
<tr>
<td>防火墙</td>
<td><strong><code>pve-firewall status</code> → <code>disabled/running</code></strong>；iptables 全链默认 ACCEPT</td>
</tr>
<tr>
<td>备份</td>
<td><strong>无 <code>jobs.cfg</code>、无 vzdump cron、<code>/var/lib/vz/dump</code> 为空 → 无任何备份</strong></td>
</tr>
<tr>
<td>时间同步</td>
<td>chrony 已同步（Stratum 3，系统时间偏差 1.9 ms）</td>
</tr>
<tr>
<td>运行时长</td>
<td>1 小时 51 分（本次启动 2026-09-11 20:31）</td>
</tr>
</tbody>
</table>
<hr />
<h2>2.3 Ubuntu 24.04 推理虚拟机（VM 102）</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>实测值</th>
</tr>
</thead>
<tbody>
<tr>
<td>VMID / 名称</td>
<td><strong>102</strong> / <code>Ubuntu24.04-desktop</code></td>
</tr>
<tr>
<td>主机名 / 系统</td>
<td><code>magicz890-Ai</code> / <strong>Ubuntu 24.04.5 LTS</strong>（Noble），内核 <code>7.0.0-31-generic</code></td>
</tr>
<tr>
<td>虚拟化</td>
<td><code>systemd-detect-virt</code> → <code>kvm</code>；机型 <code>pc-q35-11.0</code>，UEFI(OVMF)，<code>cpu: host</code></td>
</tr>
<tr>
<td>计算资源</td>
<td><strong>10 vCPU</strong> / <strong>48 GiB</strong> 内存（49152 MB），swap 8 GiB（文件，使用 0）</td>
</tr>
<tr>
<td>内存实测</td>
<td>总计 46 Gi，已用 13 Gi，可用 33 Gi；<code>pswpin/pswpout = 0</code>，<code>oom_kill = 0</code></td>
</tr>
<tr>
<td>GPU 直通</td>
<td><code>hostpci0: 0000:04:00.0,pcie=1</code> → 客体内 <code>01:00.0</code>，驱动 <code>amdgpu</code> 3.64.0</td>
</tr>
<tr>
<td>显卡规格</td>
<td>AMD <strong>Radeon AI PRO R9700</strong>（<code>1002:7551</code>，Navi 48，Sapphire），<strong>VRAM 30,576 MiB = 29.86 GiB</strong></td>
</tr>
<tr>
<td>链路</td>
<td><strong>PCIe 5.0 x16（32.0 GT/s ×16）</strong></td>
</tr>
<tr>
<td>ECC</td>
<td><code>MEM ECC is active.</code> / <code>GECC is currently enabled, which may affect performance</code></td>
</tr>
<tr>
<td>存储</td>
<td>客体 <code>sda</code> 931.5 G（<code>QEMU HARDDISK</code>，virtio-scsi）→ ext4 <code>/</code> 915 G，已用 80 G（10%）</td>
</tr>
<tr>
<td>网络</td>
<td><code>enp6s18</code> virtio（<code>bc:24:11:ac:b1:1c</code>），<strong>192.168.67.202/24</strong>，MTU 1500，单 RX/TX 队列</td>
</tr>
<tr>
<td>ROCm</td>
<td><strong>10.0.0~pre4</strong>（gfx1201 专用包），HIP <strong>7.15.26333</strong>，AMD clang 23.0.0git</td>
</tr>
<tr>
<td>llama.cpp</td>
<td><strong>0.4.0-dev (build 1, commit df03399)</strong>，GNU 13.3.0 编译，<code>/opt/llama.cpp/</code> 共 45 GB</td>
</tr>
<tr>
<td>远程桌面</td>
<td><code>gnome-remote-desktop</code> 监听 <strong>3389 / 3390</strong>（对全网段开放）</td>
</tr>
<tr>
<td>其他服务</td>
<td>GNOME 桌面、Xorg、Chrome、<code>dsh web</code>（127.0.0.1:3080）、cupsd</td>
</tr>
<tr>
<td>防火墙</td>
<td><code>ufw</code> <strong>未启用</strong>；iptables 全链 ACCEPT；<strong>未安装 fail2ban</strong> （测试机后备）</td>
</tr>
<tr>
<td>运行时长</td>
<td>1 小时 51 分（本次启动 2026-09-11 20:31:47）</td>
</tr>
</tbody>
</table>
<hr />
<h2>2.4 llama.cpp 服务配置</h2>
<pre><code>/etc/systemd/system/qwen38.service
  Description = Qwen3.8-27B llama.cpp server (R9700 HIP, MTP, 128K)
  User/Group  = magicz890
  Environment = LD_LIBRARY_PATH=/opt/llama.cpp/bin:/opt/rocm/lib
                HIP_VISIBLE_DEVICES=0
  ExecStart   = /opt/llama.cpp/bin/llama-server
                  -m  /opt/llama.cpp/models/Qwen3.8-27B-Uncensored-Q5_K_M.gguf
                  --mmproj /opt/llama.cpp/models/mmproj-Qwen3.8-27B-Uncensored-F16.gguf
                  --alias qwen3.8
                  -c 131072 --override-kv qwen35.context_length=int:131072
                  -ctk q8_0 -ctv q8_0 -fa on --load-mode none
                  --reasoning-effort medium -ngl 99 -np 1
                  --temp 0.3 --top-p 0.9
                  --spec-type draft-mtp --spec-draft-n-max 2
                  --host 0.0.0.0 --port 8080
  Restart     = on-failure / RestartSec=5 / TimeoutStopSec=60
</code></pre>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>关键参数</th>
<th>值</th>
<th>含义与影响</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>-c 131072</code> + <code>-ctk/-ctv q8_0</code></td>
<td>128K 上下文，KV 量化 8 bit</td>
<td>KV 缓存约 <strong>6.17 GiB</strong>（≈49.3 KiB/token），显存占用 25.2 GiB</td>
</tr>
<tr>
<td><code>-fa on</code></td>
<td>Flash Attention 开启</td>
<td>长上下文必需，已正确启用</td>
</tr>
<tr>
<td><code>-ngl 99</code></td>
<td>全部层卸载到 GPU</td>
<td>正确（GPU 显存充足）</td>
</tr>
<tr>
<td><strong><code>-np 1</code></strong></td>
<td><strong>仅 1 个 slot</strong></td>
<td><strong>请求串行化，无并发能力</strong>（自用测试）</td>
</tr>
<tr>
<td><code>--spec-type draft-mtp</code> + <code>n-max 2</code></td>
<td>MTP 自投机解码</td>
<td>日志显示接受率可达 0.99，是 52 t/s 高吞吐的主因</td>
</tr>
<tr>
<td><code>--host 0.0.0.0</code></td>
<td>监听全部接口</td>
<td>无鉴权暴露到局域网</td>
</tr>
</tbody>
</table>
<p dir="auto">服务启动性能：<strong>模型加载耗时约 8.2 秒</strong>（20:31:49.898 → 20:31:58），对 18.19 GiB 模型而言非常快。</p>
<hr />
<h2>2.5 评测方法</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>说明</th>
</tr>
</thead>
<tbody>
<tr>
<td>采集方式</td>
<td>通过 expect 封装的 SSH（密码登录）执行<strong>只读</strong>采集脚本；未修改任何配置、服务、文件</td>
</tr>
<tr>
<td>压测方式</td>
<td>从 macOS 客户端 <code>192.168.67.203</code> 通过局域网调用 <code>http://192.168.67.202:8080</code> 的 OpenAI 兼容与原生命令接口</td>
</tr>
<tr>
<td>功耗采集</td>
<td>VM 内 1 Hz 读取 <code>amdgpu</code> hwmon（PPT 功率、三路温度、风扇、频率、显存）；主机 1 Hz 读取 Intel RAPL 能耗计数器与 <code>coretemp</code></td>
</tr>
<tr>
<td><strong>环境</strong></td>
<td><strong>这是一台正在提供服务的生产机器。</strong> 评测期间该 llama 服务同时被 agent 工作流使用（日志可见 3 万 token 级上下文任务）。所有压测数据均在此共享背景下采集，代表"真实使用状态"而非实验室理想值</td>
</tr>
<tr>
<td><strong>副作用</strong></td>
<td>压测会<strong>冲掉 llama.cpp 的 prompt cache</strong>，可能使评测后第一个 agent 请求需要重新处理前缀；未重启、未停止任何服务；<code>n_predict</code> 超限探测请求在客户端 30 s 超时断开后，服务端正确取消了任务</td>
</tr>
<tr>
<td>未做事项</td>
<td>未做断电/掉电测试、未做 8 小时以上长稳、未做 PCIe 错误注入、未做多卡/NVLink 测试</td>
</tr>
</tbody>
</table>
<hr />
<h1>三、健康度评测</h1>
<hr />
<h2>3.1 PVE 主机健康度</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>检查项</th>
<th>实测结果</th>
<th>判定</th>
</tr>
</thead>
<tbody>
<tr>
<td>失败 systemd 单元</td>
<td><strong>0 个</strong></td>
<td>正常</td>
</tr>
<tr>
<td>核心服务状态</td>
<td><code>pveproxy</code> / <code>pvedaemon</code> / <code>pve-cluster</code> / <code>qmeventd</code> 均 active（<code>corosync</code> inactive 属正常单节点）</td>
<td>正常</td>
</tr>
<tr>
<td>MCE / 机器检查</td>
<td><code>dmesg</code> 匹配 <code>mce|machine check|hardware error</code> = <strong>0 条</strong></td>
<td>正常</td>
</tr>
<tr>
<td>PCIe AER</td>
<td>仅 10 条 <code>AER: enabled with IRQ</code> 初始化信息，<strong>无 AER 错误</strong></td>
<td>正常</td>
</tr>
<tr>
<td>EDAC 内存错误</td>
<td>未暴露 <code>ce_count/ue_count</code>（消费级平台无 ECC 也无可读计数器）</td>
<td>不可观测</td>
</tr>
<tr>
<td>CPU 热节流</td>
<td><code>core_throttle_count = 0</code>，<code>package_throttle_count = 0</code>（全部 20 核）</td>
<td>优</td>
</tr>
<tr>
<td>内核错误级日志</td>
<td>仅 6 条，全部为无害项（见下）</td>
<td>正常</td>
</tr>
<tr>
<td>内存压力</td>
<td>PSI memory some/full total = 5 µs，<strong>基本为 0</strong></td>
<td>优</td>
</tr>
<tr>
<td>交换</td>
<td><code>pswpin/pswpout = 0</code>，swap 使用 0 B</td>
<td>优</td>
</tr>
<tr>
<td>OOM</td>
<td><code>oom_kill = 0</code></td>
<td>优</td>
</tr>
<tr>
<td>负载</td>
<td>load1 = 0.14（1 分钟）/ 1.10（评测后），20 核机器非常空闲</td>
<td>优</td>
</tr>
<tr>
<td>时间同步</td>
<td>chrony Stratum 3，系统时间偏差 <strong>+1.9 ms</strong></td>
<td>优</td>
</tr>
</tbody>
</table>
<hr />
<h2>3.2 虚拟机健康度</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>检查项</th>
<th>实测结果</th>
<th>判定</th>
</tr>
</thead>
<tbody>
<tr>
<td>失败 systemd 单元</td>
<td><strong>0 个</strong>（56 个 running 单元）</td>
<td>正常</td>
</tr>
<tr>
<td>内核错误级日志</td>
<td>9 条，全部为开机瞬时告警（<code>shpchp ... pci_hp_register failed</code>、<code>snd_hda_intel: no codecs found!</code>）</td>
<td>基本无害</td>
</tr>
<tr>
<td><code>amdgpu</code> 错误</td>
<td><strong>0 条</strong> ring timeout / GPU reset / page fault</td>
<td>优</td>
</tr>
<tr>
<td>CPU 热节流</td>
<td>全部 vCPU <code>core/package_throttle_count = 0</code></td>
<td>优</td>
</tr>
<tr>
<td>内存压力</td>
<td>PSI memory some total ≈ 1 µs，full ≈ 1 µs</td>
<td>优</td>
</tr>
<tr>
<td>交换活动</td>
<td><code>pswpin/pswpout = 0</code>，swap 使用 0 B</td>
<td>优</td>
</tr>
<tr>
<td>OOM</td>
<td><code>oom_kill = 0</code></td>
<td>优</td>
</tr>
<tr>
<td>磁盘</td>
<td><code>/</code> 已用 80 G / 915 G（10%），无 md RAID 活动</td>
<td>正常</td>
</tr>
<tr>
<td>待重启</td>
<td><code>/var/run/reboot-required</code> 不存在</td>
<td>正常</td>
</tr>
</tbody>
</table>
<hr />
<h2>3.3 GPU 健康度</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>检查项</th>
<th>实测结果</th>
<th>判定</th>
</tr>
</thead>
<tbody>
<tr>
<td>型号识别</td>
<td><code>AMD Radeon AI PRO R9700</code>，Device ID <code>0x7551</code>，SKU <code>1E4990U</code>，GUID 23946</td>
<td>正常</td>
</tr>
<tr>
<td>GFX 架构</td>
<td><code>gfx1201</code>（rocminfo 同时暴露 <code>gfx12-generic</code>）</td>
<td>正常</td>
</tr>
<tr>
<td>VBIOS</td>
<td><code>113-1E4990U-S83</code></td>
<td>—</td>
</tr>
<tr>
<td>驱动</td>
<td><code>amdgpu 3.64.0</code>，SMU fw <code>104.76.0</code></td>
<td>正常</td>
</tr>
<tr>
<td>显存</td>
<td>总 <strong>30,576 MiB（29.86 GiB）</strong>，BAR 32768M，256-bit GDDR6</td>
<td>正常</td>
</tr>
<tr>
<td>显存 ECC</td>
<td><strong><code>MEM ECC is active.</code></strong> / GECC enabled</td>
<td>优</td>
</tr>
<tr>
<td>RAS 模块</td>
<td><code>UMC: ENABLED</code>、<code>DF: ENABLED</code>；SDMA/GFX/MMHUB/ATHUB/PCIE_BIF 等 <strong>DISABLED</strong></td>
<td>部分可用</td>
</tr>
<tr>
<td><strong>ECC 计数</strong></td>
<td><code>umc_err_count</code> → <strong>ue: 0，ce: 0，de: 0</strong></td>
<td><strong>优</strong></td>
</tr>
<tr>
<td>显存坏页</td>
<td><code>gpu_vram_bad_pages</code> 为空</td>
<td>优</td>
</tr>
<tr>
<td>链路状态</td>
<td>PCIe <strong>32.0 GT/s ×16</strong>（Gen5 x16），驱动 CfgCurrentLinkSpeed 16.0GT/s x16</td>
<td>优</td>
</tr>
<tr>
<td>计算单元</td>
<td>SE 4 × SH 2 × CU 8 = <strong>64 CU</strong>（active_cu_number 64）</td>
<td>正常</td>
</tr>
<tr>
<td>满载 5 分钟后复检</td>
<td>ECC 仍为 0/0/0，无 GPU reset，无 ring timeout</td>
<td><strong>优</strong></td>
</tr>
<tr>
<td>直通告警</td>
<td>QEMU 启动时 <code>vfio_container_dma_map(...) = -22 (Invalid argument)</code> 与 <code>0000:04:00.0: PCI peer-to-peer transactions on BARs are not supported.</code></td>
<td>需关注</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong>关于直通告警</strong>：QEMU 在建立 DMA 映射时出现一次 <code>-22</code>（EINVAL），并提示该卡不支持 BAR 级 P2P 事务。这两条是 <strong>VFIO 大 BAR（32 GB）映射的常见提示</strong>，在单卡直通、且不需要 GPU 间 P2P（NVLink/XGMI）的场景下<strong>不影响功能与稳定性</strong>——本次 5 分钟满载 0 错误即为佐证。</p>
</blockquote>
<h2>3.4 服务健康度</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>检查项</th>
<th>实测值</th>
</tr>
</thead>
<tbody>
<tr>
<td>服务状态</td>
<td><code>active (running)</code>，<code>enabled</code>（开机自启）</td>
</tr>
<tr>
<td><strong>重启次数 <code>NRestarts</code></strong></td>
<td><strong>0</strong></td>
</tr>
<tr>
<td>本次运行时长</td>
<td>1 小时 50 分（20:31:49 起）</td>
</tr>
<tr>
<td>主进程 PID</td>
<td>1244（<code>llama-server</code>）</td>
</tr>
<tr>
<td>内存占用</td>
<td>当前 29.5 G，<strong>峰值 36.4 G</strong>；RSS 17.9 GB</td>
</tr>
<tr>
<td>CPU 累计</td>
<td>28 分 12 秒</td>
</tr>
<tr>
<td>线程数</td>
<td>16</td>
</tr>
<tr>
<td>相关告警</td>
<td>启动时 <code>CORS is set to allow all origins ('*') and no API key is set</code>、<code>Qwen-VL models require at minimum 1024 image tokens</code></td>
</tr>
<tr>
<td>文件描述符限制</td>
<td><code>LimitNOFILE</code> 软限 <strong>1024</strong> / 硬限 524288 <strong>（需优先处理）</strong></td>
</tr>
</tbody>
</table>
<hr />
<h1>四、稳定性评测</h1>
<hr />
<h2>4.1 运行时长与重启历史</h2>
<p dir="auto"><strong>PVE 主机</strong>：今天的重启<strong>全部是我修改参数，不是崩溃</strong>。</p>
<pre><code>reboot   system boot  7.0.14-16-pve    Fri Sep 11 20:31 - still running
reboot   system boot  7.0.14-16-pve    Fri Sep 11 10:42 - 20:30  (09:48)
shutdown system down  7.0.14-16-pve    Fri Sep 11 20:30 - 20:31  (00:00)
reboot   system boot  7.0.14-16-pve    Fri Sep 11 09:02 - 10:15  (01:12)
shutdown system down  7.0.14-16-pve    Fri Sep 11 10:15 - 10:42  (00:27)
reboot   system boot  7.0.14-16-pve    Fri Sep 11 07:30 - 07:31  (00:00)
...
</code></pre>
<p dir="auto">每条 <code>reboot</code> 前都对应一条 <code>shutdown ... system down</code>，<strong>7 次启动、0 次非正常掉电</strong>。<code>journalctl --list-boots</code> 与 <code>last -x</code> 完全一致。</p>
<p dir="auto"><strong>虚拟机</strong>：同样全部为正常关机（<code>shutdown system down</code> → <code>reboot system boot</code>），最快一次仅停机 1 分钟。</p>
<p dir="auto"><strong>llama 服务</strong>：<code>NRestarts=0</code>，与主机重启时间严格对齐——主机重启时服务随虚拟机自启，未发生任何崩溃重启。</p>
<hr />
<h2>4.2 长稳压测结果</h2>
<p dir="auto">执行了<strong>两轮</strong>持续满载推理，均在 1 Hz 功耗采样覆盖下完成：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>轮次</th>
<th>时长</th>
<th>并发</th>
<th>完成请求</th>
<th>生成 token</th>
<th><strong>错误数</strong></th>
<th>单请求速度</th>
<th>标准差</th>
</tr>
</thead>
<tbody>
<tr>
<td>第一轮</td>
<td><strong>304 s</strong></td>
<td>2</td>
<td>78</td>
<td>14,976</td>
<td><strong>0</strong></td>
<td>52.46 t/s</td>
<td><strong>0.17</strong></td>
</tr>
<tr>
<td>第二轮</td>
<td><strong>252 s</strong></td>
<td>1</td>
<td>63</td>
<td>12,096</td>
<td><strong>0</strong></td>
<td>52.46 t/s</td>
<td><strong>0.13</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>关键结论：约 9.3 分钟累计满载、141 个请求，零错误、零重试、吞吐零衰减。</strong> 两轮的单请求速度中位数完全一致（52.46 t/s），说明 GPU 在持续满载下<strong>没有出现降频或性能衰减</strong>。</p>
<p dir="auto">请求延迟（192 token/请求）：中位 7.56 s，p99 7.84 s，标准差 0.46 s——排队行为可预测。</p>
<hr />
<h2>4.3 满载后的状态复检</h2>
<p dir="auto">压测结束后立即复检（22:22:44），结果：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>检查项</th>
<th>压测前</th>
<th>压测后</th>
<th>判定</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>NRestarts</code></td>
<td>0</td>
<td><strong>0</strong></td>
<td>优</td>
</tr>
<tr>
<td>GPU ECC（UE/CE/DE）</td>
<td>0/0/0</td>
<td><strong>0/0/0</strong></td>
<td>优</td>
</tr>
<tr>
<td>显存坏页</td>
<td>无</td>
<td><strong>无</strong></td>
<td>优</td>
</tr>
<tr>
<td>内核错误数</td>
<td>9</td>
<td><strong>9</strong>（无新增）</td>
<td>优</td>
</tr>
<tr>
<td>GPU reset / ring timeout</td>
<td>0</td>
<td><strong>0</strong></td>
<td>优</td>
</tr>
<tr>
<td>CPU 热节流计数</td>
<td>0/0</td>
<td><strong>0/0</strong></td>
<td>优</td>
</tr>
<tr>
<td>内核错误级日志</td>
<td>6 条</td>
<td><strong>6 条</strong>（无新增）</td>
<td>优</td>
</tr>
<tr>
<td>GPU 状态回落</td>
<td>12 W</td>
<td>18 W → 随后回落到 <strong>12 W</strong></td>
<td>优</td>
</tr>
<tr>
<td>GPU 温度回落</td>
<td>32 °C 边</td>
<td>66 °C 结温 → 冷却中</td>
<td>正常</td>
</tr>
</tbody>
</table>
<hr />
<h2>4.4 温控稳定性</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>阶段</th>
<th>GPU 结温</th>
<th>GPU 显存温度</th>
<th>CPU 封装温度</th>
<th>风扇转速</th>
<th>是否降频</th>
</tr>
</thead>
<tbody>
<tr>
<td>开机空闲</td>
<td>34 °C</td>
<td>32 °C</td>
<td>39 °C</td>
<td>891 rpm（最低）</td>
<td>否（SCLK 41 MHz 为深度空闲档）</td>
</tr>
<tr>
<td>满载 1 分钟</td>
<td>88 °C</td>
<td>90 °C</td>
<td>47 °C</td>
<td>2,370 rpm</td>
<td><strong>否</strong>（SCLK 2788 MHz）</td>
</tr>
<tr>
<td>满载稳态</td>
<td>88–90 °C</td>
<td>92–94 °C</td>
<td>51 °C</td>
<td>2,470–2,595 rpm</td>
<td><strong>否</strong>（SCLK 2730–3077 MHz）</td>
</tr>
<tr>
<td>峰值</td>
<td><strong>97 °C</strong></td>
<td><strong>94 °C</strong></td>
<td>62 °C</td>
<td>3,189 rpm</td>
<td><strong>否</strong></td>
</tr>
</tbody>
</table>
<p dir="auto">判定依据：整个压测过程 SCLK 未出现台阶式下降、生成吞吐标准差仅 0.13–0.17、CPU <code>core/package_throttle_count</code> 保持 0。<strong>平台没有热降频，但 GPU 显存温度已进入 90 °C 以上的"暖区"</strong>。</p>
<hr />
<h1>五、功耗与散热评测</h1>
<hr />
<h2>5.1 GPU 功耗与温度实测（AMD R9700）</h2>
<p dir="auto">数据来源：<code>/sys/class/drm/card1/device/hwmon/hwmon0/</code>（<code>power1_average</code>、<code>temp1/2/3_input</code>、<code>fan1_input</code>、<code>freq1/2_input</code>），1 Hz 采样，共 900 + 386 个样本。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>空闲（无客户端）</th>
<th>满载（单流持续生成）</th>
<th>功耗墙</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>PPT 封装功耗</strong></td>
<td><strong>12 W</strong>（中位/最小 11–12 W）</td>
<td><strong>299 W</strong>（中位）/ 均值 271 W / 峰值 <strong>304 W</strong></td>
<td><code>power1_cap = 300 W</code>，min 210 W</td>
</tr>
<tr>
<td>边缘温度 edge</td>
<td>32 °C</td>
<td>中位 <strong>72 °C</strong> / 峰值 73 °C</td>
<td>—</td>
</tr>
<tr>
<td><strong>结温 junction</strong></td>
<td>34 °C</td>
<td>中位 <strong>89 °C</strong> / p90 95 °C / 峰值 <strong>97 °C</strong></td>
<td>—</td>
</tr>
<tr>
<td><strong>显存温度 memory</strong></td>
<td>32 °C</td>
<td>中位 <strong>92 °C</strong> / p90 90 °C / 峰值 <strong>94 °C</strong></td>
<td>—</td>
</tr>
<tr>
<td>风扇转速</td>
<td>891 rpm（20%，最低档）</td>
<td>中位 <strong>2,473 rpm</strong> / 峰值 3,189 rpm（56%）</td>
<td>—</td>
</tr>
<tr>
<td>核心频率 SCLK</td>
<td>41 MHz</td>
<td>中位 <strong>2,822 MHz</strong> / 峰值 3,244 MHz</td>
<td>档位上限 2350 MHz（标称）/ 实测睿频可达 3.24 GHz</td>
</tr>
<tr>
<td>显存频率 MCLK</td>
<td>96 MHz</td>
<td><strong>1,258 MHz</strong>（最高档）</td>
<td>—</td>
</tr>
<tr>
<td>GPU 占用率</td>
<td>3%</td>
<td><strong>100%</strong></td>
<td>—</td>
</tr>
<tr>
<td>显存带宽占用率 mem_busy</td>
<td>0%</td>
<td>中位 <strong>63%</strong> / 峰值 87%</td>
<td>—</td>
</tr>
<tr>
<td>显存占用 VRAM</td>
<td>25.8 GB（模型+KV 常驻）</td>
<td>25.8–26.0 GB</td>
<td>总 29.86 GiB</td>
</tr>
<tr>
<td>供电电压 VDDGFX</td>
<td>1.0 mV（未激活）</td>
<td>—</td>
<td>—</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong>GPU 是绝对功耗大头</strong>：满载时 GPU 299 W vs CPU 47 W，<strong>GPU 占整机功耗约 70%，如家庭使用可降频到270W左右，影响甚微</strong>。</p>
</blockquote>
<blockquote>
<p dir="auto"><strong>空闲温度的测量口径</strong>：表中空闲值为<strong>开机后长时间无负载</strong>状态（本次为启动后 1 小时 33 分、无任何客户端）的实测值。压测结束冷却约 4 分钟后，三路温度回落到 <strong>37 / 39 / 38 °C</strong>（边缘/结温/显存），功耗回落到 <strong>11 W</strong>，风扇回到最低档 <strong>891 rpm</strong> —— 说明散热与功耗回退机制工作正常，只是机箱内部余温使读数略高于冷启动状态。</p>
</blockquote>
<hr />
<p dir="auto"><strong>GPU 功耗/温度时间线</strong></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/81bfa3e2-1cc2-4815-ad77-4ae8e6f338f3.jpeg" alt="44478484-3ae4-4c0b-b238-ab2113169d52-image.jpeg" class=" img-fluid img-markdown" /></p>
<hr />
<h2>5.2 主机 CPU 功耗与温度实测（Intel RAPL）</h2>
<p dir="auto">数据来源：<code>/sys/class/powercap/intel-rapl:0/energy_uj</code>（package）与 <code>intel-rapl:0:0</code>（core），差分法计算瓦特。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>空闲</th>
<th>满载</th>
<th>峰值</th>
</tr>
</thead>
<tbody>
<tr>
<td>CPU 封装功耗 package</td>
<td><strong>20.8 – 21.2 W</strong></td>
<td>中位 <strong>46.9 W</strong></td>
<td>69.2 W</td>
</tr>
<tr>
<td>CPU 核心功耗 core</td>
<td><strong>8.5 – 9.0 W</strong></td>
<td>中位 <strong>31.2 W</strong></td>
<td>50.4 W</td>
</tr>
<tr>
<td>CPU 封装温度</td>
<td>39–40 °C</td>
<td>中位 <strong>51 °C</strong></td>
<td>62 °C</td>
</tr>
<tr>
<td>NVMe0（Kingston，VM 根盘）</td>
<td>46.9 °C</td>
<td>47.9 °C</td>
<td>51.9 °C</td>
</tr>
<tr>
<td>NVMe1（Samsung，启动盘）</td>
<td>37.9 °C</td>
<td>36.9 °C</td>
<td>37.9 °C</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>CPU 功耗/温度时间线</strong></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/db8e5b5f-5ced-4b55-81c7-66c5e8d6935b.jpeg" alt="f5161587-e9ed-463a-84ba-3aa50834d29e-image.jpeg" class=" img-fluid img-markdown" /></p>
<blockquote>
<p dir="auto">由 RAPL 可见，<strong>CPU 在推理场景中几乎"闲着"</strong>：封装功耗仅从 21 W 升到 47 W（+26 W），远小于 GPU 的 +287 W。这与 <code>llama-server</code> 的 CPU 占用很低这一事实一致 —— <code>CPUUsageNSec</code> 折算其<strong>生命周期平均 CPU 占用约 25.4%</strong>（累计 1,692 s CPU 时间 / 6,655 s 运行时间 ÷ 10 vCPU ≈ <strong>2.5 个 vCPU</strong> 等效），且这部分开销主要来自采样与 HTTP 层，矩阵计算全部在 GPU 上完成。</p>
</blockquote>
<hr />
<h2>5.3 整机功耗估算（因消费级主板数据仅供参考）</h2>
<p dir="auto">主机<strong>无 BMC/IPMI、无机箱功耗计、无智能插座数据</strong>，因此整机瓦特数为<strong>基于实测分量的估算</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>状态</th>
<th>GPU</th>
<th>CPU 封装</th>
<th>主板+内存+2×NVMe+网卡+风扇</th>
<th>电源转换损耗(≈88%)</th>
<th><strong>整机估算（墙端）</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>空闲</strong>（服务在跑、无请求）</td>
<td>12 W</td>
<td>21 W</td>
<td>≈28 W</td>
<td>≈8 W</td>
<td><strong>≈70–90 W</strong></td>
</tr>
<tr>
<td><strong>推理满载</strong></td>
<td>299 W</td>
<td>47 W</td>
<td>≈35 W</td>
<td>≈42 W</td>
<td><strong>≈420–450 W</strong></td>
</tr>
<tr>
<td>短时峰值</td>
<td>354 W</td>
<td>69 W</td>
<td>≈35 W</td>
<td>≈50 W</td>
<td><strong>≈500 W</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>能耗折算（供参考）</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th>日耗电</th>
<th>年耗电</th>
</tr>
</thead>
<tbody>
<tr>
<td>24 h 空闲</td>
<td>≈1.9 kWh</td>
<td>≈700 kWh</td>
</tr>
<tr>
<td>24 h 满载</td>
<td>≈10.4 kWh</td>
<td>≈3,800 kWh</td>
</tr>
<tr>
<td>实际（假设每日 4 h 满载 + 20 h 空闲）</td>
<td>≈3.6 kWh</td>
<td>≈1,320 kWh</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">按 0.6 元/kWh 估算，实际混合负载约 <strong>2.2 元/天、约 790 元/年</strong>。</p>
</blockquote>
<hr />
<h2>5.4 散热评价</h2>
<p dir="auto"><strong>优点</strong></p>
<ul>
<li>GPU 空闲时风扇停在 891 rpm、12 W，功耗控制优秀。</li>
<li>满载时核心频率稳定在 2.7–3.1 GHz 且<strong>未触发降频</strong>，说明散热能力足以支撑 300 W 持续功耗。</li>
<li>CPU 与 NVMe 温度都很舒服（CPU 51 °C、NVMe ≤52 °C），机箱风道通畅。</li>
</ul>
<p dir="auto"><strong>需关注</strong></p>
<ul>
<li><strong>显存温度（92 °C 稳态 / 94 °C 峰值）高于 GPU 核心结温（89 °C 稳态）</strong>，这是典型的风冷 GDDR6 特征，也是本机<strong>最热的部件</strong>。持续高负载下会加速显存老化。</li>
<li>结温峰值 <strong>97 °C</strong> —— 虽然未触发降频（通常 AMD 结温降频阈值在 100 °C 以上），但已接近"舒适区"边界。</li>
<li>5 分钟压测尚不足以验证<strong>夏季高温环境</strong>或<strong>连续数小时满载</strong>下的表现；建议补做 1 小时以上满载测试。</li>
</ul>
<p dir="auto"><strong>散热提升建议（按性价比排序）</strong></p>
<ol>
<li>提高机箱进风量（前面板风扇），或改善 GPU 与机箱底板间距 —— 显存温度对风量最敏感。</li>
<li>使用 <code>rocm-smi --setperflevel</code> 或 <code>pp_power_profile_mode</code> 切换到 <strong>COMPUTE</strong> 档（当前为 <code>BOOTUP_DEFAULT</code>），可能获得更激进的风扇曲线（但会增加功耗/噪声）。</li>
<li>若可接受约 5–8% 的性能损失，把功耗墙从 300 W 下调到 250–270 W（<code>rocm-smi --setpoweroverdrive</code>），可显著降低显存温度。</li>
<li>夏季前重做一次散热器硅脂/导热垫维护。</li>
</ol>
<hr />
<h1>六、推理性能评测</h1>
<p dir="auto">所有测试均由局域网客户端 <code>192.168.67.203</code> 发起，目标 <code>192.168.67.202:8080</code>，属于<strong>真实端到端性能</strong>（含网络开销）。</p>
<hr />
<h2>6.1 单流生成速度（Decode）</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>生成长度</th>
<th>第 1 次</th>
<th>第 2 次</th>
<th>第 3 次</th>
<th>中位</th>
<th>标准差</th>
</tr>
</thead>
<tbody>
<tr>
<td>32 token</td>
<td>44.42</td>
<td>44.63</td>
<td>44.67</td>
<td><strong>44.63 t/s</strong></td>
<td>0.11</td>
</tr>
<tr>
<td>128 token</td>
<td>52.10</td>
<td>51.89</td>
<td>51.89</td>
<td><strong>51.89 t/s</strong></td>
<td>0.10</td>
</tr>
<tr>
<td>512 token</td>
<td>52.71</td>
<td>52.55</td>
<td>52.40</td>
<td><strong>52.55 t/s</strong></td>
<td>0.13</td>
</tr>
<tr>
<td>1024 token</td>
<td>52.76</td>
<td>52.74</td>
<td>52.73</td>
<td><strong>52.74 t/s</strong></td>
<td><strong>0.01</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><img src="https://upload.lcz.me/uploads/b4118d0b-45e0-4e62-a71d-d4c76fbf2649.jpeg" alt="a69251f5-3f40-4165-bc0e-dc03e41acef8-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><strong>结论</strong>：</p>
<ul>
<li>稳态生成速度 <strong>52.7 token/s</strong>，等价于 <strong>每 token 19.0 ms</strong>。</li>
<li>32 token 档偏低（44.6 t/s）是因为包含了首个 token 的计算与调度开销；长度 ≥128 token 后即进入稳态。</li>
<li><strong>稳定性极佳</strong>：1024 token 档三次测量标准差仅 <strong>0.01 t/s</strong>（0.02%）。</li>
<li>该成绩显著受益于 MTP 投机解码（日志显示 draft 接受率可达 <strong>0.99</strong>，<code>mean len = 2.99</code>）。</li>
</ul>
<blockquote>
<p dir="auto"><strong>注意长上下文衰减</strong>：服务自身日志显示，当上下文达到 3 万 token 时生成速度降至 <strong>30–43 t/s</strong>（例如 <code>tg_3s = 29.63 t/s</code>、<code>draft acceptance = 0.60</code>）。<strong>上下文越长，MTP 接受率越低、速度越慢</strong>，这是本地大模型推理的普遍规律。</p>
</blockquote>
<hr />
<h2>6.2 首 token 时延与流式输出</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>实测</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>TTFT 中位</strong>（前缀已缓存）</td>
<td><strong>145 ms</strong></td>
</tr>
<tr>
<td>TTFT 均值 / p90</td>
<td>199 ms / 304 ms（首次请求含冷启动）</td>
</tr>
<tr>
<td>客户端测流式吞吐</td>
<td><strong>52.15 t/s</strong>（5 次，标准差 0.05）</td>
</tr>
<tr>
<td>服务端计时吞吐</td>
<td>51.87 – 52.02 t/s（与客户端一致，说明网络无损耗）</td>
</tr>
<tr>
<td>平均 token 间隔（ITL）</td>
<td><strong>19.2 ms</strong></td>
</tr>
<tr>
<td>ITL 中位 / p90</td>
<td>0.02 ms / 55.5 ms（TCP 层聚合导致的分块到达）</td>
</tr>
<tr>
<td>256 token 总耗时</td>
<td>≈5.05 s</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong>说明</strong>：ITL 中位数 0.02 ms 而 p90 为 55 ms，是 llama.cpp 逐 token 推送被 TCP Nagle/缓冲合并的结果——客户端会以约 55 ms 的批次收到 token，但<strong>总吞吐完全一致（52.15 vs 51.9 t/s）</strong>。对于聊天类应用，体验等同于流式输出；对逐 token 实时性要求极高的场景（如语音），建议在服务端或客户端做 <code>TCP_NODELAY</code>/缓冲调优。</p>
</blockquote>
<hr />
<h2>6.3 长上下文预填充能力（关键短板）</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>输入 token 数</th>
<th>预填充速度</th>
<th><strong>实际耗时</strong></th>
<th>相对峰值</th>
</tr>
</thead>
<tbody>
<tr>
<td>256</td>
<td>669 t/s</td>
<td>0.38 s</td>
<td>70%</td>
</tr>
<tr>
<td>1,024</td>
<td>880 t/s</td>
<td>1.16 s</td>
<td>93%</td>
</tr>
<tr>
<td><strong>4,096</strong></td>
<td><strong>949 t/s（峰值）</strong></td>
<td>4.31 s</td>
<td>100%</td>
</tr>
<tr>
<td>16,384</td>
<td>809 t/s</td>
<td>20.3 s</td>
<td>85%</td>
</tr>
<tr>
<td>32,768</td>
<td>660 t/s</td>
<td>49.6 s</td>
<td>70%</td>
</tr>
<tr>
<td>65,536</td>
<td>480 t/s</td>
<td>136.6 s</td>
<td>51%</td>
</tr>
<tr>
<td><strong>98,304</strong></td>
<td><strong>376 t/s</strong></td>
<td><strong>261.2 s（4 分 21 秒）</strong></td>
<td>40%</td>
</tr>
</tbody>
</table>
<p dir="auto"><img src="https://upload.lcz.me/uploads/001428ce-962c-4b97-893b-2a386e33c8a2.jpeg" alt="30fcf742-a878-40aa-9be5-ae58653189fd-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><strong>结论</strong>：</p>
<ul>
<li>预填充速度在 <strong>4K token 时达到峰值 949 t/s</strong>，随后因注意力计算量随长度平方增长而持续下降。</li>
<li><strong>98K token 输入需要 4 分 21 秒</strong>；按实测速率外推，<strong>占满 128K 上下文约需 5.8 分钟</strong>。</li>
<li>这是"单 slot 串行"架构下最致命的问题：<strong>一个长文档分析请求会让整个服务停摆数分钟</strong>。</li>
<li>横向参考：<code>n_ctx=131072</code> 的 128K 上下文能力<strong>确实可用</strong>（无截断、无 OOM），但<strong>代价是分钟级的首 token 等待</strong>。</li>
</ul>
<hr />
<h2>6.4 前缀缓存（Prompt Cache）有效性</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th>处理 token</th>
<th>命中缓存</th>
<th>预填充速度</th>
<th><strong>耗时</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>3,400 token 首次（<code>cache_prompt=false</code>）</td>
<td>3,400</td>
<td>0</td>
<td>960 t/s</td>
<td><strong>3.55 s</strong></td>
</tr>
<tr>
<td>同一 prompt 再来一次（<code>cache_prompt=true</code>）</td>
<td>4</td>
<td>3,396</td>
<td>—</td>
<td><strong>0.24 s</strong></td>
</tr>
<tr>
<td>同一 prompt 第三次</td>
<td>4</td>
<td>3,396</td>
<td>—</td>
<td><strong>0.14 s</strong></td>
</tr>
<tr>
<td>多轮对话追加 20 token</td>
<td>340</td>
<td>3,400</td>
<td>600 t/s</td>
<td><strong>0.57 s</strong></td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong>缓存命中带来约 15 倍加速（3.55 s → 0.24 s）。</strong> 这意味着：</p>
<ol>
<li>客户端<strong>必须保证对话前缀逐字节不变</strong>，否则缓存全部失效；</li>
<li>多轮对话的追加式调用几乎无预填充开销；</li>
<li>但缓存容量有限——日志显示服务会淘汰旧条目（曾淘汰一条 <strong>6.2 GiB</strong> 的长上下文缓存），因此<strong>多客户端交替使用长上下文时缓存会互相冲刷</strong>。</li>
</ol>
</blockquote>
<hr />
<h2>6.5 并发能力（最重要的架构短板，双卡优势大）</h2>
<h3>并发压测：<code>-np 1</code> 下请求被严格串行化</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>并发客户端</th>
<th>总耗时</th>
<th>完成/失败</th>
<th>聚合吞吐</th>
<th>单请求速度</th>
<th>最后一个请求的等待时间</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>2.70 s</td>
<td>1 / 0</td>
<td><strong>47.5 t/s</strong></td>
<td>51.4 t/s</td>
<td>2.70 s</td>
</tr>
<tr>
<td>2</td>
<td>5.36 s</td>
<td>2 / 0</td>
<td><strong>47.8 t/s</strong></td>
<td>51.4 t/s</td>
<td>5.36 s</td>
</tr>
<tr>
<td>4</td>
<td>10.44 s</td>
<td>4 / 0</td>
<td><strong>49.1 t/s</strong></td>
<td>51.5 t/s</td>
<td>10.44 s</td>
</tr>
<tr>
<td>8</td>
<td>20.86 s</td>
<td>8 / 0</td>
<td><strong>49.1 t/s</strong></td>
<td>51.5 t/s</td>
<td><strong>20.86 s</strong></td>
</tr>
<tr>
<td>16</td>
<td>41.79 s</td>
<td>16 / 0</td>
<td><strong>49.0 t/s</strong></td>
<td>51.4 t/s</td>
<td><strong>41.79 s</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><img src="https://upload.lcz.me/uploads/e19477dd-c5f8-4512-9615-c48a973ab05c.jpeg" alt="e9031aa9-8d88-43d1-96a1-3225ba099076-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><strong>结论：并发 1 → 16，聚合吞吐几乎不变（47.5 → 49.0 t/s），总耗时与并发数严格成正比。</strong> 说明<strong>完全没有并行服务能力</strong>，每个请求排队串行执行。增加客户端只会等比拉长所有人的等待时间。</p>
<h3>公平性实测：小请求被长请求"饿死"</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th>请求内容</th>
<th>耗时</th>
</tr>
</thead>
<tbody>
<tr>
<td>单独执行</td>
<td>10 token 短请求</td>
<td><strong>0.40 s</strong></td>
</tr>
<tr>
<td>单独执行</td>
<td>16,200 token 长请求</td>
<td>19.99 s</td>
</tr>
<tr>
<td>并发（长请求先发 1 s）</td>
<td>10 token 短请求</td>
<td><strong>19.78 s</strong></td>
</tr>
<tr>
<td><strong>放大倍数</strong></td>
<td></td>
<td><strong>49 ×</strong></td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">一个正常情况下 0.4 秒完成的问候语请求，因为前面排队了一个 16K token 的预填充任务，实际等待了 <strong>19.8 秒</strong>。这就是 <code>-np 1</code> 的直接后果。</p>
</blockquote>
<h3>对"统一可用性"的影响（可用性长跑实测）</h3>
<p dir="auto">在评测窗口内以 2 秒间隔持续探测，共 365 次：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>探测类型</th>
<th>成功</th>
<th>失败</th>
<th>说明</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>GET /health</code></td>
<td><strong>166 / 166（100%）</strong></td>
<td>0</td>
<td>始终 &lt;8 ms 返回 200</td>
</tr>
<tr>
<td><code>GET /v1/models</code></td>
<td><strong>166 / 166（100%）</strong></td>
<td>0</td>
<td>始终快速返回</td>
</tr>
<tr>
<td><code>POST /completion</code>（4 token 生成）</td>
<td>30 / 33（91%）</td>
<td><strong>3 次超时（120 s）</strong></td>
<td>全部发生在长上下文预填充期间</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong><code>/health</code> 是"进程活着"探针，不是"服务可用"探针。</strong> 在 3 次小请求超时 120 秒的同时，<code>/health</code> 依然 100% 返回 200。<strong>任何基于 <code>/health</code> 做的负载均衡或健康检查都会失效。</strong></p>
</blockquote>
<hr />
<h2>6.6 多模态、工具调用与结构化输出</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>能力</th>
<th>实测结果</th>
<th>评价</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>视觉多模态</strong></td>
<td>一张 1000×700 JPEG 编码后触发 <strong>707 prompt token</strong>，预填充 628 t/s，6.64 s 内完成 200 token 描述</td>
<td><strong>功能可用，识别准确度差</strong>（见下）</td>
</tr>
<tr>
<td>图片识别准确度</td>
<td>把 AMD Radeon AI PRO R9700 识别为"华硕 TUF 系列 GeForce RTX 3060 Ti"，外观描述错误但格式规范</td>
<td><strong>差</strong>（模型能力问题，非部署问题）</td>
</tr>
<tr>
<td><strong>工具调用 Tool Calling</strong></td>
<td>正确返回 <code>get_weather({"city":"北京"})</code>，格式完全符合 OpenAI 规范</td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><strong>严格 JSON Schema</strong></td>
<td><code>response_format={"type":"json_schema",...}</code> 输出 <code>{"city":"北京","population":21893092}</code>，<strong>结构与类型（integer）均正确</strong></td>
<td><strong>优</strong></td>
</tr>
<tr>
<td><code>response_format={"type":"json_object"}</code></td>
<td><strong>未强制生效</strong>：模型输出被包裹在 <code>```json</code> 代码块中，<code>json.loads</code> 失败</td>
<td><strong>需注意</strong></td>
</tr>
<tr>
<td>思维链（reasoning）</td>
<td><code>reasoning_content</code> 正确分离返回（215 字符），最终答案正确（9.9 &gt; 9.11 并解释了版本号混淆）</td>
<td><strong>优</strong></td>
</tr>
<tr>
<td>关闭思维链</td>
<td><code>chat_template_kwargs={"enable_thinking": false}</code> 生效，22 token 直出答案，0.78 s</td>
<td><strong>优</strong></td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong>服务端日志同时给出的一条重要提示</strong>：<br />
<code>Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks</code> —— <strong>建议按官方建议增加 <code>--image-min-tokens 1024</code> 参数以改善图像定位任务准确度。</strong></p>
</blockquote>
<hr />
<h2>6.7 API 兼容性矩阵</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>端点</th>
<th>方法</th>
<th>状态</th>
<th>说明</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>/health</code></td>
<td>GET</td>
<td><strong>200</strong></td>
<td><code>{"status":"ok"}</code>，1 ms</td>
</tr>
<tr>
<td><code>/v1/models</code></td>
<td>GET</td>
<td><strong>200</strong></td>
<td>返回 <code>qwen3.8</code> 及完整 meta（n_params、n_ctx、ftype、size）</td>
</tr>
<tr>
<td><code>/props</code></td>
<td>GET</td>
<td><strong>200</strong></td>
<td>返回默认采样参数、<code>total_slots=1</code>、模态能力</td>
</tr>
<tr>
<td><code>/slots</code></td>
<td>GET</td>
<td><strong>200</strong></td>
<td>返回 slot 状态、剩余 token、speculative 标记</td>
</tr>
<tr>
<td><code>/v1/chat/completions</code></td>
<td>POST</td>
<td><strong>200</strong></td>
<td>OpenAI 兼容，支持流式、tools、response_format、reasoning</td>
</tr>
<tr>
<td><code>/v1/completions</code></td>
<td>POST</td>
<td><strong>200</strong></td>
<td>OpenAI 兼容补全</td>
</tr>
<tr>
<td><code>/completion</code></td>
<td>POST</td>
<td><strong>200</strong></td>
<td>原生端点，返回详细 timings（<strong>推荐用于压测</strong>）</td>
</tr>
<tr>
<td><code>/tokenize</code> / <code>/detokenize</code></td>
<td>POST</td>
<td><strong>200</strong></td>
<td>分词/反分词正常</td>
</tr>
<tr>
<td><code>/apply-template</code></td>
<td>POST</td>
<td><strong>200 / 400</strong></td>
<td>传入完整 <code>messages</code> 时 200 并返回渲染后 prompt；缺参时 400（正常）</td>
</tr>
<tr>
<td><code>/infill</code></td>
<td>POST</td>
<td><strong>500</strong></td>
<td>缺 <code>input_suffix</code> 时返回 <strong>500 而非 400</strong> —— API 规范瑕疵</td>
</tr>
<tr>
<td><code>/metrics</code></td>
<td>GET</td>
<td><strong>501</strong></td>
<td><code>Start it with '--metrics'</code> —— <strong>监控指标未开启</strong></td>
</tr>
<tr>
<td><code>/v1/embeddings</code></td>
<td>POST</td>
<td><strong>501</strong></td>
<td><code>Start it with '--embeddings'</code> —— <strong>无向量能力</strong></td>
</tr>
<tr>
<td><code>/ui</code> 与 <code>/</code></td>
<td>GET</td>
<td><strong>404</strong></td>
<td><code>/props</code> 声明 <code>"ui": true</code>，但该构建<strong>实际未提供 Web UI 端点</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>参数校验行为</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>测试</th>
<th>结果</th>
<th>评价</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>model</code> 字段填不存在的名字</td>
<td><strong>仍然正常返回结果</strong>（不校验模型名）</td>
<td>轻微不规范，多模型场景下可能造成误用</td>
</tr>
<tr>
<td><code>n_predict = -5</code></td>
<td><strong>400</strong> <code>Value must be between -1 &lt;= value &lt;= 2147483647</code></td>
<td>正确</td>
</tr>
<tr>
<td>缺少 <code>prompt</code> 字段</td>
<td><strong>400</strong> <code>key 'prompt' not found</code></td>
<td>正确</td>
</tr>
<tr>
<td><code>n_predict = 999999999</code></td>
<td><strong>接受</strong>，开始无上限生成</td>
<td><strong>风险</strong>：单请求可长时间独占唯一 slot</td>
</tr>
<tr>
<td>客户端中途断开</td>
<td><strong>服务端正确取消任务</strong>（日志 <code>stop: cancel task</code>）</td>
<td>优</td>
</tr>
</tbody>
</table>
<hr />
<h1>七、局域网连接与可用性评测（网络环境2.5G）</h1>
<hr />
<h2>7.1 网络时延、抖动与丢包</h2>
<p dir="auto">对两台目标各发送 <strong>600 个 ICMP 包（间隔 200 ms，总计 2 分钟）</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>目标</th>
<th>最小</th>
<th>中位(p50)</th>
<th>p90</th>
<th>p99</th>
<th>最大</th>
<th>均值</th>
<th>标准差</th>
<th>抖动(相邻差均值)</th>
<th>丢包</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>VM 192.168.67.202</strong></td>
<td>0.108 ms</td>
<td><strong>0.686 ms</strong></td>
<td>1.125 ms</td>
<td>1.320 ms</td>
<td>1.406 ms</td>
<td>0.741 ms</td>
<td>0.283 ms</td>
<td>0.236 ms</td>
<td><strong>0.0%</strong></td>
</tr>
<tr>
<td><strong>PVE 192.168.67.251</strong></td>
<td>0.152 ms</td>
<td><strong>0.374 ms</strong></td>
<td>0.493 ms</td>
<td>0.625 ms</td>
<td>0.963 ms</td>
<td>0.366 ms</td>
<td>0.109 ms</td>
<td>0.088 ms</td>
<td><strong>0.0%</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><img src="https://upload.lcz.me/uploads/ae832056-feb1-482c-bd78-5eaed2989709.jpeg" alt="0fe7c3a4-efbe-4d82-b2f9-6ac6eee14729-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><strong>分析</strong></p>
<ul>
<li>PVE 主机 p50 <strong>0.374 ms</strong>，是纯二层交换的理想值。</li>
<li>虚拟机 p50 <strong>0.686 ms</strong>，比主机多 <strong>0.31 ms</strong> —— 这就是 <strong>虚拟网卡（virtio → tap → fwbr → vmbr0）的固定开销</strong>，表现为一个约 0.3 ms 的双峰分布（38 个样本落在 &lt;0.25 ms，说明部分包走了更短路径）。</li>
<li><strong>零丢包、最大 1.4 ms</strong>，对 LLM 推理（单 token 19 ms）而言，<strong>网络时延完全不是瓶颈</strong>（占比 &lt;4%）。</li>
<li><code>vmbr0</code> 累计 TX dropped 仅 4 个包，可忽略。</li>
</ul>
<p dir="auto"><strong>HTTP API 层时延</strong>（100 次 <code>/health</code> 请求）</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>最小</th>
<th>中位</th>
<th>p90</th>
<th>最大</th>
<th>均值</th>
</tr>
</thead>
<tbody>
<tr>
<td>0.52 ms</td>
<td><strong>0.80 ms</strong></td>
<td>1.24 ms</td>
<td>7.38 ms</td>
<td>0.90 ms</td>
</tr>
</tbody>
</table>
<h2>7.2 局域网吞吐</h2>
<p dir="auto">使用自建 HTTP blob 服务端进行 <strong>1 GiB 定长传输</strong>，两个方向各 3 次：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>方向</th>
<th>第 1 次</th>
<th>第 2 次</th>
<th>第 3 次</th>
<th>等效带宽</th>
<th>达成率</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Mac → VM（VM 接收）</strong></td>
<td>292.99 MB/s</td>
<td>292.68 MB/s</td>
<td>293.06 MB/s</td>
<td><strong>2.344 Gbit/s</strong></td>
<td><strong>93.8%</strong>（2.5GbE 理论 2.5 Gbit/s）</td>
</tr>
<tr>
<td><strong>VM → Mac（VM 发送）</strong></td>
<td>293.28 MB/s</td>
<td>293.70 MB/s</td>
<td>293.01 MB/s</td>
<td><strong>2.344 Gbit/s</strong></td>
<td><strong>93.8%</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>结论：网络链路完全不是瓶颈。</strong> 虚拟机 virtio 网卡（单队列、TSO/GSO/GRO 全开）可以打满 2.5GbE 线速，且<strong>收发对称</strong>。对推理服务而言，即使每 token 都带 100 字节负载，52 t/s 也仅需 5 KB/s —— 带宽余量有 <strong>4 个数量级</strong>。</p>
<p dir="auto"><strong>MTU 探测</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>目标</th>
<th>payload 1472（MTU 1500）</th>
<th>payload 4000</th>
<th>payload 8972（巨帧 MTU 9000）</th>
</tr>
</thead>
<tbody>
<tr>
<td>VM</td>
<td>通过（不分片）</td>
<td>失败</td>
<td>失败</td>
</tr>
<tr>
<td>PVE</td>
<td>通过</td>
<td>—</td>
<td>失败</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">全链路 <strong>MTU = 1500，未启用巨帧</strong>。对 LLM 推理无影响，无需调整。</p>
</blockquote>
<hr />
<h2>7.3 可用性长跑与故障行为</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>场景</th>
<th>行为</th>
<th>评价</th>
</tr>
</thead>
<tbody>
<tr>
<td>服务空闲</td>
<td><code>/health</code> 0.8 ms，<code>/slots</code> 正常</td>
<td>优</td>
</tr>
<tr>
<td>单流满载 5 分钟</td>
<td>78 请求 0 失败，<code>/health</code> 始终可用</td>
<td>优</td>
</tr>
<tr>
<td>长上下文预填充期间</td>
<td><strong>小请求排队超时（120 s 内未返回）</strong>；<code>/health</code> 仍 200</td>
<td><strong>差</strong></td>
</tr>
<tr>
<td>客户端中途断开</td>
<td>服务端取消任务，资源正确释放</td>
<td>优</td>
</tr>
<tr>
<td>超大 <code>n_predict</code> 请求</td>
<td>正常接受并持续生成（<strong>可被滥用为拒绝服务</strong>）</td>
<td>风险</td>
</tr>
<tr>
<td>服务重启后恢复</td>
<td>模型加载 <strong>8.2 秒</strong> 完成并开始监听</td>
<td>优</td>
</tr>
<tr>
<td>主机重启后自恢复</td>
<td><code>onboot: 1</code> + <code>enabled</code>，自动启动，<code>NRestarts=0</code></td>
<td>优</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>实测可用性指标（评测窗口）</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>/health</code> 可用率</td>
<td><strong>100.0%</strong>（166/166）</td>
</tr>
<tr>
<td><code>/v1/models</code> 可用率</td>
<td><strong>100.0%</strong>（166/166）</td>
</tr>
<tr>
<td>真实推理请求成功率</td>
<td><strong>90.9%</strong>（30/33，3 次超时全部由槽位排队引起）</td>
</tr>
<tr>
<td>HTTP 错误率（4xx/5xx 主动返回）</td>
<td>0%</td>
</tr>
</tbody>
</table>
<hr />
<h1>八、补图</h1>
<hr />
<h2>8.1 主机图片（闷罐）</h2>
<p dir="auto"><img src="https://upload.lcz.me/uploads/311251cf-7ba6-42e7-bf97-50759ce1732b.jpeg" alt="de1bd3a6-9824-4c7a-aacf-506083e5300e-image.jpeg" class=" img-fluid img-markdown" /></p>
<hr />
<p dir="auto"><img src="https://upload.lcz.me/uploads/4a06ce08-1297-4edf-a433-d0be00e75456.jpeg" alt="cc555ede-4c42-4ccc-ae23-2a880b965872-image.jpeg" class=" img-fluid img-markdown" /></p>
<hr />
<h2>8.2 待机功率</h2>
<hr />
<p dir="auto"><img src="https://upload.lcz.me/uploads/cf0e0ef9-9505-4418-8d2d-58a0ab7ef078.jpeg" alt="c684cbb7-d3f8-4fff-aa3d-9d0c858cfc6e-image.jpeg" class=" img-fluid img-markdown" /></p>
<h2>8.3 满载功率</h2>
<hr />
<p dir="auto"><img src="https://upload.lcz.me/uploads/57c9b7ec-6781-4482-bf84-fed6788c69f8.jpeg" alt="fd31cf3f-f3fa-45e7-a443-95febb23fe8a-image.jpeg" class=" img-fluid img-markdown" /></p>
<hr />
<h1>九、实测结束</h1>
<hr />
<p dir="auto"><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f680.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--rocket" style="height:23px;width:auto;vertical-align:middle" title="🚀" alt="🚀" /> ALL IN ONE · ALL IN BOOM <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f4a5.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--boom" style="height:23px;width:auto;vertical-align:middle" title="💥" alt="💥" /><br />
<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a1.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--zap" style="height:23px;width:auto;vertical-align:middle" title="⚡" alt="⚡" /> 一台机器，全部搞定 —— 然后，炸裂全场！<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a1.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--zap" style="height:23px;width:auto;vertical-align:middle" title="⚡" alt="⚡" /></p>
<hr />
<p dir="auto"><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f3af.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--dart" style="height:23px;width:auto;vertical-align:middle" title="🎯" alt="🎯" /> ALL IN ONE —— 一机全能</p>
<p dir="auto"><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f5a5.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--desktop_computer" style="height:23px;width:auto;vertical-align:middle" title="🖥" alt="🖥" />️ 虚拟化 (PVE)	                          <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /><br />
<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f916.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--robot_face" style="height:23px;width:auto;vertical-align:middle" title="🤖" alt="🤖" /> AI 推理 (llama.cpp)	                  <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /><br />
<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f3ae.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--video_game" style="height:23px;width:auto;vertical-align:middle" title="🎮" alt="🎮" /> GPU 直通 (R9700)	                  <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /><br />
<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f310.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--globe_with_meridians" style="height:23px;width:auto;vertical-align:middle" title="🌐" alt="🌐" /> 网络服务	                                  <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /><br />
<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f4be.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--floppy_disk" style="height:23px;width:auto;vertical-align:middle" title="💾" alt="💾" /> 存储中心	                                  <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /></p>
<p dir="auto">一台机器 = 服务器 + AI 工作站 + 存储 + 网关 All in One，All the Power. <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f525.png?v=efcae6a46b1" class="not-responsive emoji emoji-android emoji--fire" style="height:23px;width:auto;vertical-align:middle" title="🔥" alt="🔥" /></p>
]]></description><link>https://lcz.me/topic/1655</link><generator>RSS for Node</generator><lastBuildDate>Wed, 23 Sep 2026 03:26:47 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1655.rss" rel="self" type="application/rss+xml"/><pubDate>Sat, 12 Sep 2026 13:06:37 GMT</pubDate><ttl>60</ttl></channel></rss>