<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[华南金牌X99-TF 2696V3 PVE直通7900XTX成功]]></title><description><![CDATA[<p dir="auto">首先感谢一下版主推荐了7900xtx，建议也很好都是个渐进的过程，小投入先玩玩看，有必要再砸钱。</p>
<p dir="auto">目标配置如下，显卡目前是7900xtx。<br />
序号	配件名称	型号	规格	备注	数量	价格<br />
1	cpu	2696V3	LGA2011-3	鸡血刷bios降压	1	0<br />
2	主板	华南金牌TF	ATX	板U套装	1	958<br />
3	内存	DDR3服务器ecc	1600MHz		4	980<br />
4	机箱	旧机箱	nzxt p410	只有240水冷		0<br />
5	硬盘	1Tssd 	sata接口			0<br />
6	cpu水冷		240规格		1	295<br />
7	显卡	7900xtx蓝宝石超白金		咸鱼		6500<br />
8	电源	酷冷至尊（CoolerMaster）X Mighty 2000W 白金牌静音电源	2000w模组		1	2799<br />
<img src="https://upload.lcz.me/uploads/a8b2c47c-f5d6-4103-af0d-4c9733c3821f.jpeg" alt="e9a943af-3d77-4b15-912c-a843563ce403-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">分别在三块硬盘上安装了三个系统<br />
第一块1个256g的nvme安装了ubuntu24.04LTS, Qwen2.5的7B和Qwen3.8的27B，没有做深度测试仅仅是跑通。十几个token上下，没有做优化，显卡用的节能模式。<br />
第二块是个sata的120Gssd硬盘，安装了Win10<br />
第三块是sata的960G的ssd，安装了PVE。一个server2022虚拟机，一个ubuntu24.04加7900xtx直通。</p>
<p dir="auto">听了版主很多视频，一再强调x99够用，最好物理安装，虚拟机很麻烦会有各种问题，但是需求也很客观，我确实需要两个机器跑数据运算，切成两个虚拟机并直通一个是最合理的分配。一个跑数据库一个跑前端运算，核心和线程可以完全匹配全部利用起来，而且比网络传输数据可靠速度也更靠谱。</p>
<p dir="auto">PVE里面server2022虚拟机和ubuntu24.04各分配了60G内存8个核心，留给pve2个核心8G内存。折腾了一天，还没安装本地大模型，但是pve的显卡直通已经搞定了。先说重点后续补充其他信息。</p>
<p dir="auto">7900xtx 是可以pve直通的。我有另外一张geforce620的亮机卡，所以可能少走了很多弯路。<br />
一、x99 BIOS准备 ，用的官方鸡血bios，版本4.04 （注意bios版本，我并没有出现后关闭导致7900xtxCSM黑屏可能与bios版本有关）<br />
二、IOMMU和VT-D开启<br />
三、CSM关闭 关闭前确认显卡选项是EFI<br />
四、开启Above 4G<br />
五、Resize Bar 开启<br />
六、SR-IOV开启<br />
<strong>重点</strong>虚拟机先不要加上7900xtx，先把虚拟机安装好，之后用命令行去直通显卡。<br />
七、之后pve会面对多次开机锁死，解决方法是在pve开始菜单按e编辑grub，之后ctrl+x启动。  后续我会补充grub代码。PVE。一个server2022虚拟机，一个ubuntu24.04加7900xtx直通。<br />
八、进入pve之后会发现虚拟机无法启动，主要问题是x99太老，导致pcie的分配混乱导致冲突引起。需要写脚本固定。后续补充代码。<br />
九、<strong>重点</strong>这是重点中的重点，x99必须要作的，提取显卡的bios并存到pve的/usr/share/kvm目录，使用gpu-z提取</p>
<p dir="auto">十、这个年代肯定都是在AI的指导下复制粘贴，总共动用了三个，Deepseek网页版，gemini网页版，Qwen网页版。最靠谱的是Deepseek在他的指导下完成。不提vbios的都是刷流氓。Qwen幻觉有点严重。</p>
<p dir="auto">十一、很晚了先占个坑，回头测一下直通的大模型情况再来更新。初步验证了直通的可行性，实际大模型表现待确认</p>
]]></description><link>https://lcz.me/topic/1781</link><generator>RSS for Node</generator><lastBuildDate>Fri, 25 Sep 2026 21:28:31 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1781.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 17 Sep 2026 16:30:39 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Tue, 22 Sep 2026 10:17:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a><br />
其实直通之后没有太多性能损失的前提下，这玩意具有非常大的实用价值：</p>
<p dir="auto">一、目前我们大部分人的算例底座都是性能过剩的，即使是X99这种老东西切了一半的cpu资源也没有任何影响，切几个核心几个G的内存百十来G的硬盘几乎没啥感觉，所以在虚拟机和直通的基础上，我们可以分裂出若干个虚拟机，用来跑agent，即避免了资源浪费噪音污染，也减少了资金投入，不然还要再找一台电脑去架设个agent。</p>
<p dir="auto">二、对提高系统的稳定性和可用性帮助也很大，pve虚拟机可以提供非常方便的冷热备份，只要有个大硬盘或者nas，对于喜欢折腾的人来说，只要及时备份虚拟机，一个系统可以在几分钟之内迅速的完全恢复。</p>
<p dir="auto">三、虚拟机是可以迁移的，如同备份。</p>
]]></description><link>https://lcz.me/post/20058</link><guid isPermaLink="true">https://lcz.me/post/20058</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Tue, 22 Sep 2026 10:17:58 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Mon, 21 Sep 2026 04:26:26 GMT]]></title><description><![CDATA[<p dir="auto">不错。直通独立显卡。在一些人的知识面还是盲区。<br />
感谢为他们科普了。</p>
]]></description><link>https://lcz.me/post/19665</link><guid isPermaLink="true">https://lcz.me/post/19665</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Mon, 21 Sep 2026 04:26:26 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Sun, 20 Sep 2026 23:38:50 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/0e2086f8-60a7-4755-ac9d-f3e1580255cb.jpg" alt="Screenshot_20260921_073458_com.tencent.mm.jpg" class=" img-fluid img-markdown" /></p>
<p dir="auto">更新一下最新的使用速度情况,用的是Hermes agent</p>
]]></description><link>https://lcz.me/post/19622</link><guid isPermaLink="true">https://lcz.me/post/19622</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Sun, 20 Sep 2026 23:38:50 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Fri, 18 Sep 2026 18:17:03 GMT]]></title><description><![CDATA[<h1>单张 RX 7900 XTX + PVE 直通：Qwen3.8-27B 本地大模型部署实录</h1>
<h2>一、最终成果</h2>
<p dir="auto">在 PVE 虚拟机中直通单张 RX 7900 XTX，跑 Qwen3.8-27B Q4_K_M：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>unsloth UD 版</th>
<th>去审查版</th>
</tr>
</thead>
<tbody>
<tr>
<td>Decode（98K）</td>
<td>41.25 t/s</td>
<td><strong>47.95 t/s</strong></td>
</tr>
<tr>
<td>Prefill</td>
<td>199.38 t/s</td>
<td>177.27 t/s</td>
</tr>
<tr>
<td>接受率</td>
<td>0.345</td>
<td>0.420</td>
</tr>
<tr>
<td>mean len</td>
<td>2.73</td>
<td>3.10</td>
</tr>
<tr>
<td>上下文</td>
<td>98304（98K）</td>
<td>98304（98K）</td>
</tr>
<tr>
<td>KV Cache</td>
<td>K q8_0 / V q4_1</td>
<td>K q8_0 / V q4_1</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>去审查版的 MTP 头是完整的第 65 层（含 attention + FFN），比 unsloth UD 版（只有 4 个 <code>nextn.*</code> 投影张量）猜得更准，decode 快约 16%。</strong></p>
<hr />
<h2>二、硬件与软件环境</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>配置</th>
</tr>
</thead>
<tbody>
<tr>
<td>主板</td>
<td>华南金牌 X99-TF</td>
</tr>
<tr>
<td>CPU</td>
<td>Xeon E5-2696V3</td>
</tr>
<tr>
<td>内存</td>
<td>128GB DDR3 1600 ECC</td>
</tr>
<tr>
<td>显卡</td>
<td>AMD Radeon RX 7900 XTX 24GB（Navi 31 / gfx1100）</td>
</tr>
<tr>
<td>虚拟化</td>
<td>PVE，Ubuntu 24.04 虚拟机，GPU 直通</td>
</tr>
<tr>
<td>后端</td>
<td>Vulkan（Mesa 25.2.8 RADV）</td>
</tr>
<tr>
<td>推理引擎</td>
<td>llama.cpp build 11037（commit 44be98f05）</td>
</tr>
<tr>
<td>模型</td>
<td>unsloth Qwen3.8-27B-UD-Q4_K_M（16.5G）/ 去审查版 Qwen3.8-27B-Uncensored-Q4_K_M（16.8G）</td>
</tr>
</tbody>
</table>
<hr />
<h2>三、编译 llama.cpp</h2>
<pre><code class="language-bash">sudo apt update
sudo apt install -y git build-essential cmake libvulkan-dev glslc spirv-headers

git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON
cmake --build build --config Release -j$(nproc)
</code></pre>
<p dir="auto"><strong>注意</strong>：<code>spirv-headers</code> 是较新版本新增的强制依赖，旧教程不会提到。缺它会在 CMake 配置阶段报 <code>Could not find SPIRV-Headers</code>。</p>
<p dir="auto">验证 Vulkan 设备：</p>
<pre><code class="language-bash">./build/bin/llama-cli --list-devices
# 应输出：Vulkan0: AMD Radeon RX 7900 XTX (RADV NAVI31) (24576 MiB, 23624 MiB free)
</code></pre>
<hr />
<h2>四、权限配置（PVE 直通必做）</h2>
<p dir="auto"><code>/dev/dri/renderD128</code> 属于 <code>root:render</code>，普通用户默认无权限访问。这会导致 Vulkan 报 <code>Permission denied (VK_ERROR_INCOMPATIBLE_DRIVER)</code>。</p>
<pre><code class="language-bash">sudo usermod -aG render,video $USER
</code></pre>
<p dir="auto"><strong>关键</strong>：加组后必须<strong>完全退出 SSH 会话重新登录</strong>，组权限才生效。验证：</p>
<pre><code class="language-bash">groups
# 应包含 render 和 video
vulkaninfo --summary | grep -A5 "GPU0"
# 应显示 AMD Radeon RX 7900 XTX (RADV NAVI31)
</code></pre>
<hr />
<h2>五、下载模型</h2>
<pre><code class="language-bash">pipx install "huggingface_hub[cli]"
# 新版命令是 hf，不是 huggingface-cli

# unsloth UD 版
hf download unsloth/Qwen3.8-27B-GGUF \
  --include "*Q4_K_M*" \
  --local-dir ~/llama.cpp

# 去审查版（推荐）
# 从社区仓库下载 Qwen3.8-27B-Uncensored-Q4_K_M.gguf + mmproj
</code></pre>
<p dir="auto">下载不稳定时可用 <code>hfd.sh</code>（基于 aria2c，支持断点续传）。</p>
<hr />
<h2>六、核心调参过程与实测数据</h2>
<h3>6.1 初始配置（失败）</h3>
<pre><code class="language-bash">--ctx-size 65536 --cache-type-k q8_0 --cache-type-v q8_0 \
--spec-type draft-mtp,ngram-map-k4v --spec-draft-n-max 3
</code></pre>
<p dir="auto">结果：<strong>decode 13.12 t/s</strong>，接受率 0.490。</p>
<p dir="auto">问题：<code>--spec-ngram-map-k4v</code> 在 Vulkan 后端产生巨大开销，反而拖慢。</p>
<h3>6.2 去掉 n-gram（改善但仍有断崖）</h3>
<pre><code class="language-bash">--ctx-size 8192 --cache-type-k q8_0 --cache-type-v q8_0 \
--spec-type draft-mtp --spec-draft-n-max 3
</code></pre>
<p dir="auto">结果：<strong>decode 51.67 t/s</strong>，接受率 0.449。</p>
<p dir="auto">但 ctx 提到 16384/32768/65536/98304 后，全部掉到 <strong>9–14 t/s</strong>。</p>
<h3>6.3 定位断崖</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>ctx</th>
<th>KV</th>
<th>decode</th>
</tr>
</thead>
<tbody>
<tr>
<td>8192</td>
<td>q8_0/q8_0</td>
<td>51.67</td>
</tr>
<tr>
<td>16384</td>
<td>q8_0/q8_0</td>
<td>14.18</td>
</tr>
<tr>
<td>32768</td>
<td>q8_0/q8_0</td>
<td>14.47</td>
</tr>
<tr>
<td>65536</td>
<td>q8_0/q8_0</td>
<td>13.12</td>
</tr>
<tr>
<td>98304</td>
<td>q8_0/q8_0</td>
<td>11–15</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>断崖在 8192 和 16384 之间。</strong></p>
<h3>6.4 第一次修复尝试：改 KV 量化</h3>
<pre><code class="language-bash">--ctx-size 16384 --cache-type-k q8_0 --cache-type-v q4_1 \
--ctx-checkpoints 2 --spec-draft-n-max 5
</code></pre>
<p dir="auto">结果：<strong>decode 46.08 t/s</strong>，接受率 0.439。</p>
<p dir="auto">V 从 q8_0 降到 q4_1 后，16384 的断崖消失。但 ctx 提到 65536/98304 后仍掉回 10–13 t/s。</p>
<h3>6.5 根因定位：256M BAR</h3>
<pre><code class="language-bash">lspci -v -s 01:00.0 | grep -i "size="
# Memory at 380000000000 (64-bit, prefetchable) [size=256M]
</code></pre>
<p dir="auto"><strong>PVE 直通环境下 BAR 只有 256M，ReBAR 未生效。</strong></p>
<p dir="auto">llama.cpp 的 Vulkan 后端默认尝试使用这 256MB 的主机可见窗口。ctx 增大后，KV cache 和计算 buffer 装不下，被迫走碎片化的低效路径，decode 断崖。</p>
<h3>6.6 最终修复：禁用 host visible VRAM</h3>
<pre><code class="language-bash">export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1
</code></pre>
<p dir="auto">让 Vulkan 后端完全放弃 256MB 窗口，所有 buffer 直接分配到设备本地 VRAM。</p>
<p dir="auto">结果：<strong>98304 ctx 下 decode 41.25 t/s，prefill 199.38 t/s</strong>，断崖彻底解决。</p>
<h3>6.7 换用去审查版（进一步提升）</h3>
<p dir="auto">去审查版的 <code>blk.64</code> 是一个<strong>完整的 Transformer 层</strong>（含 attn_q/k/v/output + ffn_gate/down/up + nextn 投影头），而 unsloth UD 版只有 4 个 <code>nextn.*</code> 张量。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>去审查版</th>
<th>unsloth UD 版</th>
</tr>
</thead>
<tbody>
<tr>
<td>decode（98K）</td>
<td><strong>47.95 t/s</strong></td>
<td>41.25 t/s</td>
</tr>
<tr>
<td>接受率</td>
<td>0.420</td>
<td>0.345</td>
</tr>
<tr>
<td>mean len</td>
<td>3.10</td>
<td>2.73</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>MTP 头的完整程度直接影响投机解码的收益</strong>，这是论坛里很少有人提到的一点。</p>
<hr />
<h2>七、最终可用配置</h2>
<h3>脚本一：unsloth UD 版（+ mmproj）</h3>
<pre><code class="language-bash">#!/bin/bash
# ~/start-llama-unsloth.sh

export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1

cd /home/kby/llama.cpp
exec ./build/bin/llama-server \
  -m /home/kby/llama.cpp/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --mmproj /home/kby/llama.cpp/Qwen3.8/mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
  --host 0.0.0.0 --port 8080 \
  --device Vulkan0 \
  -ngl 99 \
  --parallel 1 \
  --ctx-size 98304 \
  --flash-attn on \
  --cache-type-k q8_0 --cache-type-v q4_1 \
  --ctx-checkpoints 2 \
  --spec-type draft-mtp \
  --spec-draft-n-max 5 \
  --cache-ram 32768 \
  --jinja
</code></pre>
<h3>脚本二：去审查版（+ mmproj，推荐）</h3>
<pre><code class="language-bash">#!/bin/bash
# ~/start-llama-uncensored.sh

export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1

cd /home/kby/llama.cpp
exec ./build/bin/llama-server \
  -m /home/kby/llama.cpp/Qwen3.8/Qwen3.8-27B-Uncensored-Q4_K_M.gguf \
  --mmproj /home/kby/llama.cpp/Qwen3.8/mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
  --host 0.0.0.0 --port 8080 \
  --device Vulkan0 \
  -ngl 99 \
  --parallel 1 \
  --ctx-size 98304 \
  --flash-attn on \
  --cache-type-k q8_0 --cache-type-v q4_1 \
  --ctx-checkpoints 2 \
  --spec-type draft-mtp \
  --spec-draft-n-max 5 \
  --cache-ram 32768 \
  --jinja
</code></pre>
<h3>创建命令</h3>
<pre><code class="language-bash"># unsloth UD 版
cat &gt; ~/start-llama-unsloth.sh &lt;&lt; 'EOF'
#!/bin/bash
export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1
cd /home/kby/llama.cpp
exec ./build/bin/llama-server \
  -m /home/kby/llama.cpp/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --mmproj /home/kby/llama.cpp/Qwen3.8/mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
  --host 0.0.0.0 --port 8080 \
  --device Vulkan0 \
  -ngl 99 \
  --parallel 1 \
  --ctx-size 98304 \
  --flash-attn on \
  --cache-type-k q8_0 --cache-type-v q4_1 \
  --ctx-checkpoints 2 \
  --spec-type draft-mtp \
  --spec-draft-n-max 5 \
  --cache-ram 32768 \
  --jinja
EOF

# 去审查版
cat &gt; ~/start-llama-uncensored.sh &lt;&lt; 'EOF'
#!/bin/bash
export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1
cd /home/kby/llama.cpp
exec ./build/bin/llama-server \
  -m /home/kby/llama.cpp/Qwen3.8/Qwen3.8-27B-Uncensored-Q4_K_M.gguf \
  --mmproj /home/kby/llama.cpp/Qwen3.8/mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
  --host 0.0.0.0 --port 8080 \
  --device Vulkan0 \
  -ngl 99 \
  --parallel 1 \
  --ctx-size 98304 \
  --flash-attn on \
  --cache-type-k q8_0 --cache-type-v q4_1 \
  --ctx-checkpoints 2 \
  --spec-type draft-mtp \
  --spec-draft-n-max 5 \
  --cache-ram 32768 \
  --jinja
EOF

chmod +x ~/start-llama-unsloth.sh ~/start-llama-uncensored.sh
</code></pre>
<h3>使用</h3>
<pre><code class="language-bash">~/start-llama-uncensored.sh    # 推荐
# 或
~/start-llama-unsloth.sh
</code></pre>
<p dir="auto">后台运行：</p>
<pre><code class="language-bash">nohup ~/start-llama-uncensored.sh &gt; ~/llama-server.log 2&gt;&amp;1 &amp;
</code></pre>
<hr />
<h2>八、参数说明</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>参数</th>
<th>作用</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1</code></td>
<td><strong>最关键</strong>。PVE 直通下 BAR 只有 256M，此变量强制所有 buffer 走设备本地 VRAM，绕过断崖</td>
</tr>
<tr>
<td><code>--mmproj</code></td>
<td>加载视觉投影，让模型能看图片。不需要视觉功能时去掉，省约 900MB 显存</td>
</tr>
<tr>
<td><code>--device Vulkan0</code></td>
<td>多设备时明确指定显卡</td>
</tr>
<tr>
<td><code>-ngl 99</code></td>
<td>全部层上 GPU</td>
</tr>
<tr>
<td><code>--parallel 1</code></td>
<td>单槽独占 98K 上下文</td>
</tr>
<tr>
<td><code>--ctx-size 98304</code></td>
<td>98K 上下文</td>
</tr>
<tr>
<td><code>--cache-type-k q8_0 --cache-type-v q4_1</code></td>
<td>K 精度优先，V 用 q4_1（带 scale+min，比 q4_0 稳）</td>
</tr>
<tr>
<td><code>--ctx-checkpoints 2</code></td>
<td>稳定性四件套之一，不影响速度</td>
</tr>
<tr>
<td><code>--spec-type draft-mtp</code></td>
<td>启用模型内建 MTP 投机解码。<strong>不要加 ngram-map-k4v</strong>，Vulkan 下反而拖慢</td>
</tr>
<tr>
<td><code>--spec-draft-n-max 5</code></td>
<td>一次猜 5 个 token</td>
</tr>
<tr>
<td><code>--cache-ram 32768</code></td>
<td>prompt cache 开到 32GB，利用 128GB 内存，多轮对话避免重复 prefill</td>
</tr>
<tr>
<td><code>--jinja</code></td>
<td>启用 chat template</td>
</tr>
</tbody>
</table>
<hr />
<h2>九、性能对比总表</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>配置</th>
<th>ctx</th>
<th>KV</th>
<th>n_max</th>
<th>decode</th>
<th>接受率</th>
<th>判定</th>
</tr>
</thead>
<tbody>
<tr>
<td>MTP+ngram</td>
<td>65536</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>13.12</td>
<td>0.490</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> ngram 拖慢</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>8192</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>51.67</td>
<td>0.449</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /></td>
</tr>
<tr>
<td>纯 MTP</td>
<td>16384</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>14.18</td>
<td>0.551</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 断崖</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>32768</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>14.47</td>
<td>0.570</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 断崖</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>98304</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>11–15</td>
<td>—</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 断崖</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>16384</td>
<td><strong>q8_0/q4_1</strong></td>
<td>5</td>
<td><strong>46.08</strong></td>
<td>0.439</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 断崖消失</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>98304</td>
<td>q8_0/q4_1</td>
<td>5</td>
<td>11.04</td>
<td>0.377</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 仍断崖</td>
</tr>
<tr>
<td><strong>最终（unsloth）</strong></td>
<td><strong>98304</strong></td>
<td><strong>q8_0/q4_1</strong></td>
<td><strong>5</strong></td>
<td><strong>41.25</strong></td>
<td>0.345</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 禁用 host VRAM</td>
</tr>
<tr>
<td><strong>最终（去审查）</strong></td>
<td><strong>98304</strong></td>
<td><strong>q8_0/q4_1</strong></td>
<td><strong>5</strong></td>
<td><strong>47.95</strong></td>
<td>0.420</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> <strong>推荐</strong></td>
</tr>
</tbody>
</table>
<hr />
<h2>十、踩坑清单</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>#</th>
<th>坑</th>
<th>现象</th>
<th>解法</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>缺 <code>spirv-headers</code></td>
<td>CMake 报 <code>Could not find SPIRV-Headers</code></td>
<td><code>sudo apt install spirv-headers</code></td>
</tr>
<tr>
<td>2</td>
<td>用户不在 <code>render</code> 组</td>
<td>Vulkan 报 <code>Permission denied (VK_ERROR_INCOMPATIBLE_DRIVER)</code></td>
<td><code>sudo usermod -aG render,video $USER</code>，重登录</td>
</tr>
<tr>
<td>3</td>
<td>桌面会话占用 GPU</td>
<td>llama-bench 卡死、GPU 利用率 0%</td>
<td>停掉 gdm3，或让桌面用另一张卡</td>
</tr>
<tr>
<td>4</td>
<td><code>ngram-map-k4v</code> 在 Vulkan 下拖慢</td>
<td>decode 从 51 掉到 13</td>
<td>去掉该参数，只用 <code>draft-mtp</code></td>
</tr>
<tr>
<td>5</td>
<td>256M BAR 导致 ctx 断崖</td>
<td>16K 以上 decode 掉到 14</td>
<td><code>export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1</code></td>
</tr>
<tr>
<td>6</td>
<td>短 prompt 测 prefill</td>
<td>59 token 测出 20 t/s 假数字</td>
<td>用 ≥2000 token 的 prompt 测</td>
</tr>
<tr>
<td>7</td>
<td>同名 GGUF 可能没有 MTP 层</td>
<td>报 <code>model doesn't contain MTP layers</code></td>
<td>换 unsloth UD 版本，确认 <code>blk.64.nextn.*</code> 存在</td>
</tr>
<tr>
<td>8</td>
<td>mmproj 与模型不匹配</td>
<td>报张量不匹配错误</td>
<td>unsloth 版用 unsloth 的 mmproj，去审查版用去审查版的</td>
</tr>
</tbody>
</table>
<hr />
<h2>十一、关键经验</h2>
<ol>
<li><strong>PVE 直通下 ReBAR 通常不生效</strong>，BAR 只有 256M。<code>GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1</code> 是必加项，不是可选项。</li>
<li><strong>Vulkan 后端的 <code>ngram-map-k4v</code> 在 RDNA3 上表现异常</strong>，社区在 ROCm/HIP 下的收益无法复现，建议只用 <code>draft-mtp</code>。</li>
<li><strong>KV 量化不对称是正解</strong>：K q8_0 / V q4_1。K 决定注意力方向，敏感；V 是加权平均，容忍度高。</li>
<li><strong>断崖是确定性的，不是脏状态</strong>。冷启动、子分配修复都无效，只有禁用 host visible VRAM 才解决。</li>
<li><strong>测速必须固定内容类型</strong>。代码生成接受率高（0.4–0.6），散文接受率低（0.3），同一配置速度可差 2–3 倍。</li>
<li><strong>MTP 头的完整程度直接影响投机解码收益</strong>。去审查版的 <code>blk.64</code> 是完整 Transformer 层，接受率 0.420；unsloth UD 版只有 4 个投影张量，接受率 0.345。</li>
<li><strong>mmproj 会额外占约 900MB 显存</strong>。加载后如果 98K 上下文触发 OOM，降到 64K。</li>
</ol>
<hr />
<h2>十二、附录：验证命令</h2>
<pre><code class="language-bash"># 确认 Vulkan 设备可见
vulkaninfo --summary | grep -A10 "GPU0"

# 确认 BAR 大小
lspci -v -s 01:00.0 | grep -i "size="

# 确认用户组
groups

# 确认 llama.cpp 链接了 Vulkan
ldd build/bin/llama-server | grep -i vulkan

# 确认 GGUF 有 MTP 层
./build/bin/llama-gguf Qwen3.8/Qwen3.8-27B-Uncensored-Q4_K_M.gguf r 2&gt;&amp;1 | grep -iE "blk\.64|nextn" | head -20
</code></pre>
<hr />
<p dir="auto"><strong>测试日期</strong>：2026-09-18 ～ 2026-09-19<br />
<strong>环境</strong>：PVE + Ubuntu 24.04 + RX 7900 XTX 直通 + llama.cpp build 11037 + Mesa 25.2.8<br />
<strong>所有数据均为同一台机器实测，含失败组合。</strong></p>
]]></description><link>https://lcz.me/post/19214</link><guid isPermaLink="true">https://lcz.me/post/19214</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Fri, 18 Sep 2026 18:17:03 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Fri, 18 Sep 2026 17:05:52 GMT]]></title><description><![CDATA[<p dir="auto">o对了，我这个显卡目前还是节能模式，没切到oc的bios，因为还得折腾vbios就先这样了</p>
]]></description><link>https://lcz.me/post/19212</link><guid isPermaLink="true">https://lcz.me/post/19212</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Fri, 18 Sep 2026 17:05:52 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Fri, 18 Sep 2026 16:50:08 GMT]]></title><description><![CDATA[<h1>单张 RX 7900 XTX + PVE 直通：Qwen3.8-27B 本地大模型部署实录</h1>
<p dir="auto">补图，自己动手翻新了旧机箱，上岁数了吧，非常喜欢这种老情人回到18个感觉……</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/cc4f5c34-971e-44dc-9e6b-4cba5661c0aa.jpg" alt="72d9deeca311b3893403fa601370519.jpg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><strong>先说结论，PVE下X99加7900xtx基本可用，40tokens加98k上下文</strong>。与物理安装差距不大。另外大模型与基础平台的性能关系也没有预想那么大，物理配置切一半的情况下效果没有太大弱化。算是达到了购买目标。<br />
配置如下：<br />
<img src="https://upload.lcz.me/uploads/5404ae3e-d9f3-4d62-8bce-18c4da7856f8.JPG" alt="千问2.JPG" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/bc786f9d-e930-436e-87d1-b36c0bb05bbd.jpg" alt="fc5fb08e1e9183dc89cf7bafbe73ea1.jpg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/286b1ad1-64a3-402c-9c83-df06f1de6639.jpg" alt="0b7f5b580d8266705e754f009b94406.jpg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/725f88a2-9fd4-4bc5-a1bc-4f98e993c2b5.jpg" alt="8a5e00134a4798f073ddf29a47e564b.jpg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/57c3b478-9647-4d4b-89d1-2784441fcb30.jpg" alt="a32090db5289a449bb9ae92acd47f15.jpg" class=" img-fluid img-markdown" /></p>
<p dir="auto">另外发现一个很迷的问题……见图。</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/ec161ad0-d814-4d44-bb12-0d0881873dfc.JPG" alt="千问.JPG" class=" img-fluid img-markdown" /></p>
<h2>一、最终成果</h2>
<p dir="auto">在 PVE 虚拟机中直通单张 RX 7900 XTX，跑 Qwen3.8-27B Q4_K_M：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指标</th>
<th>实测值</th>
</tr>
</thead>
<tbody>
<tr>
<td>**</td>
<td>Decode</td>
<td><strong>37–41 t/s</strong></td>
</tr>
<tr>
<td>Prefill</td>
<td><strong>195–199 t/s</strong></td>
<td>**</td>
</tr>
<tr>
<td>上下文</td>
<td><strong>98304（98K）</strong></td>
</tr>
<tr>
<td>KV Cache</td>
<td>K q8_0 / V q4_1</td>
</tr>
<tr>
<td>显存占用</td>
<td>模型 + 98K KV 全部在 24GB VRAM 内</td>
</tr>
</tbody>
</table>
<hr />
<h2>二、硬件与软件环境</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>配置</th>
</tr>
</thead>
<tbody>
<tr>
<td>主板</td>
<td>华南金牌 X99-TF</td>
</tr>
<tr>
<td>CPU</td>
<td>Xeon E5-2696V3 （分配了8核心16线程）</td>
</tr>
<tr>
<td>内存</td>
<td>128GB DDR3 1600 ECC （分配了60G固定内存）</td>
</tr>
<tr>
<td>显卡</td>
<td>AMD Radeon RX 7900 XTX 24GB（Navi 31 / gfx1100）</td>
</tr>
<tr>
<td>虚拟化</td>
<td>PVE，Ubuntu 24.04 虚拟机，GPU 直通</td>
</tr>
<tr>
<td>后端</td>
<td>Vulkan（Mesa 25.2.8 RADV）</td>
</tr>
<tr>
<td>推理引擎</td>
<td>llama.cpp build 11037（commit 44be98f05）</td>
</tr>
<tr>
<td>模型</td>
<td>unsloth/Qwen3.8-27B-GGUF，Qwen3.8-27B-UD-Q4_K_M.gguf（16.5G）</td>
</tr>
</tbody>
</table>
<hr />
<h2>三、编译 llama.cpp</h2>
<pre><code class="language-bash">sudo apt update
sudo apt install -y git build-essential cmake libvulkan-dev glslc spirv-headers

git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON
cmake --build build --config Release -j$(nproc)
</code></pre>
<p dir="auto"><strong>注意</strong>：<code>spirv-headers</code> 是较新版本新增的强制依赖，旧教程不会提到。缺它会在 CMake 配置阶段报 <code>Could not find SPIRV-Headers</code>。</p>
<p dir="auto">验证 Vulkan 设备：</p>
<pre><code class="language-bash">./build/bin/llama-cli --list-devices
# 应输出：Vulkan0: AMD Radeon RX 7900 XTX (RADV NAVI31) (24576 MiB, 23624 MiB free)
</code></pre>
<hr />
<h2>四、权限配置（PVE 直通必做）</h2>
<p dir="auto"><code>/dev/dri/renderD128</code> 属于 <code>root:render</code>，普通用户默认无权限访问。这会导致 Vulkan 报 <code>Permission denied (VK_ERROR_INCOMPATIBLE_DRIVER)</code>。</p>
<pre><code class="language-bash">sudo usermod -aG render,video $USER
</code></pre>
<p dir="auto"><strong>关键</strong>：加组后必须<strong>完全退出 SSH 会话重新登录</strong>，组权限才生效。验证：</p>
<pre><code class="language-bash">groups
# 应包含 render 和 video
vulkaninfo --summary | grep -A5 "GPU0"
# 应显示 AMD Radeon RX 7900 XTX (RADV NAVI31)
</code></pre>
<hr />
<h2>五、下载模型</h2>
<pre><code class="language-bash">pipx install "huggingface_hub[cli]"
# 新版命令是 hf，不是 huggingface-cli

hf download unsloth/Qwen3.8-27B-GGUF \
  --include "*Q4_K_M*" \
  --local-dir ~/models
</code></pre>
<p dir="auto">下载不稳定时可用 <code>hfd.sh</code>（基于 aria2c，支持断点续传）。</p>
<hr />
<h2>六、核心调参过程与实测数据</h2>
<h3>6.1 初始配置（失败）</h3>
<pre><code class="language-bash">--ctx-size 65536 --cache-type-k q8_0 --cache-type-v q8_0 \
--spec-type draft-mtp,ngram-map-k4v --spec-draft-n-max 3
</code></pre>
<p dir="auto">结果：<strong>decode 13.12 t/s</strong>，接受率 0.490。</p>
<p dir="auto">问题：<code>--spec-ngram-map-k4v</code> 在 Vulkan 后端产生巨大开销，反而拖慢。</p>
<h3>6.2 去掉 n-gram（改善但仍有断崖）</h3>
<pre><code class="language-bash">--ctx-size 8192 --cache-type-k q8_0 --cache-type-v q8_0 \
--spec-type draft-mtp --spec-draft-n-max 3
</code></pre>
<p dir="auto">结果：<strong>decode 51.67 t/s</strong>，接受率 0.449。</p>
<p dir="auto">但 ctx 提到 16384/32768/65536/98304 后，全部掉到 <strong>9–14 t/s</strong>。</p>
<h3>6.3 定位断崖</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>ctx</th>
<th>KV</th>
<th>decode</th>
</tr>
</thead>
<tbody>
<tr>
<td>8192</td>
<td>q8_0/q8_0</td>
<td>51.67</td>
</tr>
<tr>
<td>16384</td>
<td>q8_0/q8_0</td>
<td>14.18</td>
</tr>
<tr>
<td>32768</td>
<td>q8_0/q8_0</td>
<td>14.47</td>
</tr>
<tr>
<td>65536</td>
<td>q8_0/q8_0</td>
<td>13.12</td>
</tr>
<tr>
<td>98304</td>
<td>q8_0/q8_0</td>
<td>11–15</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>断崖在 8192 和 16384 之间。</strong></p>
<h3>6.4 第一次修复尝试：改 KV 量化</h3>
<pre><code class="language-bash">--ctx-size 16384 --cache-type-k q8_0 --cache-type-v q4_1 \
--ctx-checkpoints 2 --spec-draft-n-max 5
</code></pre>
<p dir="auto">结果：<strong>decode 46.08 t/s</strong>，接受率 0.439。</p>
<p dir="auto">V 从 q8_0 降到 q4_1 后，16384 的断崖消失。但 ctx 提到 65536/98304 后仍掉回 10–13 t/s。</p>
<h3>6.5 根因定位：256M BAR</h3>
<pre><code class="language-bash">lspci -v -s 01:00.0 | grep -i "size="
# Memory at 380000000000 (64-bit, prefetchable) [size=256M]
</code></pre>
<p dir="auto"><strong>PVE 直通环境下 BAR 只有 256M，ReBAR 未生效。</strong></p>
<p dir="auto">llama.cpp 的 Vulkan 后端默认尝试使用这 256MB 的主机可见窗口。ctx 增大后，KV cache 和计算 buffer 装不下，被迫走碎片化的低效路径，decode 断崖。</p>
<h3>6.6 最终修复：禁用 host visible VRAM</h3>
<pre><code class="language-bash">export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1
</code></pre>
<p dir="auto">让 Vulkan 后端完全放弃 256MB 窗口，所有 buffer 直接分配到设备本地 VRAM。</p>
<p dir="auto">结果：<strong>98304 ctx 下 decode 41.25 t/s，prefill 199.38 t/s</strong>，断崖彻底解决。</p>
<hr />
<h2>七、最终可用配置</h2>
<pre><code class="language-bash">#!/bin/bash
# ~/start-llama.sh

export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1

cd /home/kby/llama.cpp
exec ./build/bin/llama-server \
  -m /home/kby/llama.cpp/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  --device Vulkan0 \
  -ngl 99 \
  --parallel 1 \
  --ctx-size 98304 \
  --flash-attn on \
  --cache-type-k q8_0 --cache-type-v q4_1 \
  --ctx-checkpoints 2 \
  --spec-type draft-mtp \
  --spec-draft-n-max 5 \
  --cache-ram 32768 \
  --jinja
</code></pre>
<pre><code class="language-bash">chmod +x ~/start-llama.sh
</code></pre>
<p dir="auto">后台运行：</p>
<pre><code class="language-bash">nohup ~/start-llama.sh &gt; ~/llama-server.log 2&gt;&amp;1 &amp;
</code></pre>
<hr />
<h2>八、参数说明</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>参数</th>
<th>作用</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1</code></td>
<td><strong>最关键</strong>。PVE 直通下 BAR 只有 256M，此变量强制所有 buffer 走设备本地 VRAM，绕过断崖</td>
</tr>
<tr>
<td><code>--device Vulkan0</code></td>
<td>多设备时明确指定显卡</td>
</tr>
<tr>
<td><code>-ngl 99</code></td>
<td>全部层上 GPU</td>
</tr>
<tr>
<td><code>--parallel 1</code></td>
<td>单槽独占 98K 上下文</td>
</tr>
<tr>
<td><code>--ctx-size 98304</code></td>
<td>98K 上下文</td>
</tr>
<tr>
<td><code>--cache-type-k q8_0 --cache-type-v q4_1</code></td>
<td>K 精度优先，V 用 q4_1（带 scale+min，比 q4_0 稳）</td>
</tr>
<tr>
<td><code>--ctx-checkpoints 2</code></td>
<td>稳定性四件套之一，不影响速度</td>
</tr>
<tr>
<td><code>--spec-type draft-mtp</code></td>
<td>启用模型内建 MTP 投机解码。<strong>不要加 ngram-map-k4v</strong>，Vulkan 下反而拖慢</td>
</tr>
<tr>
<td><code>--spec-draft-n-max 5</code></td>
<td>一次猜 5 个 token</td>
</tr>
<tr>
<td><code>--cache-ram 32768</code></td>
<td>prompt cache 开到 32GB，利用 128GB 内存，多轮对话避免重复 prefill</td>
</tr>
<tr>
<td><code>--jinja</code></td>
<td>启用 chat template</td>
</tr>
</tbody>
</table>
<hr />
<h2>九、性能对比总表</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>配置</th>
<th>ctx</th>
<th>KV</th>
<th>n_max</th>
<th>decode</th>
<th>接受率</th>
<th>判定</th>
</tr>
</thead>
<tbody>
<tr>
<td>MTP+ngram</td>
<td>65536</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>13.12</td>
<td>0.490</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> ngram 拖慢</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>8192</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>51.67</td>
<td>0.449</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /></td>
</tr>
<tr>
<td>纯 MTP</td>
<td>16384</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>14.18</td>
<td>0.551</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 断崖</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>32768</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>14.47</td>
<td>0.570</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 断崖</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>98304</td>
<td>q8_0/q8_0</td>
<td>3</td>
<td>11–15</td>
<td>—</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 断崖</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>16384</td>
<td><strong>q8_0/q4_1</strong></td>
<td>5</td>
<td><strong>46.08</strong></td>
<td>0.439</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 断崖消失</td>
</tr>
<tr>
<td>纯 MTP</td>
<td>98304</td>
<td>q8_0/q4_1</td>
<td>5</td>
<td>11.04</td>
<td>0.377</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/274c.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--x" style="height:23px;width:auto;vertical-align:middle" title="❌" alt="❌" /> 仍断崖</td>
</tr>
<tr>
<td><strong>最终</strong></td>
<td><strong>98304</strong></td>
<td><strong>q8_0/q4_1</strong></td>
<td><strong>5</strong></td>
<td><strong>41.25</strong></td>
<td>0.345</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> <strong>禁用 host VRAM 后</strong></td>
</tr>
</tbody>
</table>
<hr />
<h2>十、踩坑清单</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>#</th>
<th>坑</th>
<th>现象</th>
<th>解法</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>缺 <code>spirv-headers</code></td>
<td>CMake 报 <code>Could not find SPIRV-Headers</code></td>
<td><code>sudo apt install spirv-headers</code></td>
</tr>
<tr>
<td>2</td>
<td>用户不在 <code>render</code> 组</td>
<td>Vulkan 报 <code>Permission denied (VK_ERROR_INCOMPATIBLE_DRIVER)</code></td>
<td><code>sudo usermod -aG render,video $USER</code>，重登录</td>
</tr>
<tr>
<td>3</td>
<td>桌面会话占用 GPU</td>
<td>llama-bench 卡死、GPU 利用率 0%</td>
<td>停掉 gdm3，或让桌面用另一张卡</td>
</tr>
<tr>
<td>4</td>
<td><code>ngram-map-k4v</code> 在 Vulkan 下拖慢</td>
<td>decode 从 51 掉到 13</td>
<td>去掉该参数，只用 <code>draft-mtp</code></td>
</tr>
<tr>
<td>5</td>
<td>256M BAR 导致 ctx 断崖</td>
<td>16K 以上 decode 掉到 14</td>
<td><code>export GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1</code></td>
</tr>
<tr>
<td>6</td>
<td>短 prompt 测 prefill</td>
<td>59 token 测出 20 t/s 假数字</td>
<td>用 ≥2000 token 的 prompt 测</td>
</tr>
<tr>
<td>7</td>
<td>同名 GGUF 可能没有 MTP 层</td>
<td>报 <code>model doesn't contain MTP layers</code></td>
<td>换 unsloth UD 版本，确认 <code>blk.64.nextn.*</code> 存在</td>
</tr>
</tbody>
</table>
<hr />
<h2>十一、关键经验</h2>
<ol>
<li><strong>PVE 直通下 ReBAR 通常不生效</strong>，BAR 只有 256M。<code>GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1</code> 是必加项，不是可选项。</li>
<li><strong>Vulkan 后端的 <code>ngram-map-k4v</code> 在 RDNA3 上表现异常</strong>，社区在 ROCm/HIP 下的收益无法复现，建议只用 <code>draft-mtp</code>。</li>
<li><strong>KV 量化不对称是正解</strong>：K q8_0 / V q4_1。K 决定注意力方向，敏感；V 是加权平均，容忍度高。</li>
<li><strong>断崖是确定性的，不是脏状态</strong>。冷启动、子分配修复都无效，只有禁用 host visible VRAM 才解决。</li>
<li><strong>测速必须固定内容类型</strong>。代码生成接受率高（0.4–0.6），散文接受率低（0.3），同一配置速度可差 2–3 倍。</li>
</ol>
<hr />
<h2>十二、附录：Vulkan 设备识别验证</h2>
<pre><code class="language-bash"># 确认 Vulkan 设备可见
vulkaninfo --summary | grep -A10 "GPU0"

# 确认 BAR 大小
lspci -v -s 01:00.0 | grep -i "size="

# 确认用户组
groups

# 确认 llama.cpp 链接了 Vulkan
ldd build/bin/llama-server | grep -i vulkan
</code></pre>
<hr />
<p dir="auto"><strong>测试日期</strong>：2026-09-18<br />
<strong>环境</strong>：PVE + Ubuntu 24.04 + RX 7900 XTX 直通 + llama.cpp build 11037 + Mesa 25.2.8<br />
<strong>所有数据均为同一台机器实测，含失败组合。</strong></p>
]]></description><link>https://lcz.me/post/19201</link><guid isPermaLink="true">https://lcz.me/post/19201</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Fri, 18 Sep 2026 16:50:08 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Fri, 18 Sep 2026 14:12:21 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kang-benyuan" aria-label="Profile: Kang-Benyuan">@<bdi>Kang-Benyuan</bdi></a> 那你发帖的时候注意点，超时不允许编辑我也没办法。</p>
]]></description><link>https://lcz.me/post/19180</link><guid isPermaLink="true">https://lcz.me/post/19180</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Fri, 18 Sep 2026 14:12:21 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Fri, 18 Sep 2026 03:55:31 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a><br />
版主帖子超时无法编辑了</p>
]]></description><link>https://lcz.me/post/18999</link><guid isPermaLink="true">https://lcz.me/post/18999</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Fri, 18 Sep 2026 03:55:31 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Fri, 18 Sep 2026 02:33:42 GMT]]></title><description><![CDATA[<h1>X99 平台 Proxmox VE 直通 AMD RX 7900 XTX 完整实战指南</h1>
<blockquote>
<p dir="auto">适用平台：华南金牌 X99-TF + Intel Xeon E5 v3/v4（40 通道 CPU 最佳）<br />
目标显卡：Sapphire NITRO+ RX 7900 XTX Vapor-X（设备 ID <code>1002:744c</code>，音频 ID <code>1002:ab30</code>）<br />
PVE 版本：Proxmox VE 8.x / 9.x，内核 7.0.2-6-pve<br />
亮机卡：NVIDIA GeForce GT 620（<code>10de:1049</code>，用于宿主机显示输出）<br />
目标 Guest：Ubuntu 24.04 LTS</p>
</blockquote>
<p dir="auto"><strong>写在前面</strong>：AMD 官方<strong>明确声明 RX 7900 XTX 的 PCI Passthrough 不在支持范围内</strong>，所有可用的方案均属于社区实践。本指南记录的是在 X99 平台上<strong>实际跑通</strong>的完整配置，但不保证在其他主板上完全复现。X99 平台存在 PCIe 根端口 AER 报错这一硬件级限制，需要在软件层面尽量规避。</p>
<hr />
<h2>一、硬件准备</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>设备</th>
<th>位置</th>
<th>用途</th>
</tr>
</thead>
<tbody>
<tr>
<td>RX 7900 XTX</td>
<td>离 CPU 最近的 PCIe x16 插槽（Port 1A）</td>
<td>直通给 Guest</td>
</tr>
<tr>
<td>GT 620</td>
<td>第二条 x16 或 x4 插槽</td>
<td>PVE 宿主机显示输出</td>
</tr>
<tr>
<td>显示器</td>
<td>接在 GT 620 上</td>
<td>用于 PVE 安装和救援</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>关键原则</strong>：7900 XTX 必须插在直连 CPU 的插槽，保证 IOMMU 分组独立且 PCIe 带宽为 x16。GT 620 负责宿主机显示，两张卡在不同的 IOMMU 组（本例中 GT 620 在 Group 31，7900 XTX 在 Group 35）。</p>
<hr />
<h2>二、BIOS 设置（华南金牌 X99-TF）</h2>
<p dir="auto">进入 BIOS（Aptio Setup Utility），完成以下配置：</p>
<h3>2.1 <code>IntelRCSetup</code> 标签页</h3>
<p dir="auto">进入 <code>IIO Configuration</code>，设置：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>选项</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>Intel VT for Directed I/O (VT-d)</td>
<td><code>[Enabled]</code></td>
</tr>
<tr>
<td>ACS Control</td>
<td><code>[Enabled]</code></td>
</tr>
<tr>
<td>Interrupt Remapping</td>
<td><code>[Enabled]</code></td>
</tr>
<tr>
<td>Coherency Support (Non-Isoch)</td>
<td><code>[Enabled]</code></td>
</tr>
<tr>
<td>Coherency Support (Isoch)</td>
<td><code>[Enabled]</code></td>
</tr>
<tr>
<td>VTd Azalea Vcp Optimizations</td>
<td><code>[Disable]</code>（保持默认）</td>
</tr>
</tbody>
</table>
<p dir="auto">进入 <code>IIO0 Configuration</code>，确认两个全长插槽都是 x16：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>选项</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>IOU0 (IIO PCIe Port 2)</td>
<td><code>[x16]</code></td>
</tr>
<tr>
<td>IOU1 (IIO PCIe Port 3)</td>
<td><code>[x16]</code></td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto">注意：X99-TF 的 BIOS 中<strong>没有</strong>独立的 <code>Link Speed</code> 选项（子菜单中也没有），速率由硬件自动协商为 Gen3。</p>
</blockquote>
<h3>2.2 <code>Advanced</code> 标签页</h3>
<p dir="auto">进入 <code>PCI Subsystem Settings</code>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>选项</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>Above 4G Decoding</td>
<td><code>[Enabled]</code></td>
</tr>
<tr>
<td>Re-Size BAR Support</td>
<td><code>[Enabled]</code>（部分 BIOS 版本需关闭，见下方说明）</td>
</tr>
<tr>
<td>SR-IOV Support</td>
<td><code>[Enabled]</code></td>
</tr>
</tbody>
</table>
<p dir="auto">进入 <code>CSM Configuration</code>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>选项</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>CSM Support</td>
<td><code>[Disabled]</code>（使用 UEFI 引导）</td>
</tr>
</tbody>
</table>
<blockquote>
<p dir="auto"><strong>关于 Re-Size BAR 的取舍</strong>：如果直通后出现频繁 AER 报错或虚拟机启动不稳定，将 <code>Re-Size BAR Support</code> 改为 <code>[Disabled]</code> 再测试。这会牺牲少量性能，但能显著提升兼容性。</p>
</blockquote>
<h3>2.3 关于板载显卡选项</h3>
<p dir="auto">华南金牌 X99-TF <strong>没有物理板载显卡</strong>，BIOS 中也<strong>没有</strong> <code>Onboard VGA Control</code> 或 <code>Primary Display</code> 相关选项。直接使用 GT 620 作为宿主机显示输出即可，无需额外设置。</p>
<p dir="auto">保存退出：按 <code>F10</code>。</p>
<hr />
<h2>三、PVE 宿主机配置</h2>
<h3>3.1 安装 PVE</h3>
<p dir="auto">用 U 盘安装 PVE 8.x 或 9.x。安装时显示器接在 GT 620 上。</p>
<h3>3.2 开启 IOMMU</h3>
<p dir="auto">编辑 GRUB 配置：</p>
<pre><code class="language-bash">nano /etc/default/grub
</code></pre>
<p dir="auto">将 <code>GRUB_CMDLINE_LINUX_DEFAULT</code> 这一行<strong>完整替换</strong>为（注意：必须是一整行，不能有换行）：</p>
<pre><code class="language-text">GRUB_CMDLINE_LINUX_DEFAULT="quiet intel_iommu=on iommu=pt pci=noaer video=efifb:off initcall_blacklist=sysfb_init"
</code></pre>
<p dir="auto">保存退出后执行：</p>
<pre><code class="language-bash">update-grub
update-initramfs -u
reboot
</code></pre>
<p dir="auto">重启后验证 IOMMU：</p>
<pre><code class="language-bash">dmesg | grep -e DMAR -e IOMMU
</code></pre>
<p dir="auto">应看到 <code>DMAR: IOMMU enabled</code>。</p>
<h3>3.3 获取显卡设备 ID</h3>
<pre><code class="language-bash">lspci -nn | grep -i vga
lspci -nn | grep -i audio
</code></pre>
<p dir="auto">本例输出：</p>
<ul>
<li>7900 XTX 显卡：<code>05:00.0</code>，ID <code>1002:744c</code></li>
<li>7900 XTX 音频：<code>05:00.1</code>，ID <code>1002:ab30</code></li>
<li>GT 620：<code>01:00.0</code>，ID <code>10de:1049</code>（不需要屏蔽）</li>
</ul>
<h3>3.4 绑定 vfio-pci</h3>
<p dir="auto">创建 VFIO 配置（<strong>不要</strong>加入 <code>disable_vga=1</code>，该参数会导致 X99 平台启动死锁）：</p>
<pre><code class="language-bash">cat &lt;&lt; EOF &gt; /etc/modprobe.d/vfio.conf
options vfio-pci ids=1002:744c,1002:ab30
EOF
</code></pre>
<p dir="auto">拉黑宿主机 AMD 驱动：</p>
<pre><code class="language-bash">cat &lt;&lt; EOF &gt; /etc/modprobe.d/pve-blacklist.conf
blacklist amdgpu
blacklist radeon
EOF
</code></pre>
<p dir="auto">让 PVE 开机自动加载 VFIO 模块：</p>
<pre><code class="language-bash">echo -e "vfio\nvfio_iommu_type1\nvfio_pci\nvfio_virqfd" &gt;&gt; /etc/modules
</code></pre>
<p dir="auto">更新 initramfs 并重启：</p>
<pre><code class="language-bash">update-initramfs -u
reboot
</code></pre>
<h3>3.5 验证 vfio-pci 接管成功</h3>
<pre><code class="language-bash">lspci -nnk | grep -A 3 "VGA compatible controller"
</code></pre>
<p dir="auto">7900 XTX（<code>05:00.0</code>）下面应显示：</p>
<pre><code class="language-text">Kernel driver in use: vfio-pci
</code></pre>
<p dir="auto">GT 620（<code>01:00.0</code>）下面应显示：</p>
<pre><code class="language-text">Kernel driver in use: nvidiafb
</code></pre>
<h3>3.6 验证 IOMMU 分组独立</h3>
<pre><code class="language-bash">for d in /sys/kernel/iommu_groups/*/devices/*; do n="${d#*/iommu_groups/*}"; n="${n%%/*}"; printf 'IOMMU group %s ' "$n"; lspci -nns "${d##*/}"; done | grep -E "1002:744c|10de:1049"
</code></pre>
<p dir="auto">本例输出：</p>
<pre><code class="language-text">IOMMU group 31 01:00.0 VGA compatible controller [0300]: NVIDIA Corporation GF119 [GeForce GT 620 OEM] [10de:1049] (rev a1)
IOMMU group 35 05:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 31 [Radeon RX 7900 XT/7900 XTX/7900 GRE/7900M] [1002:744c] (rev c8)
</code></pre>
<p dir="auto">两张卡在不同 Group，说明分组干净，可以单独隔离。</p>
<hr />
<h2>四、vBIOS 文件准备</h2>
<p dir="auto"><strong>必须从自己这张显卡上提取</strong>，不要从网上下载第三方 vBIOS。</p>
<ol>
<li>将 7900 XTX 装到一台 Windows 机器上。</li>
<li>用 GPU-Z 软件导出 vBIOS，保存为 <code>7900xtx.rom</code>。</li>
<li>上传到 PVE 宿主机 <code>/usr/share/kvm/</code> 目录。</li>
<li>验证文件：</li>
</ol>
<pre><code class="language-bash">ls -lh /usr/share/kvm/7900xtx.rom
</code></pre>
<p dir="auto">应显示约 2.0M 大小。</p>
<hr />
<h2>五、创建 Ubuntu 虚拟机</h2>
<h3>5.1 创建 VM</h3>
<p dir="auto">在 PVE Web UI 中创建虚拟机，关键参数：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>项目</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>VM ID</td>
<td>101</td>
</tr>
<tr>
<td>名称</td>
<td>Ubuntu24.04</td>
</tr>
<tr>
<td>BIOS</td>
<td>OVMF (UEFI)</td>
</tr>
<tr>
<td>Machine</td>
<td>q35</td>
</tr>
<tr>
<td>CPU</td>
<td>host</td>
</tr>
<tr>
<td>核心数</td>
<td>16</td>
</tr>
<tr>
<td>内存</td>
<td>61440 MB</td>
</tr>
<tr>
<td>磁盘</td>
<td>388G</td>
</tr>
<tr>
<td>网络</td>
<td>virtio, bridge=vmbr0</td>
</tr>
<tr>
<td>显示</td>
<td>默认（先用 QEMU 虚拟显卡）</td>
</tr>
<tr>
<td>引导顺序</td>
<td>sata0;net0</td>
</tr>
</tbody>
</table>
<h3>5.2 编辑虚拟机配置文件</h3>
<pre><code class="language-bash">nano /etc/pve/qemu-server/101.conf
</code></pre>
<p dir="auto">当前跑通的配置内容：</p>
<pre><code class="language-conf">balloon: 0
bios: ovmf
boot: order=sata0;net0
cores: 16
cpu: host
efidisk0: local-lvm:vm-101-disk-0,efitype=4m,ms-cert=2023k,pre-enrolled-keys=1,size=4M
hostpci0: 0000:05:00,pcie=1,romfile=7900xtx.rom
ide2: none,media=cdrom
machine: q35
memory: 61440
meta: creation-qemu=11.0.0,ctime=1789626016
name: Ubuntu24.04
net0: virtio=BC:24:11:7B:0F:27,bridge=vmbr0,firewall=1
numa: 0
ostype: l26
sata0: local-lvm:vm-101-disk-1,size=388G
scsihw: virtio-scsi-single
smbios1: uuid=e144b18b-b6fb-4350-bc4c-b1c9659bfe59
sockets: 1
vmgenid: afd026f7-60b3-4c56-a1ca-5d67d88245b7
</code></pre>
<p dir="auto"><strong>关键说明</strong>：</p>
<ul>
<li><code>hostpci0: 0000:05:00,pcie=1,romfile=7900xtx.rom</code> — 使用 <code>pcie=1</code> 启用 PCIe 直通，加载 vBIOS</li>
<li><strong>不要加</strong> <code>args: -fw_cfg name=opt/ovmf/X-PciMmio64Mb,string=65536</code>，X99 平台会因此触发 AER 报错</li>
<li><strong>不要加</strong> <code>disable_vga=1</code>，会导致启动死锁</li>
<li><code>x-vga=1</code> 可以视 Guest 内输出情况决定是否添加，本例最终<strong>未使用</strong></li>
</ul>
<h3>5.3 给虚拟机添加串口（调试用）</h3>
<pre><code class="language-bash">qm set 101 -serial0 socket
</code></pre>
<p dir="auto">之后可以用 <code>qm terminal 101</code> 通过串口查看引导输出。</p>
<hr />
<h2>六、安装 Ubuntu 24.04</h2>
<h3>6.1 先不加直通，用虚拟显卡装完系统</h3>
<p dir="auto"><strong>重要</strong>：先用最简配置完成 Ubuntu 安装，避免直通干扰安装流程。</p>
<pre><code class="language-bash"># 临时去掉直通
qm set 101 -delete hostpci0
</code></pre>
<p dir="auto">挂载 Ubuntu ISO：</p>
<pre><code class="language-bash">ls /var/lib/vz/template/iso/
qm set 101 -ide2 local:iso/你的ubuntu镜像.iso,media=cdrom
qm set 101 -boot order=ide2;sata0;net0
</code></pre>
<p dir="auto">启动 VM：</p>
<pre><code class="language-bash">qm start 101
</code></pre>
<p dir="auto">通过 PVE Web UI → VM 101 → Console → noVNC 完成 Ubuntu 安装。</p>
<h3>6.2 安装时确认</h3>
<ul>
<li>设置用户名（本例为 <code>kby</code>）和密码</li>
<li>配置网络（DHCP 或静态 IP）</li>
<li>安装过程中勾选 <code>Install OpenSSH server</code></li>
</ul>
<p dir="auto">安装完成后，记录 Guest 的 IP 地址（本例为 <code>192.168.0.58</code>）。</p>
<h3>6.3 在 Guest 内配置 SSH</h3>
<p dir="auto">SSH 进 Guest：</p>
<pre><code class="language-bash">ssh kby@192.168.0.58
</code></pre>
<p dir="auto">安装并启用 SSH：</p>
<pre><code class="language-bash">sudo apt update
sudo apt install openssh-server -y
sudo systemctl enable --now ssh
sudo systemctl status ssh
</code></pre>
<p dir="auto">看到 <code>active (running)</code> 即可。</p>
<hr />
<h2>七、加回直通并配置 Guest 显示</h2>
<h3>7.1 在 PVE 宿主机加回直通</h3>
<pre><code class="language-bash">qm set 101 -hostpci0 0000:05:00,pcie=1,romfile=7900xtx.rom
</code></pre>
<p dir="auto">重启 VM：</p>
<pre><code class="language-bash">qm stop 101
qm start 101
</code></pre>
<p dir="auto">启动时可能出现以下警告：</p>
<pre><code class="language-text">error writing '1' to '/sys/bus/pci/devices/0000:05:00.0/reset': Inappropriate ioctl for device
failed to reset PCI device '0000:05:00.0', but trying to continue as not all devices need a reset
</code></pre>
<p dir="auto"><strong>这是正常警告，不会阻断启动</strong>，QEMU 会继续执行。</p>
<h3>7.2 Guest 内验证显卡识别</h3>
<p dir="auto">SSH 进 Guest：</p>
<pre><code class="language-bash">ssh kby@192.168.0.58
lspci -nnk | grep -i vga -A3
</code></pre>
<p dir="auto">应看到：</p>
<pre><code class="language-text">01:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 31 [Radeon RX 7900 XT/7900 XTX/7900M] [1002:744c] (rev c8)
        Subsystem: Sapphire Technology Limited NITRO+ RX 7900 XTX Vapor-X [1da2:e471]
        Kernel driver in use: amdgpu
        Kernel modules: amdgpu
</code></pre>
<p dir="auto">以及 <code>00:01.0</code> 的 QEMU 虚拟显卡（<code>bochs-drm</code>）。</p>
<h3>7.3 检查 amdgpu 驱动加载日志</h3>
<pre><code class="language-bash">sudo dmesg | grep -i amdgpu | tail -n 80
</code></pre>
<p dir="auto">关键成功标志：</p>
<pre><code class="language-text">amdgpu: 24560M of VRAM memory ready
[drm] Initialized amdgpu 3.64.0 for 0000:01:00.0 on minor 1
</code></pre>
<p dir="auto">可能出现的非致命报错：</p>
<pre><code class="language-text">amdgpu: PCIe atomic ops is not supported
amdgpu: [drm] Failed to setup vendor infoframe on connector HDMI-A-2: -22
</code></pre>
<p dir="auto"><code>-22</code> 是 HDMI infoframe 格式不合法，会影响 HDMI 音频通道建立，<strong>但不影响显示输出的核心功能</strong>。</p>
<h3>7.4 确认 /dev/dri 设备节点</h3>
<pre><code class="language-bash">ls -l /dev/dri/
</code></pre>
<p dir="auto">应看到：</p>
<pre><code class="language-text">card0（bochs 虚拟显卡）
card1（7900 XTX）
renderD128（7900 XTX 渲染节点）
</code></pre>
<hr />
<h2>八、强制 Xorg 使用 7900 XTX 作为主 GPU</h2>
<p dir="auto"><strong>这一步是关键</strong>：Ubuntu 默认把 <code>card0</code>（bochs 虚拟显卡）当作主显示输出，导致 7900 XTX 虽然驱动加载成功，但拿不到 GDM 的显示绑定。</p>
<h3>8.1 创建 Xorg 配置文件</h3>
<p dir="auto">在 Guest 内：</p>
<pre><code class="language-bash">sudo nano /etc/X11/xorg.conf
</code></pre>
<p dir="auto">填入：</p>
<pre><code class="language-conf">Section "ServerLayout"
    Identifier     "Layout0"
    Screen      0  "Screen0"
EndSection

Section "Device"
    Identifier     "AMD0"
    Driver         "amdgpu"
    BusID          "PCI:1:0:0"
    Option         "PrimaryGPU" "yes"
EndSection

Section "Screen"
    Identifier     "Screen0"
    Device         "AMD0"
    DefaultDepth   24
EndSection
</code></pre>
<p dir="auto">保存：<code>Ctrl+O</code> → <code>Enter</code> → <code>Ctrl+X</code>。</p>
<h3>8.2 强制 GDM 使用 Xorg（禁用 Wayland）</h3>
<pre><code class="language-bash">sudo nano /etc/gdm3/custom.conf
</code></pre>
<p dir="auto">找到 <code>#WaylandEnable=false</code>，去掉前面的 <code>#</code>，变成：</p>
<pre><code class="language-text">WaylandEnable=false
</code></pre>
<p dir="auto">保存退出。</p>
<h3>8.3 更新并重启</h3>
<pre><code class="language-bash">sudo update-initramfs -u
sudo reboot
</code></pre>
<p dir="auto">重启后，<strong>7900 XTX 屏幕应出现 GDM 登录界面</strong>。</p>
<hr />
<h2>九、HDMI 音频配置</h2>
<p dir="auto">如果 HDMI 音频未自动生效，按以下顺序排查：</p>
<h3>9.1 检查系统声音设置</h3>
<p dir="auto">打开 Ubuntu 的 <strong>设置 → 声音</strong>，在“输出设备”中查找 <strong>HDMI / DisplayPort Output</strong> 或 <strong>Navi 31 HDMI Audio</strong>，选中它。</p>
<h3>9.2 检查音频设备直通状态</h3>
<p dir="auto">确保 7900 XTX 的音频设备（<code>05:00.1</code>，ID <code>1002:ab30</code>）也一同直通给 Guest。本例中 <code>hostpci0: 0000:05:00</code> 写的是 <code>05:00.0</code>，由于 <code>05:00.0</code> 和 <code>05:00.1</code> 在同一 IOMMU 组，QEMU 会自动带上音频设备。</p>
<p dir="auto">在 Guest 内验证：</p>
<pre><code class="language-bash">lspci -nnk | grep -i audio
</code></pre>
<p dir="auto">应看到：</p>
<pre><code class="language-text">01:00.1 Audio device [0403]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 31 HDMI/DP Audio [1002:ab30]
</code></pre>
<h3>9.3 若声音断断续续，禁用 WirePlumber 自动挂起</h3>
<pre><code class="language-bash">mkdir -p ~/.config/wireplumber/wireplumber.conf.d/
nano ~/.config/wireplumber/wireplumber.conf.d/disable-suspend.conf
</code></pre>
<p dir="auto">填入：</p>
<pre><code class="language-conf">wireplumber.profiles = {
  main = {
    hooks.node.suspend = disabled
  }
}
</code></pre>
<p dir="auto">然后重启音频服务：</p>
<pre><code class="language-bash">systemctl --user restart pipewire
systemctl --user restart wireplumber
</code></pre>
<h3>9.4 若 HDMI 音频始终异常</h3>
<p dir="auto"><strong>换用 DisplayPort 线</strong>。<code>-22</code> infoframe 报错是 HDMI 通道特有的兼容性问题，DP 通道通常能直接绕过。这是最省心的解决方案。</p>
<hr />
<h2>十、X99 平台长期使用注意事项</h2>
<h3>10.1 <code>failed to reset PCI device</code> 警告</h3>
<p dir="auto">每次 <code>qm start</code> 都可能出现，<strong>不是致命错误</strong>，QEMU 会继续启动。如果遇到 VM 启动后 7900 XTX 不亮：</p>
<pre><code class="language-bash">qm stop 101
# 若 stop 失败
qm stop 101 --skiplock
qm start 101
</code></pre>
<p dir="auto">若仍失败，<strong>冷启动 PVE 主机</strong>（长按电源键关机再开机）恢复 PCIe 状态。</p>
<h3>10.2 AER 错误</h3>
<p dir="auto">X99 平台 PCIe 根端口会持续报 <code>AER: Uncorrectable (Non-Fatal)</code>。这是硬件级限制，已通过 <code>pci=noaer</code> 和 <code>initcall_blacklist=sysfb_init</code> 在内核层面尽量抑制。高负载下若链路崩溃，需冷启动主机。</p>
<h3>10.3 备份关键配置</h3>
<p dir="auto">跑通后立刻备份：</p>
<pre><code class="language-bash">cp /etc/pve/qemu-server/101.conf ~/101.conf.backup
cp /etc/default/grub ~/grub.backup
cp /etc/modprobe.d/vfio.conf ~/vfio.conf.backup
cp /etc/modprobe.d/pve-blacklist.conf ~/pve-blacklist.conf.backup
cp /usr/share/kvm/7900xtx.rom ~/7900xtx.rom.backup
</code></pre>
<h3>10.4 关于鼠标键盘</h3>
<p dir="auto"><strong>不需要专门直通键鼠</strong>。QEMU 默认提供虚拟 USB 控制器和模拟键鼠，物理键鼠在切换显示窗口时通常能直接使用。只有在追求极低延迟（如游戏）或虚拟机需要完全独占输入时，才考虑单设备 USB 直通（不要直通整个 USB 控制器，X99 上有风险）。</p>
<hr />
<h2>十一、完整配置清单速查</h2>
<p dir="auto"><strong>PVE 宿主机</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>文件</th>
<th>内容</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>/etc/default/grub</code></td>
<td><code>GRUB_CMDLINE_LINUX_DEFAULT="quiet intel_iommu=on iommu=pt pci=noaer video=efifb:off initcall_blacklist=sysfb_init"</code></td>
</tr>
<tr>
<td><code>/etc/modprobe.d/vfio.conf</code></td>
<td><code>options vfio-pci ids=1002:744c,1002:ab30</code></td>
</tr>
<tr>
<td><code>/etc/modprobe.d/pve-blacklist.conf</code></td>
<td><code>blacklist amdgpu</code> / <code>blacklist radeon</code></td>
</tr>
<tr>
<td><code>/etc/modules</code></td>
<td><code>vfio</code> / <code>vfio_iommu_type1</code> / <code>vfio_pci</code> / <code>vfio_virqfd</code></td>
</tr>
<tr>
<td><code>/usr/share/kvm/7900xtx.rom</code></td>
<td>从本卡导出的 vBIOS</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>Guest 虚拟机（101.conf）关键行</strong>：</p>
<pre><code class="language-conf">hostpci0: 0000:05:00,pcie=1,romfile=7900xtx.rom
</code></pre>
<p dir="auto"><strong>Guest 内关键配置</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>文件</th>
<th>内容</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>/etc/X11/xorg.conf</code></td>
<td>强制 amdgpu 作为 PrimaryGPU，BusID <code>PCI:1:0:0</code></td>
</tr>
<tr>
<td><code>/etc/gdm3/custom.conf</code></td>
<td><code>WaylandEnable=false</code></td>
</tr>
</tbody>
</table>
<hr />
<h2>十二、最终验证</h2>
<p dir="auto">在 Guest 内执行：</p>
<pre><code class="language-bash">glxinfo | grep "OpenGL renderer"
</code></pre>
<p dir="auto">应显示：</p>
<pre><code class="language-text">OpenGL renderer string: AMD Radeon RX 7900 XTX (radeonsi, navi31, ...)
</code></pre>
<p dir="auto">看到这一行，说明整个直通链路——BIOS → vfio-pci → QEMU → amdgpu → Xorg → GDM → 显示输出 → HDMI 音频——全部打通。</p>
<p dir="auto"><strong>这套配置在华南金牌 X99-TF + RX 7900 XTX 上已实际验证可用。</strong></p>
]]></description><link>https://lcz.me/post/18987</link><guid isPermaLink="true">https://lcz.me/post/18987</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Fri, 18 Sep 2026 02:33:42 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Thu, 17 Sep 2026 19:03:15 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kang-benyuan" aria-label="Profile: Kang-Benyuan">@<bdi>Kang-Benyuan</bdi></a> 直通成功不错。X99 + 2696V3 这套有几点可以顺手记下：</p>
<ul>
<li>7900XTX 不支持 SR-IOV，只能整卡直通；两卡要分别落在独立 IOMMU group 里。X99 上如果分组不干净，要么换插槽，要么用 ACS override（有安全代价，内网可接受）。</li>
<li>BIOS 里 Above 4G Decoding、Resizable BAR、IOMMU 都要开；7900 系有经典的 reset bug，直通 VM 重启前先确认宿主机用 vendor-reset 或内核参数兜住，不然第二次启动会挂。</li>
<li>PVE 里把 /dev/kfd 和 /dev/dri 正确映射给 guest；ROCm 6.x 在 VM 里能跑 llama.cpp/vLLM，但性能比裸机差一截，能用 LXC 就别用 KVM（ROCm 在 LXC 需要手动映射设备）。</li>
<li>单卡 ReBAR 对推理收益有限，主要看显存映射；你以后切 NV 是对的，但 5090/4090 的价差也要算进去。</li>
</ul>
]]></description><link>https://lcz.me/post/18928</link><guid isPermaLink="true">https://lcz.me/post/18928</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 17 Sep 2026 19:03:15 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Thu, 17 Sep 2026 16:38:29 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a></p>
<p dir="auto">1推荐的7900xtx<br />
2……内存便宜忍了<br />
3没拍……回头补上</p>
]]></description><link>https://lcz.me/post/18901</link><guid isPermaLink="true">https://lcz.me/post/18901</guid><dc:creator><![CDATA[Kang Benyuan]]></dc:creator><pubDate>Thu, 17 Sep 2026 16:38:29 GMT</pubDate></item><item><title><![CDATA[Reply to 华南金牌X99-TF 2696V3 PVE直通7900XTX成功 on Thu, 17 Sep 2026 16:32:58 GMT]]></title><description><![CDATA[<p dir="auto">1，我特么从来没推荐华南金牌，我只是说性价比不错，不叫推荐，推荐是大家闭眼买就对了。<br />
2，你单卡开启rebar有个毛用啊，主要是看能否两张XTX，这个很显然不行。<br />
3，你发这个帖子不上图？</p>
]]></description><link>https://lcz.me/post/18899</link><guid isPermaLink="true">https://lcz.me/post/18899</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Thu, 17 Sep 2026 16:32:58 GMT</pubDate></item></channel></rss>