<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告]]></title><description><![CDATA[<h1>Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告</h1>
<p dir="auto">(這是我自己打的, 不是AI Report)<br />
首先，我是一個layman， 一個不懂IT技術的香港普通人。年輕時也曾沉迷打機，能動手自己"砌機"(即組合和安裝電腦硬件)。不過隨著年齡增長，結了婚，有了小朋友，已有二十多年没去深水埗黃金電腦商場。買電腦會考慮買I-mac, mac mini之類的, 說白了就是收入好了, 不再折騰研究組裝，買一體機慳水慳力，可以多留一些時間給老婆仔女。</p>
<p dir="auto">今年Open Claw大熱，我也手癢在Kimi上買了$199月費試玩。没有研究電腦發展多年的我，驚訝地發現AI Agent(或Harness)+LLM的生產能力已遠超過我想像。孩子大了，時間多了，折騰本質又回來。</p>
<p dir="auto">當初找資料時，知道Mac Mini M4 16G可以部署試玩，花了$5000HKD，在二手市場買了一個回來日日夜夜研究。結果可想而知，用HKU Nanobot(香港大學的免費開源AI Agent Bot)，部署個什麼4B 8B模型，正中裴多菲那一首詩，「世上有電腦千萬顆，我卻只有一個蘋果；就這一個也已經太多，她早晚要氣死我……」(對原詩句有興趣的話，大家可以去找一找，很像Terry的幽默調侃的風格)。真的，問個天氣都能氣死你。</p>
<p dir="auto">有句說話「成長就是客觀世界打破主觀世界的一個過程」。接受給水文"4B 7B Token自由"騙了的現實從新出發，在Nanobot上接入Deepseek V4 Flash，終於能實現生產力了。</p>
<p dir="auto">但我就是不甘心，腰骨好似有股氣一樣，站起來說一定要實現本地部署的，有生產力的方案。坦白說，因為我是一個layman，Deepseek V4 Flash對我來說是完全overkill，我自覺根本用不到他的核心能力。我的需求也很明確，公司有一些資料需要蒐集、整理、分析，小孩被"教育恐怖主義"脅持，參加課外活動比上班更加地獄，我是需要一個有生產力的AI助手幫忙處理他們的日程，才能繼續遊戲人間。</p>
<p dir="auto">其實調用deepseek api也很便宜，就是要時時想著"慳住洗"(省著用)不爽。本地部署可以隨便用，一家人亂用一通都得。</p>
<p dir="auto">被水文騙了一次不要緊，精益求精再次做資料蒐集，看了"掄錘者"頻道，發現Terry說的內容都是很juicy，例如"Qwen 3.8 27B 是本地部署模型最好的模型，没有之一。"重點是什麼?不是這一句話"Qwen 3.8 27B 是本地部署模型最好的模型"，這一句隨處可見!重點是他有硬件參數，配置參數，實測接入AI Agent的效果，具體到幾多輪幾多步工作不掉線，能力是deepseek v4 flash 的8成功力。這才是真正的用家體驗。</p>
<p dir="auto">8月28日，掄錘者出了一個叫"AMD W7900/W7800 48G 性價比如何?...."的視頻，我看了Terry的介紹，消化了在論壇裡多個用家報告AMD Radeon 7900XTX 24G的體驗，立即去網上看一下香港W7800 48G的價錢，發現這個性價比非常適合我。當時標價$24999HKD，我拿到報價單時已是$26999(折合人民幣23100-23200)，印象跟Terry說的差不多，於是果斷下單了，不是買單卡，是買一整部電腦。(買的W7800 48G過程也有一些疑慮，搜遍全網也找不到W7800 48G的體驗報告，我寫的這一篇可能是全網第一篇用家體驗)</p>
<p dir="auto">配置如下：<br />
AMD Ryzen 5 7500F CPU (Tray, AM5, 6Core/12Threads, 3.7/5.gGHz, 6MB L2, 32MB L3, TDP)<br />
ASUS華碩 TUF GAMING B850I WIFI NEO AM5 ITX MB<br />
Klevv DDR5 FIT V 黑色 6000MHz 32GB Kit (16GB*2, CLS 28-36-36-76)<br />
Gigabyte技嘉 Radeon PRO W7800 AI TOP 48G Graphic Cards (GV-W7800 AI TOP-48G)<br />
NZXT C Series C850 SFX Gold 850W Fully Modular ATX 3.1 Power Supply (80 Plus Gold, Cybenetics A-, 600W 12V-2x6 PCIe 5.1, 10yr warranty)<br />
WD Black M.2. SN7100 1TB SSD (R:7250MB/s, W6900, WDS100T4X0E)<br />
Fractal Design Ridge White ITX Case 白色(PCIe 4.0, FD-C-RID1N-12)<br />
Thermalright AXP120-X67 ARGB Black CPU Cooler(下吹式散熱設計)</p>
<p dir="auto">你可能會問，為何不買大ATX機箱，散熱較好呀?</p>
<p dir="auto">香港住屋的面積小早已享譽國際。對一般人來說大一點機箱没什麼問題，對我來說每1cm都是戰場。所以雙卡不在考慮之列，大機箱不在考慮之列，因此最終成品就是這個小型運算中心。</p>
<p dir="auto">請大家也不用問我硬件軟件資料是怎麼取捨。我的知識面停留在GTR機箱，EDO Ram和DDR Ram，CPU 386 486等等，你折騰我也折不出什麼來，反正我跟商家要求就是三點，機箱要細，散熱要好，電源要穩定。</p>
<p dir="auto">現在我的工作硬件是這樣的：</p>
<p dir="auto">Mac Mini M4 16g 256g hardisk + 綠聯外置Docker一個NVME 4TB Hardisk +  獨立Inference Server (W7800 48G)</p>
<ul>
<li>mac mini 24/7運行，裝Deeepseek Harness及所有Project</li>
<li>W7800 電腦 24/7 運行，只做Inference Server</li>
<li>mac mini 透過家中的hub調用Inference server</li>
</ul>
<p dir="auto">錢學森說"系統是指由多個相互關聯，相互作用，相互影響的的部份所組成的具有特定功能的有機整體。"</p>
<p dir="auto">好了，經過簡短的介紹後，後面就是用以上配置，DSH + Qwen 3.8 27B (xhigh thinking)，3輪69步AI生成的AI Inference Server Report，歡迎大家隨便使用。</p>
<p dir="auto">(AI生成Report)</p>
<blockquote>
<p dir="auto"><strong><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a0.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--warning" style="height:23px;width:auto;vertical-align:middle" title="⚠" alt="⚠" />️ 9/7 更新（先睇呢段）：</strong> 發文後我連續跑咗 48 小時（1,916 個請求、約 8,500 萬 tokens），將所有結論重新實測過。<strong>原文有三個結論被推翻</strong>：(1) W7800 上面 <strong>ROCm 反而快過 Vulkan</strong>（prefill +38%、MTP n=8 做到 65.1 t/s）；(2) <strong>KV cache 用全精度 f16 最快</strong>（q8_0 慢 3%、慳 4GB 但唔值得）；(3) 最終生產配置係 <strong>Q6_K + 2 個平行 slot（各 128K）+ f16 KV</strong>，唔係 Q8 單 slot。另外新增：SGLang 實測（W7800 上用唔到）、RAM 86% 謎團真相（要加 <code>--no-cache-idle-slots</code>）、兩日實戰數據同溫度記錄。詳細內容喺文末「9/7 更新」一節，原文保留作記錄。</p>
</blockquote>
<hr />
<h2>我係邊個</h2>
<p dir="auto">呢篇帖子就係我嘅實測記錄。如果你都係同我一樣，唔係技術人員但想試下本地 AI，希望呢份報告對你有幫助。</p>
<hr />
<h2>硬件配置</h2>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>部件</th>
<th>型號</th>
</tr>
</thead>
<tbody>
<tr>
<td>GPU</td>
<td>AMD Radeon Pro W7800 48GB（70 CU · 384-bit · 864 GB/s）</td>
</tr>
<tr>
<td>CPU</td>
<td>AMD Ryzen 7 7500F</td>
</tr>
<tr>
<td>主機板</td>
<td>ASUS B650M</td>
</tr>
<tr>
<td>記憶體</td>
<td>32GB DDR5</td>
</tr>
<tr>
<td>存儲</td>
<td>1TB NVMe SSD</td>
</tr>
<tr>
<td>系統</td>
<td>Ubuntu 24.04</td>
</tr>
<tr>
<td>驅動</td>
<td>Mesa Vulkan（RADV）</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>BIOS 重要設定：</strong></p>
<ul>
<li>Above 4G Decoding（ASUS 新版叫 Above 4G MMIO Limit）→ <strong>Enabled</strong></li>
<li>Resizable BAR → <strong>Enabled</strong></li>
</ul>
<hr />
<h2>軟件棧</h2>
<ul>
<li><strong>llama.cpp</strong>：Vulkan build（cmake -DGGML_VULKAN=ON）</li>
<li><strong>模型</strong>：Qwen3.8-27B UD-Q4_K_M（Unsloth Dynamic 量化，16.4GB）</li>
<li><strong>推理引擎</strong>：llama-server（port 8080）</li>
<li><strong>前端</strong>：Hermes Agent（Mac Mini 遙距控制）</li>
</ul>
<blockquote>
<p dir="auto"><strong>[9/7 更新]</strong> 而家生產已轉去 <strong>ROCm 7.2.4 build</strong>（實測更快，見文末 9/7 更新），模型升級到 <strong>UD-Q6_K（22GB）</strong>，用 systemd 服務自動開機啟動、崩咗自動重啟。Vulkan build 留返做備份。</p>
</blockquote>
<p dir="auto"><strong>依賴套件（缺一不可）：</strong></p>
<pre><code class="language-bash">sudo apt install -y mesa-vulkan-drivers vulkan-tools libvulkan-dev \
  glslc spirv-headers libshaderc-dev git build-essential cmake ninja-build
</code></pre>
<p dir="auto"><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a0.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--warning" style="height:23px;width:auto;vertical-align:middle" title="⚠" alt="⚠" />️ <strong>glslc 唔係 glslang-tools</strong> —— 兩個唔同嘅 package。我 first build 時就錯裝咗 glslang-tools，configure 失敗，查咗好耐先發現。</p>
<hr />
<h2>推理服務配置（生產環境）</h2>
<blockquote>
<p dir="auto"><strong>[9/7 更新]</strong> 下面係 9/4 首日版本（Vulkan + Q4 + 單 slot 131K），已退役。而家嘅生產版本（ROCm + Q6_K + 2 slot × 128K + f16 KV）見文末「9/7 更新」第 6 節。</p>
</blockquote>
<pre><code class="language-bash">#!/bin/bash
exec ~/llama.cpp/build-vulkan/bin/llama-server \
  -m ~/models/Qwen3.8-27B-UD-Q4_K_M.gguf \
  --alias qwen3.8-27b --device Vulkan0 \
  --fit off -ngl -1 --no-mmap \
  --spec-type draft-mtp --spec-draft-n-max 2 \
  -c 131072 -ub 512 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --parallel 1 --cache-ram 32768 --flash-attn on \
  --mmproj ~/models/mmproj-F16.gguf \
  --image-min-tokens 1024 \
  --host 0.0.0.0 --port 8080 --jinja --reasoning off
</code></pre>
<p dir="auto"><strong>重點參數解讀（9/4 首日版）：</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>參數</th>
<th>值</th>
<th>原因</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>--parallel</code></td>
<td>1</td>
<td>48GB VRAM 只夠跑 1 個 131K ctx 實例</td>
</tr>
<tr>
<td><code>--cache-ram</code></td>
<td>32768</td>
<td>prompt cache 放 system RAM，長對話 98% hit rate</td>
</tr>
<tr>
<td><code>--spec-draft-n-max</code></td>
<td>2</td>
<td>實測 sweep 2 最優（37.4 pps / 0.63 接受率）</td>
</tr>
<tr>
<td><code>--cache-type-k/v</code></td>
<td>q8_0</td>
<td>KV cache 量化，慳 VRAM（<strong>9/7 更新：實測 f16 全精度反而更快，已改，見文末</strong>）</td>
</tr>
<tr>
<td><code>--flash-attn</code></td>
<td>on</td>
<td>必須開，唔開會慢好多</td>
</tr>
</tbody>
</table>
<hr />
<h2>實測數據</h2>
<h3>1. Prefill 速度（llama-bench）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>Prompt 長度</th>
<th>速度 (t/s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>1K</td>
<td>504</td>
</tr>
<tr>
<td>4K</td>
<td>460</td>
</tr>
<tr>
<td>8K</td>
<td>445</td>
</tr>
<tr>
<td>16K</td>
<td>439</td>
</tr>
<tr>
<td>32K</td>
<td>399</td>
</tr>
<tr>
<td>64K</td>
<td>331</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>解讀：</strong></p>
<ul>
<li>1K-16K 範圍內速度衰減唔多（~12%），日常使用最舒服</li>
<li>64K 要 ~200 秒先吐首字</li>
</ul>
<h3>2. Decode 速度</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>場景</th>
<th>速度 (t/s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>純 decode（無 MTP）</td>
<td>29</td>
</tr>
<tr>
<td>MTP n-max 2（agent 負載）</td>
<td>37.4</td>
</tr>
<tr>
<td>散文/創作類</td>
<td>27</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>MTP 接受率：</strong> 0.63（n-max 2）</p>
<p dir="auto"><strong>對比：</strong> 7900 XTX 篇文（<a href="http://lcz.me/topic/1164%EF%BC%89MTP" rel="nofollow ugc">lcz.me/topic/1164）MTP</a> 開 73.4 t/s（工具調用負載）。我部 W7800 之前測出 37.4 t/s，以為慢一倍 — 後來照 topic 1164 公開嘅測試題目原樣複製先發現真相（見下節）。</p>
<h3>3. 同題目對比實測（2026-09-04，最重要數據）</h3>
<p dir="auto">照 topic/1164 嘅 4 條工具調用題目 + 8 個 mock tools、n-max 5 原樣跑（每題 ×2）：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>情境</th>
<th>W7800 實測</th>
<th>7900 XTX 原文</th>
<th>比例</th>
</tr>
</thead>
<tbody>
<tr>
<td>單一工具呼叫</td>
<td>56.7 t/s</td>
<td>81.6 t/s</td>
<td>0.69</td>
</tr>
<tr>
<td>平行呼叫兩個工具</td>
<td>54.8 t/s</td>
<td>75.7 t/s</td>
<td>0.72</td>
</tr>
<tr>
<td>shell 任務</td>
<td>55.9 t/s</td>
<td>70.8 t/s</td>
<td>0.79</td>
</tr>
<tr>
<td>長參數工具</td>
<td>41.9 t/s</td>
<td>65.6 t/s</td>
<td>0.64</td>
</tr>
<tr>
<td><strong>平均</strong></td>
<td><strong>52.3 t/s</strong></td>
<td><strong>73.4 t/s</strong></td>
<td><strong>0.71</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>結論：52.3 / 73.4 = 71%，幾乎完全等於 CU 數比例 70/96 = 73%！</strong></p>
<ul>
<li>差距係<strong>純硬件</strong>（7900 XTX 多 26 個 CU），唔係 config 問題</li>
<li>之前測到 37.4 t/s 係因為用「agent JSON + markdown 混合」負載 — MTP 接受率得 0.63</li>
<li>純工具調用（格式固定）接受率上到 0.78-0.96，速度即刻上返 52 t/s</li>
</ul>
<h3>4. KV cache 全精度 + 大 Context 實測（48GB 夠唔夠？）</h3>
<p dir="auto"><strong>Qwen3.8-27B 混合架構發現（讀 GGUF metadata）：</strong> 65 層入面得 <strong>17 層係 full attention</strong>（idx 3, 7, 11 … 63, 64 — 每 4 層一層 + 最後層），其餘 48 層係 SSM（linear attention），<strong>完全唔食 KV cache</strong>。所以 KV cache 比其他 65 層全 attention 模型細約 4 倍 — hybrid 架構嘅最大好處。</p>
<p dir="auto">KV 計法：每 full-attn 層 = 4 kv-heads × 256 dim × 2 (K+V) = 2,048 elements/token；全 model = 34,816 elements/token（f16 = 68 KB/token）。</p>
<p dir="auto"><strong>實測 VRAM（全部真機 loaded：Q4_K_M 16.4GB + mmproj 0.9GB + compute buffer）：</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>配置</th>
<th>VRAM 用量 / 51.5GB</th>
<th>結果</th>
</tr>
</thead>
<tbody>
<tr>
<td>131K q8_0 KV（現行 production）</td>
<td>23.0 GB</td>
<td>—</td>
</tr>
<tr>
<td>160K f16（全精度）</td>
<td>29.4 GB（free 22）</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 冇 OOM</td>
</tr>
<tr>
<td>160K f32（最盡全精度）</td>
<td>40.7 GB（free 11）</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 冇 OOM</td>
</tr>
<tr>
<td>256K f16</td>
<td>36.9 GB（free 14.6）</td>
<td><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 冇 OOM</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>256K f16 照跑 topic 1164 題目（n-max 5，每題 ×2）：</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>情境</th>
<th>256K f16</th>
<th>131K q8_0</th>
<th>分別</th>
</tr>
</thead>
<tbody>
<tr>
<td>單一工具呼叫</td>
<td>59.0 t/s</td>
<td>56.7 t/s</td>
<td>+4%</td>
</tr>
<tr>
<td>平行呼叫兩工具</td>
<td>52.4 t/s</td>
<td>54.8 t/s</td>
<td>-4%</td>
</tr>
<tr>
<td>shell 任務</td>
<td>58.2 t/s</td>
<td>55.9 t/s</td>
<td>+4%</td>
</tr>
<tr>
<td>長參數工具</td>
<td>48.7 t/s</td>
<td>41.9 t/s</td>
<td><strong>+16%</strong></td>
</tr>
<tr>
<td><strong>平均</strong></td>
<td><strong>54.6 t/s</strong></td>
<td><strong>52.3 t/s</strong></td>
<td><strong>+4.4%</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>結論：</strong></p>
<ul>
<li>48GB 上 160K 全精度（f16）甚至 f32 都冇 OOM；256K f16 都得</li>
<li>f16 全精度 KV 比 q8_0 <strong>快</strong>（慳 dequant），大 context + 全精度係「免費升級」</li>
<li>推算 256K f32（~54GB）先會 OOM — 呢個先係 48GB 卡嘅物理上限</li>
</ul>
<h3>5. 模型精度升級實測（Q6_K / Q8_K_XL，2026-09-04）</h3>
<p dir="auto">48GB 仲有空間升模型精度，實測三隻 quant（全部 topic 1164 題目、n-max 5 純 MTP）：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>模型（weights）</th>
<th>Config</th>
<th>VRAM</th>
<th>平均 decode</th>
</tr>
</thead>
<tbody>
<tr>
<td>UD-Q4_K_M（16.5GB）</td>
<td>131K q8_0</td>
<td>23.0 GB</td>
<td>52.3 t/s</td>
</tr>
<tr>
<td>UD-Q4_K_M</td>
<td>256K f16</td>
<td>36.9 GB</td>
<td>54.6 t/s</td>
</tr>
<tr>
<td><strong>UD-Q6_K（22.0GB）</strong></td>
<td><strong>256K f16</strong></td>
<td><strong>42.2 GB</strong></td>
<td><strong>52.3 t/s</strong></td>
</tr>
<tr>
<td>UD-Q8_K_XL（31.5GB）</td>
<td>160K f16</td>
<td>44.3 GB</td>
<td>32.9 t/s</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>分析：</strong></p>
<ul>
<li><strong>Q6_K = 免費升級</strong>：質素明顯高過 Q4（6.5bpw vs 4.85bpw），但速度完全一樣（52.3 t/s），256K context 都保得住</li>
<li><strong>Q8_K_XL 懲罰好大</strong>：-37% 速度（32.9 t/s）— 權重 31.5GB，decode 每 token 要讀晒全部權重，撞正記憶體頻寬天花板（864 GB/s ÷ 31.5GB ≈ 27 t/s，實測 32.9 已算好）</li>
<li>頻寬模型驗證：Q4 16.5GB → 864/16.5 ≈ 52 t/s（實測 52.3 ✓）</li>
<li><strong>結論：Q6_K @ 256K f16 = 甜點</strong>（42.2GB，free 9.3GB）；Q8_K_XL 只適合唔介意速度、要最高精度嘅場景</li>
</ul>
<h3>6. 實際使用體驗</h3>
<ul>
<li><strong>日常對話（8-16K ctx）</strong>：流暢，每秒 30-40 字</li>
<li><strong>長文處理（32K+）</strong>：首字要等 1-2 分鐘，但可以接受</li>
<li><strong>Vision（mmproj）</strong>：成功讀到相入面嘅品牌/價錢/型號</li>
<li><strong>多輪對話</strong>：98% cache hit，唔會重新 prefill</li>
</ul>
<hr />
<h2>踩過嘅坑（重要！）</h2>
<h3>坑 1：glslc 錯裝</h3>
<p dir="auto"><strong>症狀：</strong> cmake configure 失敗，報 <code>Could NOT find Vulkan (missing: glslc)</code><br />
<strong>解決：</strong> <code>sudo apt install glslc</code>（唔係 glslang-tools）</p>
<h3>坑 2：headless GPU runtime PM</h3>
<p dir="auto"><strong>症狀：</strong> RAM 突然爆（27/32GB + swap）、VRAM 顯示 0GB、decode 由 31 跌到 14 t/s<br />
<strong>根因：</strong> AMD GPU 冇插 monitor，runtime PM <code>auto</code> → idle 6 秒就 suspend → VRAM 內容搬去 system RAM<br />
<strong>解決：</strong></p>
<pre><code class="language-bash"># 立即修復
echo on &gt; /sys/class/drm/card0/device/power/control

# 持久化（systemd unit）
sudo nano /etc/systemd/system/gpu-pm-fix.service
</code></pre>
<p dir="auto"><strong><img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/26a0.png?v=2fb7360d8c6" class="not-responsive emoji emoji-android emoji--warning" style="height:23px;width:auto;vertical-align:middle" title="⚠" alt="⚠" />️ 先睇速度先好落結論：</strong> ReBAR 開咗之後 driver 會將部分 allocation 放 GTT，單憑 <code>mem_info_vram_used</code> 細 <strong>唔可以</strong>判斷 model 搬咗 RAM。正確判準：① 實測 decode t/s 有冇跌 ② <code>runtime_status</code> 係咪 <code>suspended</code> ③ host RAM 有冇爆。</p>
<h3>坑 4：cache-ram 唔夠 / RAM 爆</h3>
<p dir="auto"><strong>症狀：</strong> log 出現 <code>making room for prompt cache entry</code>；system RAM 長期 86%<br />
<strong>影響：</strong> 長對話每輪都要重新 prefill，體感慢<br />
<strong>解決：</strong> <code>--cache-ram</code>（食 system RAM，唔係 VRAM）<br />
<strong>[9/7 更新] 根因補充：</strong> <code>making room</code> 本身只係 cache pool 輪替（唔係 OOM），但 RAM 真係爆嘅元兇係 <code>--cache-idle-slots</code> 預設開 —— 每個 slot 嘅 KV state 都 copy 落 system RAM。一定要加 <code>--no-cache-idle-slots</code>，RSS 即刻由 21GB 返到 2-3GB，cache hit 唔受影響（詳見 9/7 更新第 4 節）</p>
<h3>坑 5：pkill -f 殺錯</h3>
<p dir="auto"><strong>症狀：</strong> SSH session 斷線、server 冇起到<br />
<strong>根因：</strong> <code>pkill -f llama-server</code> 會連自己條 SSH command line 一齊殺<br />
<strong>解決：</strong> 用 <code>pkill -x llama-server</code>（exact process name match）</p>
<hr />
<h2>我嘅建議</h2>
<ol>
<li><strong>唔好直接買新卡</strong> —— 先睇下有冇現成 GPU 可以跑</li>
<li><strong>48GB 係甜點</strong> —— 24GB 唔夠（131K ctx 會爆），96GB 太貴</li>
<li><strong>[9/7 更新] Ubuntu + ROCm 7.2.4 係最穩</strong> —— 原文寫「ROCm 支援未完善」係舊版經驗；A/B 實測 ROCm prefill 快 38%、MTP n=8 做到 65.1 t/s（Vulkan n=8 跌到 33.7），而家生產已轉 ROCm，Vulkan 留返做備份</li>
<li><strong>一定開 Resizable BAR</strong> —— 呢個係 80% 嘅速度</li>
<li><strong>cache-ram 唔好慳</strong> —— 長對話體感差異好大；<strong>[9/7 更新] 同時一定要加 <code>--no-cache-idle-slots</code></strong>，不然 system RAM 會食到 86%（見文末）</li>
<li><strong>headless 機一定要做 GPU PM fix</strong> —— 否則會中 runtime suspend 嘅坑</li>
<li><strong>[9/7 新增] 結論唔好照搬，自己機實測</strong> —— 7900 XTX 嘅 KV 量化結論喺 W7800 完全相反；SGLang 官方「支援 Radeon」但實測用唔到</li>
</ol>
<hr />
<h2>9/7 更新：兩日實戰回饋（09-05 09:36 → 09-07 11:00，HKT）</h2>
<p dir="auto">發文之後我冇停手。我裝咗一個小小「數據記錄員」（每 2 秒睇一次 server log，自動記低每個請求嘅速度、tokens、MTP 接受率），又加咗一個溫度監察（每 30 秒讀一次 GPU 溫度、風扇、功率）。48 小時不停跑完，一共 <strong>1,916 個請求、約 8,540 萬 tokens 輸出</strong>，我將所有舊結論重新實測過。以下係結果。</p>
<h3>1. ROCm 原來快過 Vulkan（推翻原文結論）</h3>
<p dir="auto">我原本跟住社區主流「Vulkan 單卡全面贏」，對 ROCm 有少少懷疑。9/4 夜至 9/5 朝做咗完整 A/B（同一 commit、同一模型、同一參數，只換 backend）：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>測試</th>
<th>Vulkan</th>
<th>ROCm 7.2.4</th>
</tr>
</thead>
<tbody>
<tr>
<td>Prefill 512 tokens</td>
<td>503 t/s</td>
<td><strong>693 t/s（+38%）</strong></td>
</tr>
<tr>
<td>Prefill 4096 tokens</td>
<td>482 t/s</td>
<td><strong>668 t/s（+39%）</strong></td>
</tr>
<tr>
<td>Decode 128 tokens（無 MTP）</td>
<td>30.6 t/s</td>
<td>30.7 t/s（平手）</td>
</tr>
<tr>
<td>MTP n=8 工具調用</td>
<td>33.7 t/s（大跌）</td>
<td><strong>65.1 t/s</strong></td>
</tr>
<tr>
<td>MTP n=2 工具調用</td>
<td><strong>54.3 t/s</strong></td>
<td>50.2 t/s</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>結論：W7800（gfx1100）+ 新版 llama.cpp + ROCm 7.2.4，prefill 快 38%、decode 平手、MTP 長 draft 快兩倍。</strong> 「ROCm 支援未完善」係舊版本嘅經驗，在此更正。而家生產已轉 ROCm，Vulkan 留返做備份。</p>
<h3>2. KV cache：全精度 f16 最快（同 7900 XTX 相反）</h3>
<p dir="auto">7900 XTX 篇文（topic 1164）實測 KV 量化幾乎唔慢。W7800 實測<strong>完全相反</strong>（ROCm、n=2、工具調用負載）：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>KV 精度</th>
<th>速度</th>
<th>VRAM 慳</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>f16</strong></td>
<td><strong>50.9 t/s</strong></td>
<td>基準</td>
</tr>
<tr>
<td>q8_0</td>
<td>49.3 t/s（-3%）</td>
<td>4.2 GB</td>
</tr>
<tr>
<td>q4_0</td>
<td>48.1 t/s（-5.5%）</td>
<td>6.7 GB</td>
</tr>
<tr>
<td>q5_0/q4_1</td>
<td>38.7 t/s（-24%！）</td>
<td>7 GB</td>
</tr>
</tbody>
</table>
<p dir="auto">精度越高越快、線性下跌，q5_0 更暴跌 24%。所以原文嘅 q8_0 已改 f16 —— 呢個係「唔好照搬結論、要自己機實測」嘅第二個例證。</p>
<h3>3. 質素升級：Q8 唔係一定好（頻寬係王）</h3>
<p dir="auto">我做咗三級質素梯級測試（同 45K ctx、128 tokens、每級 20 次）：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>量化</th>
<th>檔案大小</th>
<th>速度</th>
<th>VRAM</th>
</tr>
</thead>
<tbody>
<tr>
<td>Q4_K_M</td>
<td>15.3 GiB</td>
<td>38 t/s</td>
<td>25.5 GiB</td>
</tr>
<tr>
<td>Q6_K</td>
<td>22.0 GiB</td>
<td>36 t/s</td>
<td>39.3 GiB（2 slot）</td>
</tr>
<tr>
<td>Q8_K_XL</td>
<td>31.5 GiB</td>
<td>34 t/s</td>
<td>41.1 GiB</td>
</tr>
</tbody>
</table>
<p dir="auto">每升一級質素，速度跌 ~5% —— 因為 decode 係頻寬競爭，檔案越大每 step 要讀越多資料。<strong>最終生產我揀 Q6_K + 2 個平行 slot（各 128K）</strong>：質素同 Q8 相差唔遠，但同時可以應付兩個對話，單流 ~36 t/s、雙流同時各 ~24-25 t/s。如果你得一個人用，Q8 單 slot（34 t/s）都係好選擇。</p>
<h3>4. RAM 86% 謎團揭開：<code>--no-cache-idle-slots</code></h3>
<p dir="auto">原文講「cache-ram 唔好慳」，但冇寫我踩咗一個坑：30GB RAM 長期食到 86%。9/6 查到元兇：llama.cpp 嘅 <code>--cache-idle-slots</code> <strong>預設係開</strong>，每個 slot 嘅 KV state 都會被 copy 落 system RAM，成個 cache pool 食滿 21GB RSS、仲行 swap。加咗 <code>--no-cache-idle-slots</code> 之後 RSS 穩定返 2-3GB，cache hit 完全唔受影響。<strong>如果你用 cache-ram，一定同時加呢個 flag。</strong></p>
<h3>5. SGLang 實測：「列咗名但未 ready」</h3>
<p dir="auto">9/5 我亦試咗 SGLang（AMD 官方話支援 Radeon）。結果：</p>
<ul>
<li>官方 FP8 量化：server 起得到，但 decode 只有 <strong>0.62 t/s</strong>（比 llama.cpp 慢 ~100 倍）—— RDNA3 冇原生 FP8 core，要模擬，GPU 100% busy 都係每 token 1.6 秒</li>
<li>三條 INT4 路線：全部失敗（loader 唔認 / SSM kernel dtype bug）</li>
<li>最新 python-native 版本：起得到、tool call 全部正確，但 decode 只有 4.9 t/s</li>
<li>對照組 Qwen3-8B dense：30.2 t/s 完全正常 —— 證明問題唔係 SGLang 本身，係 27B hybrid 架構 + FP8/INT4 支援不足</li>
</ul>
<p dir="auto">查咗 AMD 官方 Day-0 文章，支援對象係 <strong>Instinct 數據中心卡（MI300X 等）</strong>，Radeon 只係「列咗名」。<strong>W7800 繼續用 llama.cpp ROCm；想 AMD 行 SGLang，正路係 Instinct。</strong></p>
<h3>6. 兩日實戰數據（09-05 09:36 → 09-07 11:00）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指標</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>總請求</td>
<td>1,916</td>
</tr>
<tr>
<td>總輸出</td>
<td>~8,540 萬 tokens</td>
</tr>
<tr>
<td>Decode 速度</td>
<td>中位 33.3 t/s（p10 22.6 / p90 54.3 / max 120）</td>
</tr>
<tr>
<td>MTP 接受率</td>
<td>中位 0.70（純工具調用可上 0.91）</td>
</tr>
<tr>
<td>最長單次輸出</td>
<td>77,000 tokens（9/7 10:45，連生成 5,652 tokens、27.9 t/s 無跌速）</td>
</tr>
<tr>
<td>崩潰 / 致命錯誤</td>
<td><strong>0</strong></td>
</tr>
<tr>
<td>溫度</td>
<td>高峰 junction 101°C（全係 busy 時段）、風扇 max 2,571 rpm、功率 max 230W</td>
</tr>
<tr>
<td>Prompt cache 重用</td>
<td><strong>91% 請求命中 cache</strong>（f_sim &gt; 0.9），長對話幾乎唔使重新讀 prompt</td>
</tr>
</tbody>
</table>
<p dir="auto">現役配置（9/6 17:15 起）已<strong>連續運行 19 個半小時無重啟</strong>：VRAM 40/48 GiB、RAM 15/30 GB、GPU busy 時 junction ~93°C。</p>
<p dir="auto"><strong>最終生產配置（9/7 版）：</strong></p>
<pre><code class="language-bash">exec ~/llama.cpp-rocm/build-rocm/bin/llama-server \
  -m ~/models/Qwen3.8-27B-UD-Q6_K.gguf \
  --mmproj ~/models/mmproj-F16.gguf --image-min-tokens 1024 \
  --alias qwen3.8-27b --device ROCm0 --fit off -ngl -1 --no-mmap \
  --spec-type draft-mtp --spec-draft-n-max 2 \
  -c 262144 -ub 512 \
  --cache-type-k f16 --cache-type-v f16 \
  --parallel 2 --cache-ram 15360 --no-cache-idle-slots \
  --flash-attn on -t 12 -tb 6 \
  --host 0.0.0.0 --port 8080 --jinja --reasoning off
</code></pre>
<p dir="auto">（2 個 slot 各 128K、f16 KV、MTP n=2、cache pool 15G、<code>--no-cache-idle-slots</code> 防止 RAM 爆；systemd 自動重啟）</p>
<h3>7. 呢兩日學到嘅嘢</h3>
<ol>
<li><strong>結論唔好照搬</strong> —— KV 量化嘅結果 7900 XTX 同 W7800 完全相反；我信咗「Vulkan 全面贏」都係喺自己機上推翻</li>
<li><strong>先實測先結論</strong> —— SGLang 紙面「支援」，實測用唔到；ROCm 紙面「未完善」，實測快 38%</li>
<li><strong>要記錄數據</strong> —— 冇呢兩日嘅記錄，我唔知自己跑咗幾多、101°C 高峰係咪常態</li>
<li><strong>頻寬係王</strong> —— 27B decode 係純頻寬競爭，質素同速度係一條 spectrum，搵自己嘅平衡點</li>
</ol>
<hr />
<h2>最後</h2>
<p dir="auto"><strong>祝你好運。</strong></p>
<hr />
]]></description><link>https://lcz.me/topic/1538</link><generator>RSS for Node</generator><lastBuildDate>Mon, 07 Sep 2026 22:18:19 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1538.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 07 Sep 2026 07:37:04 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 12:45:04 GMT]]></title><description><![CDATA[<p dir="auto">48GB 大顯存 , 稀有珍品</p>
]]></description><link>https://lcz.me/post/16450</link><guid isPermaLink="true">https://lcz.me/post/16450</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Mon, 07 Sep 2026 12:45:04 GMT</pubDate></item><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 12:44:16 GMT]]></title><description><![CDATA[<p dir="auto">我係你全部弱化版..<br />
DSH 係MAC PRO 2017<br />
行7900XTX 香港8千</p>
<p dir="auto">都係睇佢講,所以吸引左行LOCAL QWEN 27B</p>
<p dir="auto">估唔到香港都有人一齊受影響</p>
<p dir="auto">不過講真, 我覺得本地都是VRAM 愈多愈好<br />
7900xtx 夠快, 短問答是不錯<br />
但長時間CODING agnet , 精度跌就好難用, 其本上行幾輪上下文就會影響到</p>
<p dir="auto">我而家試梗用unsloth 係7900XTX 下,盡量用更少RAM<br />
不過真係多RAM 就好多野做到...可惜~  BTW,邊到仲有得買?</p>
<p dir="auto">@小兒子,幫我轉做簡體中文</p>
]]></description><link>https://lcz.me/post/16449</link><guid isPermaLink="true">https://lcz.me/post/16449</guid><dc:creator><![CDATA[hoyin258]]></dc:creator><pubDate>Mon, 07 Sep 2026 12:44:16 GMT</pubDate></item><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 12:27:33 GMT]]></title><description><![CDATA[<p dir="auto">这么看的话7800  还是弱了点相比7900 的话。</p>
]]></description><link>https://lcz.me/post/16446</link><guid isPermaLink="true">https://lcz.me/post/16446</guid><dc:creator><![CDATA[张光璞]]></dc:creator><pubDate>Mon, 07 Sep 2026 12:27:33 GMT</pubDate></item><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 12:04:52 GMT]]></title><description><![CDATA[<p dir="auto">非常好的评测，W7800看起来还不错，你可以尝试下SG-Lang，论坛有帖子，体验会好很多。</p>
]]></description><link>https://lcz.me/post/16433</link><guid isPermaLink="true">https://lcz.me/post/16433</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Mon, 07 Sep 2026 12:04:52 GMT</pubDate></item><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 08:16:52 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kirin-h." aria-label="Profile: Kirin-H.">@<bdi>Kirin-H.</bdi></a></p>
<p dir="auto"><a href="https://www.price.com.hk/product.php?p=622502&amp;srsltid=AfmBOopm4JPC3wdPSQoM8-n9Xhq6TDSL3filqtBGDzv7V7wfhYudBvla" rel="nofollow ugc">https://www.price.com.hk/product.php?p=622502&amp;srsltid=AfmBOopm4JPC3wdPSQoM8-n9Xhq6TDSL3filqtBGDzv7V7wfhYudBvla</a></p>
<p dir="auto">W7900 48G 香港叫價$37800HKD左右，香港物價確實是非常貴</p>
]]></description><link>https://lcz.me/post/16388</link><guid isPermaLink="true">https://lcz.me/post/16388</guid><dc:creator><![CDATA[Wing Wah Law]]></dc:creator><pubDate>Mon, 07 Sep 2026 08:16:52 GMT</pubDate></item><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 08:13:03 GMT]]></title><description><![CDATA[<p dir="auto">香港的卖这么贵啊还以为香港电子产品会便宜点，大陆这边全新的W7900才21999人民币</p>
]]></description><link>https://lcz.me/post/16385</link><guid isPermaLink="true">https://lcz.me/post/16385</guid><dc:creator><![CDATA[Kirin H.]]></dc:creator><pubDate>Mon, 07 Sep 2026 08:13:03 GMT</pubDate></item><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 08:10:57 GMT]]></title><description><![CDATA[<p dir="auto">我覺得不吵，待機没聲音</p>
<p dir="auto">電腦滿載時，對比小米即熱飲水機在冷凍水時，小米更吵</p>
]]></description><link>https://lcz.me/post/16384</link><guid isPermaLink="true">https://lcz.me/post/16384</guid><dc:creator><![CDATA[Wing Wah Law]]></dc:creator><pubDate>Mon, 07 Sep 2026 08:10:57 GMT</pubDate></item><item><title><![CDATA[Reply to Gigabyte Radeon Pro W7800 AI TOP 48GB + Qwen3.8-27B 本地部署實測報告 on Mon, 07 Sep 2026 08:04:45 GMT]]></title><description><![CDATA[<p dir="auto">评测的很全面细致，这卡待机和满载的时候噪音分贝表现如何，吵不吵？</p>
]]></description><link>https://lcz.me/post/16381</link><guid isPermaLink="true">https://lcz.me/post/16381</guid><dc:creator><![CDATA[liu big]]></dc:creator><pubDate>Mon, 07 Sep 2026 08:04:45 GMT</pubDate></item></channel></rss>