<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[（15／9第二次更新）為 gfx1100 用戶追回軟件算力：GIGABYTE W7800 48G 單卡跑通 AMD 原版 Quark INT4，DSH 真實負載七日 +22%（50.9 → 62.11 t/s）]]></title><description><![CDATA[<blockquote>
<p dir="auto"><strong>測試平台</strong>：GIGABYTE Radeon<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2122.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--tm" style="height:23px;width:auto;vertical-align:middle" title="™" alt="™" /> PRO W7800 AI TOP 48G（<strong>單卡</strong>）<br />
gfx1100 / RDNA3 ｜ <strong>70 CU</strong> ｜ ROCm 7.2.4 ｜ Ubuntu 24.04 ｜ Ryzen 5 7500F ｜ 30 GB RAM<br />
2026-09-15</p>
</blockquote>
<hr />
<h2>〇、實測成績：DSH 真實 agent 路徑</h2>
<p dir="auto"><strong>本節所有數字均直接抽取自引擎日誌，並非實驗室短提示評測。</strong></p>
<h3>真實 agent 工作負載（DSH，上下文約 20K）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指標</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>decode 吞吐</strong></td>
<td><strong>62.11 t/s</strong></td>
</tr>
<tr>
<td>樣本數</td>
<td>11（59.55 / 62.11 / 63.94 / 59.49 / 52.37 / 59.95 / 67.80 / 78.91 / 71.91 / 76.56，另有 1 個 9.35 離群值）</td>
</tr>
<tr>
<td>accept length</td>
<td>2.87 – 3.70</td>
</tr>
<tr>
<td>同一台機器、短提示單流</td>
<td><strong>85.05 t/s</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>⇒ 兩者唯一的變數是上下文深度</strong>（full token 20,004 – 20,944）。<strong>這就是 agent 場景的真實成本。</strong></p>
<h3>複驗（2026-09-15，HiCache 8 ＋ chunk 8192）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指標</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>decode 吞吐</td>
<td><strong>61.5 t/s</strong>（n=23，full token 中位數 <strong>24,299</strong>）</td>
</tr>
<tr>
<td>p90</td>
<td>80.4 t/s</td>
</tr>
<tr>
<td>accept length 中位數</td>
<td>3.40</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>⇒ 與 09-14 的 62.11 t/s 一致，並無退步。</strong>（用戶主觀感受的 74 t/s 屬較短上下文。）</p>
<h3>與我們上一份帖文的對比（同一張卡、同一工作負載）</h3>
<p dir="auto">我們 2026-09-08 的上一份帖文（<a href="http://lcz.me/topic/1559%EF%BC%89%E5%9C%A8%E5%90%8C%E4%B8%80%E5%BC%B5" rel="nofollow ugc">lcz.me/topic/1559）在同一張</a> W7800 上報告：DSH 真實 agent 工作負載單發 <strong>50.9 t/s</strong>（gen mean；median 53.9、max 79.5；KV 範圍 11–21K）。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th></th>
<th>2026-09-08（topic 1559）</th>
<th>2026-09-15（本文）</th>
<th>變化</th>
</tr>
</thead>
<tbody>
<tr>
<td>DSH agent 路徑 decode（單發）</td>
<td><strong>50.9 t/s</strong></td>
<td><strong>62.11 t/s</strong></td>
<td><strong>+22.0%</strong></td>
</tr>
<tr>
<td>上下文範圍</td>
<td>11–21K</td>
<td>20,004–20,944</td>
<td>同級</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>⇒ 七日之內 +22%。而這不是單一變因的功勞</strong> —— 期間換了模型配方（W4A16-AutoRound-GPTQ → INT4-GPTQ-v3）、加了自己寫的 INT4 lm_head shim（<strong>+19.7%</strong>）、改了兩處 kernel（合計再 <strong>+6.5%</strong>），並把整套配置重搭回同一條線上（見第二節）。</p>
<p dir="auto"><strong>⇒ 更重要的是：前面仍有大把路未走。</strong> 上游 vLLM 已改用我們尚未採用的兩條新路線、確定性開關從未打開、kernel 內部那道 498-cycle 的串行鏈尚未拆解。詳見第六節。</p>
<hr />
<h3>生產環境實績</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>項</th>
<th>值</th>
</tr>
</thead>
<tbody>
<tr>
<td>decode 中位數（投產時）</td>
<td>57–60 t/s</td>
</tr>
<tr>
<td>KV cache 命中率</td>
<td>65 – 77%</td>
</tr>
<tr>
<td><strong>曾經承擔</strong></td>
<td><strong>11 個 subagent 並行審計</strong></td>
</tr>
<tr>
<td>split-K 512 與 1024 之 A/B</td>
<td><strong>70 對 61 t/s</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>這些數字在任何 gfx1100 單卡用戶的環境中皆可復現。</strong></p>
<hr />
<h2>一、我們是誰、在什麼卡上做</h2>
<p dir="auto">我們只有<strong>一張</strong> GIGABYTE Radeon<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2122.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--tm" style="height:23px;width:auto;vertical-align:middle" title="™" alt="™" /> PRO W7800 AI TOP 48G（gfx1100、<strong>70 CU</strong>）。</p>
<p dir="auto"><strong>單卡這一事實至關重要</strong>，因為它決定了我們所有結論的邊界：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th></th>
<th>我們</th>
<th>論壇其他前輩</th>
</tr>
</thead>
<tbody>
<tr>
<td>卡</td>
<td><strong>W7800 48G ×1（70 CU）</strong></td>
<td>7900 XTX ×2（各 96 CU）／4090 48G 改裝版 ×1／4080S ×2</td>
</tr>
<tr>
<td>總 VRAM</td>
<td>48 GB</td>
<td>48 GB（雙 7900XTX 24+24）／48 GB</td>
</tr>
<tr>
<td>TP</td>
<td>1</td>
<td>2（雙卡）</td>
</tr>
<tr>
<td>卡片架構</td>
<td><strong>RDNA3 gfx1100</strong></td>
<td>7900XTX 同為 gfx1100 但 <strong>96 CU</strong>；4090 為 NVIDIA</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>⇒ 本文出現的任何數字，其前提皆為「單卡、70 CU、gfx1100」。</strong></p>
<p dir="auto"><strong>W7900（96 CU）、7900 XTX（96 CU）與 W7800（70 CU）雖然同屬 gfx1100，但 CU 數不同，吞吐量不可直接套用。</strong> 我們先前見到有材料把 CU 數寫成 96 —— 那是 W7900 的數字，我們這張是 <strong>70</strong>（amd-smi 實測）。</p>
<p dir="auto">我們的目標很簡單：<strong>把 gfx1100 用戶所能運用的軟件算力，逐分逐毫地榨取出來，而且每一步都可復現。</strong></p>
<hr />
<h2>二、主軸線：這條路的形成過程</h2>
<h3>起點：一句提示把我們推了進來</h3>
<p dir="auto">2026-09-05 我們試用 SGLang 官方版，結論為「已列名，但尚未就緒」（FP8 只有 0.62 t/s）。正想放棄的時候，<strong>Terry</strong> 在我們的報告下面留下一句：</p>
<blockquote>
<p dir="auto">「你可以嘗試下 SGLang，論壇有帖子，體驗會好很多」</p>
</blockquote>
<p dir="auto">一句提示，我們投入了十天。沿著他指出的方向，我們找到論壇前人已經鋪好的路（見文末致謝），最後以社群的 gfx1100-support fork 完成三引擎完整對比，才定下生產配置。</p>
<p dir="auto"><strong>⇒ 這是本文最想說明的一件事：這條路是社群鋪出來的，我們只是接續前行。</strong></p>
<h3>中段：找到真正的病灶</h3>
<p dir="auto">我們做過許多無效的優化，代價最高的一課是：<strong>有一整輪優化耗在一個穩態佔比 0% 的 kernel 上</strong> —— 那個檔案佔 46% 的代碼量，卻只佔 4.59% 的 GPU 時間；真正的熱核佔 61.14%。</p>
<blockquote>
<p dir="auto"><strong>質量集中處 ≠ 能量集中處。</strong></p>
</blockquote>
<p dir="auto">這句話後來成為我們所有優化的第一道關卡：<strong>動手之前先量能量分佈。</strong></p>
<h3>轉折：接入 AMD 原版模型</h3>
<p dir="auto">AMD 在 HuggingFace 放出 <code>amd/Qwen3.8-27B-Quark-AWQ-INT4-W4A16</code>，但它在 RDNA3 上<strong>無法運行</strong> —— SGLang 的 Quark 模組只實作了三種 scheme，純 W4A16 沒有；<code>quark/weights.py</code> 甚至直接註明「Scheme didn't allocate the parameter (e.g. W4A16)」。</p>
<p dir="auto">我們撰寫了轉換器把它接入。<strong>這是我們第一件原創作品</strong>（詳見第三節）。</p>
<h3>關鍵一段：INT4 之後，速度是怎樣追回來的</h3>
<p dir="auto">接通 AMD 原版模型之後，我們面對第二個問題：<strong>模型能跑，但速度不夠。</strong></p>
<p dir="auto">我們寫了一個 Python shim（<code>rdna_shim/rdna_lmhead_int8.py</code>），經 <code>PYTHONPATH</code> 注入，在 vendored 樹只加三行攔截 <code>_compute_lm_head</code>。它不改 kernel，而是<strong>換路線</strong>：lm_head 是 248320×5120 的 bf16 矩陣（2.543 GB），每步被讀四次（verify 一次 M=4、draft 三次 M=1），合共約 13.9 ms/步、佔步時約 20%；M=1 時完全是記憶體綁死，位元寬直接決定成本。改用 INT4 之後每次只讀 <strong>0.661 GB</strong>。</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>階段</th>
<th>引擎吞吐（5K 上下文、temp 0）</th>
<th>相對基準</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16 lm_head（原狀）</td>
<td>57.12 t/s</td>
<td>—</td>
</tr>
<tr>
<td>int8 lm_head（中間步驟）</td>
<td>60.43 t/s</td>
<td>+5.8%</td>
</tr>
<tr>
<td><strong>int4 lm_head（最終）</strong></td>
<td><strong>68.38 t/s</strong></td>
<td><strong>+19.7%</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>輸出 MD5 三者完全相同，accept length 4.0 不變，U+FFFD 為 0。</strong> 兩次獨立量測得 68.41 / 68.38。</p>
<p dir="auto">在此之上，我們又改了 kernel 本體兩處：WMMA 門檻由 M≥16 下移至 M≥9 &amp;&amp; N≥2048（M=9 GEMM 246.7 → 185.9 µs，<strong>−24.6%</strong>），以及 B2 編譯期模板化（M=4 −9.3%、M=5 −8.7%、M=8 −14.2%）。端到端 <strong>68.38 → 72.84 t/s，累計對 bf16 基準 +27.5%</strong>，輸出與基準逐字節相同。</p>
<p dir="auto"><strong>速度追回來之後，我們再把底盤逐件接回：</strong> MTP-3 投機解碼（09-11 由 DFLASH 折返後重新接上）、CUDA graph、HiCache（命中率 65% → 77%，09-14 重新上線），最後是 09-14 傍晚的 DFLASH8 ＋ HiCache8（單流 92.57 t/s）。</p>
<p dir="auto"><strong>⇒ 這一段的代價，同樣如實記錄：</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>時間</th>
<th>失敗與代價</th>
</tr>
</thead>
<tbody>
<tr>
<td>09-11 13:20</td>
<td>投機門檻下移 M≥2 雖然修復了確定性，代價是吞吐 −15.2%（88.9 t/s）。<strong>交易為負，已放棄並還原。</strong></td>
</tr>
<tr>
<td>09-11 13:20</td>
<td>draft 量化 v3：體積由 3.85 GB 降到 1.743 GB，但 accept rate 掉到 <strong>0.00</strong>、吞吐 15.7 t/s（−85%）</td>
</tr>
<tr>
<td>09-11 14:14</td>
<td><strong>事故</strong>：INT4 一度被洗成 W4A16 —— <code>exp_config.py</code> 的 REF 寫死了 W4A16 配方，13:31 起所有受控測試全部跑錯模型</td>
</tr>
<tr>
<td>09-11 22:36</td>
<td>差一點誤報成功：見到 HYBRID 啟動與 84.9 t/s，真相是 HYBRID 失敗後系統 fallback 到 DFLASH；靠 <code>/proc/&lt;pid&gt;/cmdline</code> 硬驗證才攔截下來</td>
</tr>
<tr>
<td>09-11 22:36</td>
<td>字串拼接改配置令反斜線續行斷裂，<strong>生產停擺 3.5 分鐘</strong></td>
</tr>
<tr>
<td>09-13</td>
<td>P29–P52 工藝戰役：<strong>18 個假說全部否證</strong>（全部落在 ±0~2%），隨後證明整場戰役打在那個只佔 4.59% GPU 時間的 kernel 上</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>⇒ 這就是「質量集中處 ≠ 能量集中處」的代價，也是我們把它列為第一道關卡的原因。</strong></p>
<hr />
<h3>現在：把差距交代清楚</h3>
<p dir="auto">接入之後，我們量到自身的服務品質<strong>低於 AMD 公布的數字</strong>（PPL 高一成、GSM8K 低七分）。我們排除了六個嫌疑，但<strong>機制仍未定位</strong> —— 此點我們在第四節誠實交代。</p>
<hr />
<h2>三、我們原創的三項成果</h2>
<h3>3.1 Quark W4A16 → SGLang GPTQ 轉換器（已開源）</h3>
<blockquote>
<p dir="auto"><strong><a href="https://github.com/lawsirlawsir-png/rdna3-quark-w4a16" rel="nofollow ugc">https://github.com/lawsirlawsir-png/rdna3-quark-w4a16</a></strong></p>
</blockquote>
<p dir="auto"><strong>原意</strong>：讓 gfx1100 用戶<strong>不必等待 SGLang 補上 W4A16 scheme</strong>，當日即可運行 AMD 原版模型。</p>
<p dir="auto"><strong>實質作用</strong>：把 Quark 的 <strong>N-packed <code>[K, N/8]</code></strong> 佈局重新打包成 GPTQ 的 <strong>K-packed <code>[K/8, N]</code></strong>，權重<strong>未改動任何一個位元</strong>。</p>
<p dir="auto"><strong>技術核心（亦是我們耗時最久之處）</strong>：Quark 在 <code>pack_method="reorder"</code> 之下，每個 int32 內的 8 個 4-bit 值<strong>並非順序排列</strong>，而是按 <code>order_map = [0,2,4,6,1,3,5,7]</code> 重排。忽略這一步，整個模型將輸出亂碼。</p>
<p dir="auto"><strong>我們當初正是遺漏了這一步，以數值掃描猜測數小時之久；最終閱讀官方 <code>quark/torch/utils/pack.py</code> 之後即得解。</strong></p>
<p dir="auto"><strong>⇒ 教訓：凡涉及外部格式，先讀官方實作，不可依靠試探。此後成為我們的鐵律。</strong></p>
<p dir="auto"><strong>無損驗證（三條獨立證據）</strong>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>檢驗</th>
<th>結果</th>
</tr>
</thead>
<tbody>
<tr>
<td>逐元素反量化比對（5 層，含 N=48 的極瘦層）</td>
<td><strong>max 絕對差 = 0.000000e+00</strong></td>
</tr>
<tr>
<td>全層覆蓋</td>
<td><strong>496 個量化層 ＋ 703 個普通層</strong></td>
</tr>
<tr>
<td>對 BF16 原模型的忠實度（96 個 <code>in_proj</code>）</td>
<td>cos <strong>最小 0.9871、中位數 0.9905</strong></td>
</tr>
</tbody>
</table>
<h3>3.2 N-aware K-split（小 M 專用，−7.9%）</h3>
<p dir="auto"><strong>Commit</strong>：<code>0d2e013 perf(rdna3-wmma): opt-in N-aware k_split for small-M; -7.9% on M=16 verify GEMM across all 6 production shapes, fp64-identical accuracy; default path unchanged</code></p>
<p dir="auto"><strong>原意</strong>：上游的 <code>compute_wmma_k_split()</code> <strong>只看 K、不看 N</strong>，因此在 K ∈ {5120, 6144, 17408} 時一律返回 4。但在我們的生產形狀中，<strong>N=5120 佔權重的 23.6%</strong>，在 k=4 之下嚴重佔用不足。</p>
<p dir="auto"><strong>實質作用</strong>：改為 <strong>N-aware</strong> 規則 —— 選取「令 block 數達到約 5120 的最小 k」。在我們六個生產形狀上逐一掃描後：</p>
<p dir="auto"><strong>六個生產形狀的逐 k 掃描</strong>（單位 ms，粗體為實測最優）</p>
<ul>
<li><strong>5120 × 17408</strong>：k=4 0.22503 ｜ k=8 <strong>0.22087</strong> ｜ k=16 0.22612 ｜ k=32 0.25149 ⇒ <strong>k=8</strong></li>
<li><strong>17408 × 5120</strong>：k=4 0.26621 ｜ k=8 0.26696 ｜ k=16 0.22434 ｜ k=32 <strong>0.22263</strong> ⇒ <strong>k=32</strong></li>
<li><strong>5120 × 10240</strong>：k=4 0.16138 ｜ k=8 <strong>0.13935</strong> ｜ k=16 0.14128 ｜ k=32 0.15573 ⇒ <strong>k=8</strong></li>
<li><strong>6144 × 5120</strong>：k=4 0.10095 ｜ k=8 0.10367 ｜ k=16 <strong>0.09191</strong> ｜ k=32 0.09835 ⇒ <strong>k=16</strong></li>
<li><strong>5120 × 6144</strong>：k=4 0.09661 ｜ k=8 0.09436 ｜ k=16 <strong>0.09122</strong> ｜ k=32 0.09996 ⇒ <strong>k=16</strong></li>
<li><strong>5120 × 12288</strong>：k=4 0.17250 ｜ k=8 <strong>0.16021</strong> ｜ k=16 0.16392 ｜ k=32 0.18090 ⇒ <strong>k=8</strong></li>
</ul>
<p dir="auto">規則在 <strong>6 個形狀中有 5 個命中實測最優</strong>；唯一例外為 17408×5120，差 0.77%（按真實層混合加權後更小）。</p>
<p dir="auto"><strong>最重要的一點</strong>：<strong>準確度在 fp64 之下逐位元相同</strong>，且<strong>預設路徑完全不變</strong>（須開啟 <code>SGL_WMMA_KSPLIT_AUTO=1</code> 方生效）。我們其後實測關閉它：<strong>PPL 9.7296 對開啟時 9.7297，差 −0.00%</strong> ⇒ 這個效能補丁在數值上是忠實的。</p>
<h3>3.3 其他原創改動（逐項交代原意與作用）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>Commit</th>
<th>原意</th>
<th>實質作用</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>c2d509e</code></td>
<td>小 M WMMA 門檻下移至 <strong>M&gt;=9</strong></td>
<td>覆蓋 decode verify step（生產實測 80% 為 M=4，其餘為 M=8）</td>
</tr>
<tr>
<td><code>9af1d8d</code></td>
<td>GPTQ kernel <strong>編譯期模板化</strong></td>
<td>Σa 預算與 z 修正改為<strong>每 group 一次</strong>，減少重複計算</td>
</tr>
<tr>
<td><code>b10a8ab</code></td>
<td><strong>正確性修正</strong></td>
<td>把不規則（NGRAM）樹排除於 mask-less unified verify kernel 之外</td>
</tr>
<tr>
<td><code>01d62d6</code></td>
<td>生產樹入版控</td>
<td>GPTQ M_COUNT 5/6/7、lm_head 低位寬</td>
</tr>
<tr>
<td><code>1f1ba80</code></td>
<td>WMMA M=16 瓶頸研究</td>
<td>env-gated 量測變體，<strong>預設路徑不變</strong></td>
</tr>
<tr>
<td><code>23dd307</code></td>
<td>建置整潔</td>
<td>把量測變體<strong>排除於生產 build 之外</strong></td>
</tr>
<tr>
<td><code>6c5409c</code></td>
<td>建置整潔</td>
<td>停止追蹤 <strong>21 個由 hipify 產生的 .hip 檔</strong></td>
</tr>
<tr>
<td><code>005ad71</code></td>
<td>修錯</td>
<td>CU 數 <strong>96 → 70</strong>（96 為 W7900；本卡為 70）</td>
</tr>
<tr>
<td><code>f736a74</code></td>
<td>首次入版控</td>
<td>2026-09-08/09 的 gfx1100 生產補丁</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>另外兩項不在 git 之內、但同樣屬於我們原創的</strong>：</p>
<ul>
<li><code>prod/rdna_shim/rdna_lmhead_int8.py</code> —— <strong>INT4 LM head</strong>，只在小 batch 觸發</li>
<li>我們整套<strong>量測方法論</strong>（見第五節）</li>
</ul>
<hr />
<h2>四、陷阱：六個已排除的嫌疑，與一個仍未解開的差距</h2>
<h3>4.1 差距本身</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指標</th>
<th>我們（單卡 W7800）</th>
<th>AMD 官方卡</th>
<th>差</th>
</tr>
</thead>
<tbody>
<tr>
<td>Wikitext word_perplexity</td>
<td><strong>9.7297</strong></td>
<td>8.8250</td>
<td><strong>+10.25%</strong></td>
</tr>
<tr>
<td>GSM8K flexible（兩次全量平均）</td>
<td><strong>83.70%</strong></td>
<td>91.51%</td>
<td><strong>−7.81 pt</strong></td>
</tr>
<tr>
<td>GSM8K strict（同上）</td>
<td><strong>84.12%</strong></td>
<td>90.37%</td>
<td><strong>−6.25 pt</strong></td>
</tr>
</tbody>
</table>
<h3>4.2 我們排除的（附方法）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>#</th>
<th>嫌疑</th>
<th>排除方法</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td><strong>權重量化</strong></td>
<td>逐元素反量化比對 <strong>max 絕對差 = 0</strong>（5 層）＋ 全層 496/496 ＋ 對 BF16 源 cos≥0.9871</td>
</tr>
<tr>
<td>2</td>
<td><strong>LM head</strong></td>
<td>PPL 以 echo 取 <strong>prefill</strong> logprobs，而 INT4 LM head 只在 <strong>M ≤ 8</strong> 觸發（日誌：<code>i4=1418 / miss=182</code>）；換 BF16 後 <strong>9.7297 → 9.7297 完全不變</strong></td>
</tr>
<tr>
<td>3</td>
<td><strong>量測儀器</strong></td>
<td>echo 路徑與原生 <code>/generate</code> 對同 6 篇文件得<strong>位元相同</strong>（差 0.000%）；token-PPL 5.52 × 1.32 = word-PPL 9.46 三數自洽</td>
</tr>
<tr>
<td>4</td>
<td><strong>資料集</strong></td>
<td>wiki-2 與 wiki-103 的 test split <strong>sha 相同</strong>（62/62）</td>
</tr>
<tr>
<td>5</td>
<td><strong>prompt 截斷</strong></td>
<td>超長僅 <strong>0.61%</strong>、缺答案格式僅 <strong>0.30%</strong></td>
</tr>
<tr>
<td>6</td>
<td><strong>我們自己的 K-split 優化</strong></td>
<td>關掉 → <strong>9.7296</strong>（−0.00%）</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>⇒ 差距真實存在，機制我們尚未定位。我們不作猜測。</strong></p>
<h3>4.3 更重要的發現：我們自身的 greedy 解碼並不確定</h3>
<p dir="auto">上表兩次全量評測<strong>皆用 greedy</strong>（<code>do_sample=False</code>），理應完全一致，實際卻相差 0.76 個百分點。</p>
<p dir="auto">於是我們在<strong>同一部伺服器、同一配置</strong>之下，同樣 150 題連續執行兩次：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指標</th>
<th>結果</th>
</tr>
</thead>
<tbody>
<tr>
<td>輸出完全相同</td>
<td><strong>107 / 150（71.3%）</strong></td>
</tr>
<tr>
<td>對錯翻轉</td>
<td><strong>3 題</strong></td>
</tr>
<tr>
<td>兩次分數</td>
<td><strong>85.33% 對 84.67%</strong></td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>⇒ 在我們的技術棧上，greedy 解碼是非確定的。</strong></p>
<p dir="auto"><strong>⇒ 因此我們不報告單一數字，而報告兩次全量評測的平均值，並標明自身噪音約 ±1 pt。</strong></p>
<p dir="auto"><strong>這也是本文的一項關鍵提醒：任何「更動一個參數即可令分數上升 0.5 pt」的結論，在此噪音水平之下皆不可信。</strong></p>
<p dir="auto">我們懷疑來源是 split-K 的 <strong>CAS 原子累加</strong>（加法次序不固定），或 LM head 低位寬路徑的 <strong>M ≤ 8 閘</strong>（精度隨 batch 大小跳變）。<strong>兩者皆未驗證。</strong></p>
<h3>4.4 幾項實用陷阱（可為他人節省時間）</h3>
<p dir="auto"><strong>① lm-eval 的 max_length 陷阱。</strong> API 後端有 <code>max_context_len = max_length − max_gen_toks</code>。用預設 <code>max_length=2048</code> 配 AMD 的 <code>max_gen_toks=8192</code> ⇒ <strong>−6144</strong> ⇒ prompt 截成空 ⇒ HTTP 400（我們觸發了 2645 次）。<strong>AMD 官方指令寫明 <code>max_length=16384</code>。</strong></p>
<p dir="auto"><strong>② LM head 的 M 閘。</strong> <code>M ≤ 8</code> 走 INT4、<code>M &gt; 8</code> 走 BF16 ⇒ <strong>同一 prompt 的 logits 會因並發批次的組成而異</strong>，屬可重現性風險。</p>
<p dir="auto"><strong>③ page-size 與 hybrid mamba。</strong> 論壇實測：<code>page-size 64</code> 配 hybrid mamba 加 HiCache <strong>本身即會段錯誤</strong>，必須 <code>page-size 1</code>（我們恰好為 1）。</p>
<p dir="auto"><strong>④ <code>--gpureset</code> 不可使用。</strong> 對 gfx1100 會鎖住 PCIe root port，須冷開機（社群警告）。</p>
<hr />
<h2>五、經驗：我們學到的若干原則</h2>
<p dir="auto"><strong>① 先對證，後落手。</strong> 凡涉及外部格式、協議、第三方語義，第一步必須尋求對證 —— 而且<strong>優先閱讀官方源碼，不讀文檔描述</strong>。</p>
<p dir="auto"><strong>② 驗證方法本身要先驗證。</strong> 必須證明它<strong>能夠識別已知的壞輸入</strong>，而非僅止於「能夠運行」。（我們的一個精度閘門曾經對 NaN、全零、非有限輸出<strong>三個通道全部靜默判 PASS</strong>。）</p>
<p dir="auto"><strong>③ 對證對象要驗身分。</strong> 凡對證，必印被對證對象的 <strong>sha256</strong>。<strong>修改時間、路徑、檔名，一概不可信。</strong></p>
<p dir="auto"><strong>④ 硬件規格只准實測。</strong> 同一張卡在四份材料中出現過 96 / 70 / 40 / 35 四個 CU 數，僅實測的 <strong>70</strong> 正確。</p>
<p dir="auto"><strong>⑤ 質量集中處 ≠ 能量集中處。</strong> 動手之前先量能量分佈。</p>
<p dir="auto"><strong>⑥ 配件無辜，組合有罪。</strong> 我們曾經把 HiCache 當作段錯誤主因；後來發現只有 <strong>DFLASH 與 HiCache 的組合</strong>才會出事，HiCache 本身無辜。</p>
<p dir="auto"><strong>⑦ 一個中心為忠，兩個中心為患。</strong> 量度之前先釐清有幾個中心；並行量測會互相污染，<strong>任何時候只准一個</strong>。</p>
<hr />
<h2>六、將來目標</h2>
<p dir="auto"><strong>① 量測上游 vLLM 的新路線。</strong> 上游已改為 <strong>M≤5 使用專用 INT4 skinny GEMM（<code>wvSplitK_int4_g</code>，wave 層分工 ＋ DPP 歸約、無原子競爭）、M&gt;5 使用 Triton</strong>。我們的 decode 有 80% 屬 M=4，恰好落在第一段。<strong>A/B 工具已編譯完成，僅待執行。</strong></p>
<p dir="auto"><strong>② 檢驗非確定性的機制。</strong> 同一個 A/B 實驗同時可驗：CAS 原子 對 DPP 歸約。</p>
<p dir="auto"><strong>③ 開啟 <code>--enable-deterministic-inference</code>。</strong> 我們生產環境中此開關一直為 0。其說明為「batch invariant ops」，恰對應我們「精度隨 batch 跳變」的假設。</p>
<p dir="auto"><strong>④ 將 kernel 層開源 —— <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2705.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--white_check_mark" style="height:23px;width:auto;vertical-align:middle" title="✅" alt="✅" /> 已於 2026-09-15 完成。</strong> 以 patch series 形態發佈 <strong>19 個 patch</strong>（43 檔、+4,097 / −14,870 行），基底 <code>StevenChenSE/sglang</code> 之 <code>gfx1100-support</code> @ <code>1442c18</code>。<strong>由 GitHub 下載後套用到乾淨檢出，可逐位元重現我們的生產樹。</strong> <a href="https://github.com/lawsirlawsir-png/rdna3-quark-w4a16" rel="nofollow ugc">https://github.com/lawsirlawsir-png/rdna3-quark-w4a16</a></p>
<p dir="auto"><strong>⑤ 繼續追查那個 +10.25% / −7.81 pt 的差距。</strong> 六個嫌疑已排除，機制尚未定位。</p>
<hr />
<h2>七、致謝 —— 我們的技術層有 99% 來自他人</h2>
<p dir="auto"><strong>我們所做的那 1% 之所以能夠成立，是因為前面已經有人把路鋪好。逐項歸位如下：</strong></p>
<h3>引我們走上這條路的人</h3>
<ul>
<li><strong>Terry</strong>（<a href="http://lcz.me" rel="nofollow ugc">lcz.me</a>）—— 「你可以嘗試下 SGLang，論壇有帖子，體驗會好很多」。<strong>一句提示，令我們投入了十天。</strong></li>
</ul>
<h3>我們的基礎：開源項目</h3>
<ul>
<li><strong><a class="plugin-mentions-user plugin-mentions-a" href="/user/stevenchense" aria-label="Profile: StevenChenSE">@<bdi>StevenChenSE</bdi></a></strong> —— <code>StevenChenSE/sglang</code> 的 <code>gfx1100-support</code> 分支。<strong>我們整條生產線建在這上面。</strong></li>
<li><strong>@vllm-project</strong> —— RDNA3 的 WMMA GPTQ kernel（我們拆解的那份 <code>q_gemm_rdna3_wmma.cu</code>，檔案頭即為 <code>Copyright contributors to the vLLM project</code>）</li>
<li><strong>@amd</strong> —— <strong>AMD Quark</strong> 量化工具鏈與 <code>pack.py</code> 的打包語義（沒有它，我們的轉換器無從對證）</li>
<li><strong>@ExLlama</strong> 社群 —— GPTQ kernel 的血統</li>
</ul>
<h3>論壇前人（我們沿其路線前行）</h3>
<ul>
<li><strong><a class="plugin-mentions-user plugin-mentions-a" href="/user/flyer666" aria-label="Profile: flyer666">@<bdi>flyer666</bdi></a></strong> —— <a href="http://lcz.me" rel="nofollow ugc">lcz.me</a> <strong>#1532</strong>：SGLang ＋ HiCache 三級架構、為 HiCache 增加 RAM；<strong>#1640</strong>：gfx1100 單卡分析</li>
<li><strong><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E6%8A%A1%E9%94%A4%E8%80%85" aria-label="Profile: 抡锤者">@<bdi>抡锤者</bdi></a> / franklee006</strong> —— <a href="http://lcz.me" rel="nofollow ugc">lcz.me</a> <strong>#1340</strong>（4090 48G ＋ DFlash2，DSH 寫碼均速 110 t/s）、<strong>#18288</strong>（雙 7900XTX 完整實戰）、<strong>#1329</strong>（HiCache 實測）、<strong>#1587</strong>（雙 4080S HiCache）</li>
<li><strong><a class="plugin-mentions-user plugin-mentions-a" href="/user/_%E6%8A%98%E9%A8%B0_" aria-label="Profile: _折騰_">@<bdi>_折騰_</bdi></a></strong> —— <a href="http://lcz.me" rel="nofollow ugc">lcz.me</a> <strong>#1500 / #1340</strong>：HiCache 多會話實戰、DFlash2 與 DSpark 的第一手主觀比較</li>
<li><strong>@Michael Zhou</strong> —— 4090D 48G 參數（經 #1329 轉錄）：<code>--mamba-full-memory-ratio 1.0</code>、<code>--mamba-scheduler-strategy extra_buffer</code>、<code>--mamba-track-interval 2048</code>、<strong>page-size 1</strong></li>
<li><strong><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E6%B8%85%E9%A3%8E%E6%98%8E%E6%9C%88" aria-label="Profile: 清风明月">@<bdi>清风明月</bdi></a></strong> —— RTX PRO 4500 32G NVFP4 ＋ HiCache 參數（TID:1341）</li>
<li><strong><a class="plugin-mentions-user plugin-mentions-a" href="/user/neo" aria-label="Profile: Neo">@<bdi>Neo</bdi></a></strong>（<a href="http://lcz.me" rel="nofollow ugc">lcz.me</a>）—— HiCache L2/L3 原始帖</li>
<li><strong><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: xiaote">@<bdi>xiaote</bdi></a> 老特</strong> —— 參數調校與三方對照建議</li>
<li><strong><a href="http://lcz.me" rel="nofollow ugc">lcz.me</a> #1252、#30599 系列</strong> —— vLLM／SGLang 在 RDNA3 的早期實測</li>
</ul>
<h3>技術文章</h3>
<ul>
<li><strong>智源 FlagOpen / FlagGems</strong> —— C++ Wrapper 與 Triton 算子呼叫鏈的實測文章（小算子 wrapper 開銷 60–100 µs 的量級，與我們量到的每步固定開銷同階）</li>
</ul>
<p dir="auto"><strong>若有遺漏，敬請在下方指出 —— 這是我們的疏失，絕非本意。</strong></p>
<hr />
<h2>八、技術附錄（供其他 AI 與工程師取用）</h2>
<p dir="auto"><strong>本節供機器讀取，人類讀者可直接略過。</strong></p>
<h3>8.1 生產配置指紋</h3>
<pre><code>啟動腳本 sha256 : 7f52d3b0bfe876ad34d5c55c25e5e9a42b83afbae01b8a85891eac905f27a114

生產 .so sha256 : 7847e12eaf3d8b104f0227ab779dcfdef4248e843668d461334522598f2f5625
                  （29,823,720 bytes = BLOCK_KN_SIZE 512）
</code></pre>
<h3>8.2 生效參數</h3>
<pre><code>--model-path .../Qwen3.8-27B-INT4-GPTQ-v3
--tp-size 1 --quantization gptq --dtype bfloat16 --mamba-ssm-dtype bfloat16
--kv-cache-dtype bf16 --attention-backend triton
--chunked-prefill-size 8192 --context-length 131072 --mem-fraction-static 0.88
--speculative-algorithm NEXTN    （內部解析為 EAGLE／MTP-3：steps 3 / topk 1 / draft 4）
--cuda-graph-max-bs 32 --sleep-on-idle
--max-running-requests 4 --max-queued-requests 6
--enable-hierarchical-cache --hicache-ratio 1.0 --hicache-size 8
--hicache-write-policy write_through --hicache-io-backend kernel
--hicache-mem-layout page_first --enable-cache-report
--default-chat-template-kwargs {"enable_thinking": false}
（無 --page-size ⇒ page_size = 1）
</code></pre>
<h3>8.3 環境變數（全部，實讀自 /proc/pid/environ）</h3>
<pre><code>SGL_RDNA_CUSTOM_AR=0  SGL_RDNA_NO_FUSED=1  SGL_RDNA_GEMMA_TRITON=1
SGL_RDNA_VLLM_VERIFY=1  SGL_RDNA_LMHEAD_INT4=1  SGL_RDNA_LMHEAD_INT4_GS=128
SGL_RDNA_LMHEAD_INT8=0  SGL_WMMA_KSPLIT_AUTO=1
SGL_DTYPE=bfloat16  SGLANG_PHASE_TIMING=1  SGLANG_ENABLE_HEALTH_ENDPOINT_GENERATION=0
SGLANG_ENABLE_DETERMINISTIC_INFERENCE=0
SGLANG_MAMBA_SSM_DTYPE=bfloat16  SGLANG_CPROFILE_SPEC=0
TVM_FFI_DISABLE_TORCH_C_DLPACK=1
PYTHONPATH=.../prod/rdna_shim
</code></pre>
<h3>8.4 關鍵量測（附儀器）</h3>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>量</th>
<th>值</th>
<th>儀器</th>
</tr>
</thead>
<tbody>
<tr>
<td>decode（DSH agent 路徑 ~20K ctx）</td>
<td>62.11 t/s</td>
<td>引擎 log gen throughput，running-req=1，n=11</td>
</tr>
<tr>
<td>decode（短提示單流）</td>
<td>85.05 t/s</td>
<td>同上</td>
</tr>
<tr>
<td>每步固定開銷（gap &gt; 20 µs）</td>
<td><strong>2.4%</strong></td>
<td>09-11 decode 段 817.12 ms 逐段</td>
</tr>
<tr>
<td>GPU busy（滿載）</td>
<td>89.2 – 94%</td>
<td>同兩次</td>
</tr>
<tr>
<td>GPU 滿載溫度 EDGE / HOTSPOT / MEM</td>
<td>79 / 99 / 94 °C</td>
<td>amd-smi（門檻 100/110/105）</td>
</tr>
<tr>
<td>節流</td>
<td><strong>無</strong>（GFX CLK 2148 ＝ MAX_CLK）</td>
<td>amd-smi metric</td>
</tr>
<tr>
<td>每 token 有效權重頻寬</td>
<td>~385 GB/s（74 t/s × 19 GB ÷ accept 3.65）</td>
<td>推導</td>
</tr>
</tbody>
</table>
<h3>8.5 生產 kernel 的 ISA 拆解（M≥16 WMMA 路徑）</h3>
<p dir="auto">反匯編自生產 .so 的 <code>.hip_fatbin</code> 內嵌 ELF，目標 <code>gemm_q4_wmma_kernel_16x16_1w&lt;__hip_bfloat16&gt;</code>：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>指令</th>
<th>靜態條數</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>v_wmma</code></td>
<td><strong>1</strong></td>
</tr>
<tr>
<td><code>s_waitcnt</code></td>
<td><strong>49</strong></td>
</tr>
<tr>
<td>└ <code>lgkmcnt(0)</code></td>
<td><strong>27</strong></td>
</tr>
<tr>
<td><code>ds_load_u16_d16</code> ＋ <code>_hi</code></td>
<td><strong>8 ＋ 8 ＝ 16</strong>（對應源碼 <code>for (i=0..15) b_frag[i] = b_tile[i][lane_lo]</code>）</td>
</tr>
<tr>
<td><code>ds_bpermute</code></td>
<td>8</td>
</tr>
<tr>
<td><code>global_atomic_cmpswap</code></td>
<td>8</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>編譯器（AMD ROCm LLVM fork，clang 22.0.0git）並非無能</strong> —— 它在其他位置懂得使用 <code>lgkmcnt(1)</code>×98、<code>(2)</code>×20、<code>(7)</code>×24 等局部等待。<strong>是這段存取模式令它只能選擇最保守的方案。</strong></p>
<h3>8.6 上游 vLLM 的新路線（我們尚未採用）</h3>
<p dir="auto"><code>vllm/model_executor/kernels/linear/mixed_precision/rdna_hybrid_w4a16.py</code>：</p>
<pre><code>M &lt;= MAX_SKINNY_BATCH_SIZE (=5) : HIP skinny GEMM (wvSplitK_int4_g)
M &gt;  MAX_SKINNY_BATCH_SIZE      : Triton W4A16 fused dequant GEMM
</code></pre>
<p dir="auto"><strong>我們的 decode 有 80% 屬 M=4，恰好落在第一段；而我們目前兩段皆使用 GPTQ kernel。</strong></p>
<h3>8.7 已發佈的技術層（供取用）</h3>
<pre><code>轉換器 : tools/convert_quark_int4_to_gptq_v3.py（Quark W4A16 → GPTQ，無損）
補丁   : patches/0001..0019.patch（基底 StevenChenSE/sglang gfx1100-support @ 1442c18）
 shim   : shim/rdna_lmhead_int8.py（INT4 LM head，端到端 +19.7%）
 repo   : https://github.com/lawsirlawsir-png/rdna3-quark-w4a16  (tag v1.0.0)

驗證方式：以 HTTPS 下載全部 19 個 patch → 套用到 1442c18 的乾淨檢出 →
          git am 全數通過 → 與生產建置來源 git diff --quiet 無輸出（9,034 檔相同）。
</code></pre>
<hr />
<h3>8.8 硬體身份（實測）</h3>
<pre><code>MARKET_NAME   : AMD Radeon PRO W7800 48GB
SUBVENDOR_ID  : 0x1458  (GIGABYTE)
DEVICE_ID     : 0x7449   SUBSYSTEM_ID: 0x2428
NUM_COMPUTE_UNITS: 70     TARGET_GRAPHICS_VERSION: gfx1100
PCIE          : Gen 4 x16（MAX_PCIE_WIDTH 16 / MAX_PCIE_SPEED 16 GT/s）
vBIOS         : W7800 48G/F1/1133
</code></pre>
]]></description><link>https://lcz.me/topic/1718</link><generator>RSS for Node</generator><lastBuildDate>Fri, 25 Sep 2026 22:14:51 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1718.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 15 Sep 2026 07:56:57 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to （15／9第二次更新）為 gfx1100 用戶追回軟件算力：GIGABYTE W7800 48G 單卡跑通 AMD 原版 Quark INT4，DSH 真實負載七日 +22%（50.9 → 62.11 t/s） on Wed, 16 Sep 2026 03:11:19 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/chia-an-yang" aria-label="Profile: CHIA-AN-YANG">@<bdi>CHIA-AN-YANG</bdi></a> 其實我也測出9X T／S， 不過那個數據没有用， 始終大家都是用AGENT HARNESS，我在DSH測出這個數字也非常夠用了， 我的目標是在AMD的卡上追趕NVIDIA，有60－70％的效果我個人覺得已相當不錯。</p>
]]></description><link>https://lcz.me/post/18511</link><guid isPermaLink="true">https://lcz.me/post/18511</guid><dc:creator><![CDATA[Wing Wah Law]]></dc:creator><pubDate>Wed, 16 Sep 2026 03:11:19 GMT</pubDate></item><item><title><![CDATA[Reply to （15／9第二次更新）為 gfx1100 用戶追回軟件算力：GIGABYTE W7800 48G 單卡跑通 AMD 原版 Quark INT4，DSH 真實負載七日 +22%（50.9 → 62.11 t/s） on Wed, 16 Sep 2026 03:08:37 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 謝謝讚賞<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f605.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--sweat_smile" style="height:23px;width:auto;vertical-align:middle" title="😅" alt="😅" /> 不過，我真的不知道AI做了什麼的， 我不是程序員PROGRAMMER，數學也没學好。做1％工作就是這樣，小突破也不錯。我10天玩出一個花樣也算是為這社群做一次有1％原創性的貢獻。</p>
<p dir="auto">我投產了，將來有機會再玩一下。謝謝你提供這個平台，祝TERRY兄和其他兄弟姐妹生活順順利利<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/2728.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--sparkles" style="height:23px;width:auto;vertical-align:middle" title="✨" alt="✨" /></p>
]]></description><link>https://lcz.me/post/18510</link><guid isPermaLink="true">https://lcz.me/post/18510</guid><dc:creator><![CDATA[Wing Wah Law]]></dc:creator><pubDate>Wed, 16 Sep 2026 03:08:37 GMT</pubDate></item><item><title><![CDATA[Reply to （15／9第二次更新）為 gfx1100 用戶追回軟件算力：GIGABYTE W7800 48G 單卡跑通 AMD 原版 Quark INT4，DSH 真實負載七日 +22%（50.9 → 62.11 t/s） on Wed, 16 Sep 2026 01:02:31 GMT]]></title><description><![CDATA[<p dir="auto">@CHIA AN YANG 这篇 AI 翻译有一处要纠正：7900 XTX 比 W7800 快，不是「96 CU vs 70 CU 所以快 37%」。</p>
<p dir="auto">gfx1100 上 decode 是带宽瓶颈不是算力瓶颈：7900 XTX 960 GB/s，W7800 48G 只有 576 GB/s，光带宽就差 1.6×。所以 27B INT4 长上下文 decode 上 XTX 会更快，但幅度由带宽决定，不是 CU 数；CU 数主要影响 prefill（compute-bound）。</p>
<p dir="auto">另外 62 t/s @ 20–24K 在单卡 gfx1100 上不算「生产力不够」，是正常区间：decode 每步要读全部权重（27B INT4 ~14GB）+ KV（fp16 20K 约 2–3GB），带宽吃满就是这个量级。想再快只能降 KV 精度、缩上下文或上投机解码，加 CU 没用。</p>
]]></description><link>https://lcz.me/post/18475</link><guid isPermaLink="true">https://lcz.me/post/18475</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Wed, 16 Sep 2026 01:02:31 GMT</pubDate></item><item><title><![CDATA[Reply to （15／9第二次更新）為 gfx1100 用戶追回軟件算力：GIGABYTE W7800 48G 單卡跑通 AMD 原版 Quark INT4，DSH 真實負載七日 +22%（50.9 → 62.11 t/s） on Tue, 15 Sep 2026 22:44:27 GMT]]></title><description><![CDATA[<p dir="auto">讀完,,讓ai翻譯,20-24k上下文,感覺生產力不夠?</p>
<p dir="auto">直接幫你抽取出「實際開多少上下文」與「實際輸出速度（Token/s）」，對比清清楚楚：</p>
<hr />
<p dir="auto">一、他到底測了什麼？（平台環境）</p>
<ul>
<li>顯示卡：AMD 專業卡 Radeon PRO W7800 48GB（單卡，跟你的 7900 XTX 一樣是 RDNA3 / gfx1100 架構，但你的 7900 XTX 運算單元比他強！他 70 CU，你的 7900 XTX 是 96 CU 滿血版）。</li>
<li>跑的模型：Qwen3.8-27B（INT4 量化版）。</li>
<li>跑的框架：SGLang（透過魔改支援 AMD ROCm / RDNA3）。</li>
</ul>
<hr />
<p dir="auto">二、實際「上下文」開多少？實際「Token 速度」多少？</p>
<p dir="auto">作者測了 兩種情境（這就是為什麼文章裡數字跳來跳去）：</p>
<ol>
<li>【真實 Agent 重度工作負載】（對話很長時，也就是你目前關心的指標）</li>
</ol>
<ul>
<li>實際上下文深度（Context）：約 20,000 ~ 24,000 Tokens (20K~24K)<br />
(原文數據：full token 20,004 – 20,944，複驗中位數 24,299)</li>
<li>實際解碼輸出速度（Decode Throughput）：<br />
<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f449.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--point_right" style="height:23px;width:auto;vertical-align:middle" title="👉" alt="👉" /> 61.5 ~ 62.11 t/s（Token/秒）<br />
(這是在吃滿 2 萬多字上下文的情況下，每秒還能吐出 62 個字！)</li>
</ul>
<ol start="2">
<li>【短提示單流測試】（問一句簡短問題時）</li>
</ol>
<ul>
<li>實際上下文深度（Context）：短提示（極小上下文）</li>
<li>實際解碼輸出速度（Decode Throughput）：<br />
<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f449.png?v=ecb7c61779a" class="not-responsive emoji emoji-android emoji--point_right" style="height:23px;width:auto;vertical-align:middle" title="👉" alt="👉" /> 85.05 t/s（Token/秒）<br />
(如果開啟他改寫的 DFLASH 投機解碼，短文最快甚至能飆到 92.57 t/s)</li>
</ul>
<hr />
<p dir="auto">三、三行白話文總結他整篇在講什麼：</p>
<ol>
<li>上下文越長越慢：上下文在短的時候能跑 85 t/s，但當上下文吃到 20K~24K 時，速度會掉到 62 t/s。</li>
<li>AMD 顯卡跑大模型有搞頭：他證明了 AMD RDNA3（gfx1100）單張顯卡跑 27B 大模型，即使上下文塞了 2 萬字，依然能有 62 t/s 的極高實用速度。</li>
<li>對店長 7900 XTX 的意義：你的 7900 XTX 算力核心（96 CU）比這張 W7800（70 CU）高出近 37%，只要顯存放得下，在類似架構下的理論推論速度只會比他更快！</li>
</ol>
]]></description><link>https://lcz.me/post/18461</link><guid isPermaLink="true">https://lcz.me/post/18461</guid><dc:creator><![CDATA[CHIA AN YANG]]></dc:creator><pubDate>Tue, 15 Sep 2026 22:44:27 GMT</pubDate></item><item><title><![CDATA[Reply to （15／9第二次更新）為 gfx1100 用戶追回軟件算力：GIGABYTE W7800 48G 單卡跑通 AMD 原版 Quark INT4，DSH 真實負載七日 +22%（50.9 → 62.11 t/s） on Tue, 15 Sep 2026 19:28:50 GMT]]></title><description><![CDATA[<p dir="auto">非常牛逼的分享，只是这个玩法太小众</p>
]]></description><link>https://lcz.me/post/18446</link><guid isPermaLink="true">https://lcz.me/post/18446</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 15 Sep 2026 19:28:50 GMT</pubDate></item><item><title><![CDATA[Reply to （15／9第二次更新）為 gfx1100 用戶追回軟件算力：GIGABYTE W7800 48G 單卡跑通 AMD 原版 Quark INT4，DSH 真實負載七日 +22%（50.9 → 62.11 t/s） on Tue, 15 Sep 2026 10:03:19 GMT]]></title><description><![CDATA[<p dir="auto">数据很扎实，尤其把每步固定开销（gap&gt;20µs）、accept length 和吞吐一起报，比只给 t/s 可信得多。</p>
<p dir="auto">两点建议，都能在你现有配置上做单变量 A/B：</p>
<ol>
<li>你 8.6 提到上游 vLLM 的 rdna_hybrid_w4a16 路线——你 80% 的 decode 落在 M=4，正好是 MAX_SKINNY_BATCH_SIZE=5 的 HIP skinny GEMM（wvSplitK_int4_g）分支。可以只把 decode 段切过去、prefill 仍留现有 GPTQ kernel，成本最低。</li>
<li>page-size=1 + write_through + hicache-size 8 在 30GB RAM 下是合理的；但命中率只有 65–77% 时，write_through 的 host 写放大可能被低估，可以对比 write_back 看 decode 段是否更平稳。</li>
</ol>
<p dir="auto">另外你把 W7800（70CU）和 7900XTX（96CU）、W7900 的边界写得很清楚，这点最容易被抄参数的人忽略。</p>
]]></description><link>https://lcz.me/post/18348</link><guid isPermaLink="true">https://lcz.me/post/18348</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 15 Sep 2026 10:03:19 GMT</pubDate></item></channel></rss>