<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比]]></title><description><![CDATA[<pre><code>   发现论坛里玩3090的朋友不少，但似乎都是各忙各的，每次有人晒跑分，底下想对比都没法比——基准不一样，配置也没标注，还有致命的一点是，只有一个t/s，想抄作业 ，但是又怕下载回来折腾一阵，只能跑32K 64K，实际跑不起来生产力，热闹一阵就散了。
</code></pre>
<p dir="auto">所以想发起一个小活动，把大家的经验和数据攒到一起，攒出一份属于3090玩家的真实对比。分两个阶段，都很简单：</p>
<p dir="auto">第一阶段：报名 + 投票（本帖）<br />
想参加的，直接在主帖跟贴 留言 (其实也不限3090），选择你认可的评测套件，每张显卡一票（顺便说下自己的配置和NVLink情况更好）。</p>
<p dir="auto">然后用投票决定统一用哪个测试套件。候选先列几个，欢迎补充：</p>
<p dir="auto">zx-bench：<a href="https://github.com/suncityldp/zx-bench" rel="nofollow ugc">https://github.com/suncityldp/zx-bench</a> 楼主今天凌晨已经测了一下27B，感觉时间挺长。一套配置就 要3小时 。。。<br />
tooleval<br />
benchlocal<br />
其他你觉得合适的（GAIA、AgentBench之类也行）<br />
投票方式：回帖“投票 + 套件名 + 真实硬件照片”，少数服从多数。<br />
截止日期：下周一9月21日 晚上23点59分</p>
<p dir="auto">第二阶段：新帖晒成果<br />
投票确定下来后我会另开一个新帖（链接到时候发在这里）。大家每周在周日前，把自己当时觉得最强的配置和跑分结果 ，发到新帖里，内容包括但不限于：</p>
<p dir="auto">跑分 + 截图（有图有真相）<br />
最满意的作品也可以晒（token efficiency 更佳）<br />
附上硬件环境（有无NVLink）和操作系统环境 及选择原因等。<br />
本活动接受各种建议。</p>
<p dir="auto">附则：投票权重折算表（以3090=1票为基准）<br />
投票权重按各卡稠密FP16张量算力相对3090折算。计算公式：票数 = FP16稠密 / 71，四舍五入到0.5的倍数。<br />
未列出的卡可自行计算：查该卡FP16张量算力（稠密值），除以71即可。</p>
<p dir="auto">消费级游戏卡</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>显卡</th>
<th>稠密FP16（TFLOPS）</th>
<th>折算票数</th>
<th>参考价格</th>
<th></th>
<th></th>
<th></th>
<th></th>
</tr>
</thead>
<tbody>
<tr>
<td>RTX 3060 12G</td>
<td>~25</td>
<td>0.5</td>
<td>1.5-2k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 4060 / 4060 Ti</td>
<td>~30-44</td>
<td>0.5</td>
<td>2-3k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 3070 / 3070 Ti</td>
<td>~41-44</td>
<td>0.5</td>
<td>2-3k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 3080 / 3080 Ti</td>
<td>~60-68</td>
<td>1</td>
<td>3-5k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 3090</td>
<td>71</td>
<td>1（基准）</td>
<td>5-8k二手</td>
<td>RTX 3090 Ti</td>
<td>~80</td>
<td>1</td>
<td>6-9k</td>
</tr>
<tr>
<td>RTX 4070 / 4070 Ti</td>
<td>~58-80</td>
<td>1</td>
<td>3-6k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 4080</td>
<td>~97</td>
<td>1.5</td>
<td>6-8k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 5070</td>
<td>~124</td>
<td>1.5</td>
<td>4-5k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 4090</td>
<td>~165</td>
<td>2.5</td>
<td>15-20k二手</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 5070 Ti</td>
<td>~176</td>
<td>2.5</td>
<td>5-7k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 5080</td>
<td>~225</td>
<td>3</td>
<td>7-9k</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>RTX 5090</td>
<td>~419</td>
<td>（÷71 ≈ 6）</td>
<td>?</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
</tbody>
</table>
]]></description><link>https://lcz.me/topic/1748</link><generator>RSS for Node</generator><lastBuildDate>Fri, 18 Sep 2026 12:46:30 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1748.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 16 Sep 2026 08:26:16 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 12:39:02 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/eric-su" aria-label="Profile: Eric-Su">@<bdi>Eric-Su</bdi></a> 不支持dflash 比较可惜了</p>
]]></description><link>https://lcz.me/post/19151</link><guid isPermaLink="true">https://lcz.me/post/19151</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Fri, 18 Sep 2026 12:39:02 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 12:36:30 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/applejuice" aria-label="Profile: applejuice">@<bdi>applejuice</bdi></a> 正在測...但看來效能可能會很慘(GPU沒吃滿，另不支持DFlash)<br />
結果::<br />
pp2×tp2 POC 實測結果(27B)</p>
<p dir="auto">可行性 — 過了<br />
pp2×tp2 在 Qwen3.8-27B 混合架構(GDN mamba+MTP)上能穩定起、能推理。最大的未知雷——mamba 層跨 pp 段——沒炸:每 rank Mamba Cache 正常分配、TritonGDNKernel 正常 dispatch、四 rank(PP0/PP1×TP0/TP1)排布正確。這是 POC 要驗的核心,達成。</p>
<p dir="auto">速度 — 27B 上全面輸現役(如預期)</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>口徑</th>
<th>pp2×tp2(無spec)</th>
<th>現役 tp2+DFLASH</th>
</tr>
</thead>
<tbody>
<tr>
<td>單路代碼</td>
<td>71 t/s</td>
<td>260 t/s(-73%)</td>
</tr>
<tr>
<td>單路中文散文</td>
<td>70</td>
<td>73</td>
</tr>
<tr>
<td>併發4路代碼</td>
<td>192</td>
<td>361(-47%)</td>
</tr>
</tbody>
</table>
<p dir="auto">主因兩層:</p>
<ol>
<li>丟了 DFLASH spec(pp≠1 硬互斥)——tok/chunk 從 ~6 掉到 1.00,單路少吐一堆 token,這是大頭。</li>
<li>pp bubble + 每層仍走 PCIe。</li>
</ol>
<p dir="auto"><img src="https://upload.lcz.me/uploads/5a23b6ca-9618-42e1-aded-8fe87763e4b0.jpeg" alt="e9348f43-ee5d-4422-adca-ed78ac7e2678-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/19150</link><guid isPermaLink="true">https://lcz.me/post/19150</guid><dc:creator><![CDATA[Eric Su]]></dc:creator><pubDate>Fri, 18 Sep 2026 12:36:30 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 12:15:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/eric-su" aria-label="Profile: Eric-Su">@<bdi>Eric-Su</bdi></a> 大佬 测一测 pp=2 × tp=2</p>
<p dir="auto">我想玩很久了 所以之前问ai 问了很多<br />
据ai 单发 可能比较慢<br />
但是多发就快了<br />
而且有96gb vram</p>
<p dir="auto">开个4卡3090帖子吧<br />
分享下硬件 psu 之类的<br />
我想了解多点<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f606.png?v=0650a1064dd" class="not-responsive emoji emoji-android emoji--laughing" style="height:23px;width:auto;vertical-align:middle" title=":laughing:" alt="😆" /></p>
]]></description><link>https://lcz.me/post/19148</link><guid isPermaLink="true">https://lcz.me/post/19148</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Fri, 18 Sep 2026 12:15:58 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 11:59:12 GMT]]></title><description><![CDATA[<p dir="auto">也是，除非有 四路的 NVLINK...<br />
不過記錄喚起我的記憶，也是因為 SGLang 不支持會出CustomAllreduce 異常，當下就沒繼續鑽研了</p>
]]></description><link>https://lcz.me/post/19145</link><guid isPermaLink="true">https://lcz.me/post/19145</guid><dc:creator><![CDATA[Eric Su]]></dc:creator><pubDate>Fri, 18 Sep 2026 11:59:12 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 11:56:28 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/eric-su" aria-label="Profile: Eric-Su">@<bdi>Eric-Su</bdi></a> 有nvlink 一样慢，因为还是要经过pcie</p>
<p dir="auto">问过ai 的确decode可能变慢 但是prefill 应该好点<br />
跟两张nvlink 一样</p>
]]></description><link>https://lcz.me/post/19143</link><guid isPermaLink="true">https://lcz.me/post/19143</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Fri, 18 Sep 2026 11:56:28 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 12:11:17 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/applejuice" aria-label="Profile: applejuice">@<bdi>applejuice</bdi></a> 有，因 PCIE 頻寬頻頸(沒NVLink)，效率低下(雖然96G顯存很吸引人)...<br />
剛剛又重新跟AI對話下，更新結論</p>
<p dir="auto">3090(無 NVLink,全程 PCIe)並行配置比較表<br />
數字全取自 skill 已固化實測(2026-09-15,社群 uni_bench 口徑,代碼場景)</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>比較項</th>
<th>tp=2(單實例,用2卡)</th>
<th>tp=4(四卡合一)</th>
<th>tp=2×2(雙實例,各2卡)</th>
</tr>
</thead>
<tbody>
<tr>
<td>卡間鏈路</td>
<td>PCIe</td>
<td>PCIe</td>
<td>PCIe</td>
</tr>
<tr>
<td>all-reduce kernel</td>
<td>custom(較省)</td>
<td>退 NCCL 通用(較慢)</td>
<td>custom(較省)</td>
</tr>
<tr>
<td>通信範圍</td>
<td>2 卡對傳</td>
<td>4 卡集合,含跨 PHB 組慢跳</td>
<td>2 卡/實例,零跨組</td>
</tr>
<tr>
<td>單路 decode</td>
<td>249.7 t/s</td>
<td>227.9 t/s(↓9%)</td>
<td>261.7 / 271.8 t/s(各實例)</td>
</tr>
<tr>
<td>併發 4 路 agg</td>
<td>361.0 t/s</td>
<td>228.4 t/s(↓37%)</td>
<td>342.7 / 367.1 t/s(各實例)</td>
</tr>
<tr>
<td>系統總吞吐</td>
<td>361.0 t/s(2卡用,2卡閒)</td>
<td>228.4 t/s</td>
<td>709.8 t/s(兩實例相加,↑97%)</td>
</tr>
<tr>
<td>單模型可用顯存</td>
<td>48GB</td>
<td>96GB(單一池)</td>
<td>48GB × 2(獨立)</td>
</tr>
<tr>
<td>卡使用</td>
<td>2 用 / 2 閒</td>
<td>4 全用</td>
<td>4 全用(2 實例)</td>
</tr>
</tbody>
</table>
<p dir="auto">口徑警告(避免誤讀)</p>
<ul>
<li>709.8 是兩實例併發吞吐「相加」的系統總產能,不是單一請求變快。談速度看單路那列。</li>
<li>單路對照(公平比速度):tp=4 227.9 比 tp=2 249.7 還慢 9% —— tp=4 純速度單路、併發兩頭都輸。</li>
<li>tp=2×2 贏在系統產能(多開一實例吃第二組卡),非單路更快;單路它 ~265,與 tp=2 同級(略高,因兩實例各獨佔一組 PHB 無干擾)。</li>
</ul>
<p dir="auto">一句話結論</p>
<ul>
<li>現役 27B(~19GB,48GB 綽綽有餘):tp=2×2 系統產能最高(709),定案最優。</li>
<li>tp=4 唯一價值 = 96GB 單一顯存池,只有在「模型/context 塞不進 48GB」時才選,代價是單路 -9%、併發 -37%。速度維度全輸。</li>
</ul>
]]></description><link>https://lcz.me/post/19140</link><guid isPermaLink="true">https://lcz.me/post/19140</guid><dc:creator><![CDATA[Eric Su]]></dc:creator><pubDate>Fri, 18 Sep 2026 12:11:17 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 11:40:55 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/eric-su" aria-label="Profile: Eric-Su">@<bdi>Eric-Su</bdi></a> 大佬 有没有测试tp=4？</p>
]]></description><link>https://lcz.me/post/19137</link><guid isPermaLink="true">https://lcz.me/post/19137</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Fri, 18 Sep 2026 11:40:55 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Fri, 18 Sep 2026 11:02:28 GMT]]></title><description><![CDATA[<p dir="auto">最近中秋前夕太忙哩，本來就想支持的，拖到現在<br />
感謝 <a class="plugin-mentions-user plugin-mentions-a" href="/user/starryskyknight" aria-label="Profile: starryskyknight">@<bdi>starryskyknight</bdi></a> 大神讓我少走很多彎路</p>
<p dir="auto">架構：tp=2×2 雙實例 + sglang-router</p>
<p dir="auto">對外入口   sglang-router → :8000 → round_robin 分流到 8001/8002<br />
實例 A     sglang-a → :8001 (GPU0/1, tp=2)  healthy<br />
實例 B     sglang-b → :8002 (GPU2/3, tp=2)  healthy</p>
<p dir="auto">模型<br />
target: /models/Qwen3.8-27B-AWQ-MTP (twolven abliterated)<br />
draft:  /models/Qwen3.8-27B-DFlash2<br />
served-name: qwen3.8-27b，架構 Qwen3_5ForConditionalGeneration</p>
<p dir="auto">兩實例共用啟動配方（一模一樣）<br />
--tp-size 2  --context-length 262144<br />
--mem-fraction-static 0.86<br />
--kv-cache-dtype fp8_e5m2<br />
--speculative-algorithm DFLASH  --speculative-num-draft-tokens 8  (draft=DFlash2)<br />
--mamba-ssm-dtype bfloat16  --max-mamba-cache-size 37   ← P3 修的 mamba cap<br />
--max-running-requests 4<br />
--chunked-prefill-size 2048  --max-prefill-tokens 8192<br />
--enable-hierarchical-cache  --hicache-ratio 6.0  --hicache-write-policy write_back  ← 真復用<br />
enable_thinking=false（全域關 think）<br />
reasoning-parser qwen3 / tool-call-parser qwen3_coder</p>
<p dir="auto">硬體 / 系統<br />
4× RTX 3090，全卡 260W + persistence=Enabled，目前 idle（util 0%，溫 29-56°C）<br />
顯存：四卡各佔滿 ~23GB/24GB（模型常駐）<br />
RAM：247G，用 190G / 剩 56G（雙實例 hicache 各 ~43G）</p>
<p dir="auto">DFLASH2 壓測 — 對齊 <a href="http://lcz.me/topic/1502" rel="nofollow ugc">lcz.me/topic/1502</a> uni_bench 口徑</p>
<p dir="auto">測試規格（帖子統一口徑，強度加倍版）<br />
temp=0, streaming+include_usage, enable_thinking=false<br />
rounds=6（帖子3）, max_tokens=1024（帖子512）, 並發4路（帖子2）<br />
decode=(comp-1)/(total-ttft)，只信 usage.completion_tokens</p>
<p dir="auto">A. 單實例 :8001 (tp=2, GPU0/1) — 與帖子同構（2×3090 tp=2 單實例）</p>
<p dir="auto">單路 decode<br />
代碼生成      269.5 t/s   (tok/chunk 6.01, TTFT 0.075s)<br />
英文技術論述  169.8 t/s   (3.81)<br />
中文常規對話  119.4 t/s   (2.68)<br />
中文散文       70.0 t/s   (1.57)</p>
<p dir="auto">並發4路 aggregate<br />
同 prompt (radix 命中)   714.0 t/s<br />
不同 prompt              388.0 t/s</p>
<p dir="auto">A.對照帖子（NVLink 王者版 / 無NVLink applejuice，都 rounds=3 max512 並發2）</p>
<p dir="auto">場景          ws-ai(我)  樓主NVLink  applejuice無NVLink<br />
代碼單路        269.5      237/246.7      141/235.9<br />
英文技術        169.8         166           140.8<br />
中文對話        119.4         123            99.0<br />
中文散文         70.0          92            61.4<br />
並發同prompt   434.6(2路)     440           429.7<br />
並發4路同       714.0          —              —</p>
<p dir="auto">判讀</p>
<ul>
<li>單路全項贏或持平 applejuice（同為無 NVLink 同硬體），且贏樓主 NVLink 版的代碼/英文——無 NVLink 非瓶頸再次坐實。中文散文 70 稍低於樓主 92，是 accept len 低場景（tok/chunk 1.57），draft 命中差異，非環境問題。</li>
<li>並發同 prompt：2路434.6 對齊帖子 430/440；4路衝到 714，radix 命中下近線性放大（mrr=4 吃滿）。</li>
</ul>
<p dir="auto">B. 全系統 router :8000（雙實例 tp=2×2, 四卡全用）</p>
<p dir="auto">單路 decode（經 router 分流，與單實例幾乎一致）<br />
代碼生成      265.1 t/s<br />
英文技術論述  172.2 t/s<br />
中文常規對話  123.1 t/s<br />
中文散文       71.6 t/s</p>
<p dir="auto">並發8路 aggregate（雙實例 mrr=8 全壓）<br />
同 prompt (radix 命中)   1387.1 t/s<br />
不同 prompt               663.9 t/s</p>
<p dir="auto">系統總產能全景（同代碼場景，同口徑）</p>
<p dir="auto">配置              同prompt理想上界   不同prompt實戰<br />
單實例 tp=2  2路      434.6              433.0<br />
單實例 tp=2  4路      714.0              388.0<br />
雙實例 8路（全系統）  1387.1             663.9</p>
<p dir="auto">判讀</p>
<ul>
<li>系統峰值 1387 t/s（8路同prompt），對比單實例4路714，雙實例近乎線性翻倍（1.94×）——四卡 tp=2×2 佈局的產能翻倍在系統級 uni_bench 口徑下坐實，跟 skill 記的 709 t/s（舊 rounds3/max512 版）同型放大。</li>
<li>不同 prompt 8路 663.9：這是最貼近來福真實負載的數字——8 張不同 issue 並發，無共享 radix，overlap 打折後系統仍能吐 664 t/s。對比單實例4路不同prompt 388，雙實例約1.71×。</li>
<li>單路 decode 經 router（265/172/123/72）vs 直打單實例（270/170/119/70）幾乎無損，round_robin 分流沒吃掉單路速度。</li>
</ul>
<p dir="auto"><img src="https://upload.lcz.me/uploads/8d504f09-a1fa-46e6-aaf1-3ef77a4ea648.jpeg" alt="image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/19129</link><guid isPermaLink="true">https://lcz.me/post/19129</guid><dc:creator><![CDATA[Eric Su]]></dc:creator><pubDate>Fri, 18 Sep 2026 11:02:28 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 17:05:37 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/applejuice" aria-label="Profile: applejuice">@<bdi>applejuice</bdi></a> 所以我建议晚上在算力空闲的时候测。</p>
]]></description><link>https://lcz.me/post/18911</link><guid isPermaLink="true">https://lcz.me/post/18911</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Thu, 17 Sep 2026 17:05:37 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 14:29:14 GMT]]></title><description><![CDATA[<p dir="auto">对我来说 太麻烦了<br />
ai 跟我说 测一次 2-3个小时 我立刻就没心继续了</p>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/starryskyknight" aria-label="Profile: starryskyknight">@<bdi>starryskyknight</bdi></a> 你之前那个贴 的成绩也是用375w 跑出来的？</p>
]]></description><link>https://lcz.me/post/18872</link><guid isPermaLink="true">https://lcz.me/post/18872</guid><dc:creator><![CDATA[applejuice]]></dc:creator><pubDate>Thu, 17 Sep 2026 14:29:14 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 13:01:48 GMT]]></title><description><![CDATA[<p dir="auto">数据收到，回几点 + 两个建议复测的点。</p>
<p dir="auto">1）单请求 138.8 t/s 基本就是这套配置的带宽天花板。AWQ-4bit 27B 权重约 14–15GB，TP=2 每卡读一半，按 3090 936 GB/s 折算理论 decode 约 125–135 t/s；跑到 138.8 说明权重读取已经贴住利用率上限，调度没拖后腿，再往上只能靠并发把 kernel 缝隙填满。</p>
<p dir="auto">2）2 并发 260.9（1.88×）正常；但 4 并发合计反而掉到 243.4，比 2 并发低，这点值得查。262K ctx + fp8 KV 在 24G×2 上余量很薄，怀疑是 KV 装不下触发了 preempt/retract，或 chunked prefill 抢了 decode 的算力。建议看 SGLang 日志里 <code>#retract</code>/<code>#preempt</code> 计数和 mem_fraction_static，把 ctx 收到实际用量再复测，4 并发应能回到接近 2 并发的 1.8–1.9× 线性区。</p>
<p dir="auto">3）NVLink「400 字回答搬 10.1GB」我持保留。TP=2 每 token 每层 all-reduce 约 2×hidden×2B≈20KB，60+ 层就是每 token 约 1.2–1.5MB，400 token 大致 0.5GB 量级。10.1GB 高了近 20×，更像把同一进程之前的 prefill 和 MTP 多步 draft forward 一起算进去了。建议清空计数后跑固定 prompt，测「增量字节 / 生成 token」，prefill 与 decode 分开报。</p>
<p dir="auto">4）「平均利用率 0.06% 所以 NVLink 不是瓶颈」方向对，但推理要小心：TP 的 all-reduce 是短促突发、延迟敏感的，用占峰值百分比衡量必然很低，低利用率不等于没用。该看每层 all-reduce 延迟，以及 <code>nvidia-smi topo -m</code> 是否真出 NV4、NCCL 是否走 P2P（Z890 上这步常被 ACS/驱动吃掉）。这两样贴出来对双卡党最有价值。</p>
<p dir="auto">投票 zx-bench 没问题，请按你说的两层交：zx-bench 长协议打底 + 一组固定 p/n、pp/tg 分开的短协议，长短都交，跨卡才真能对照。</p>
]]></description><link>https://lcz.me/post/18855</link><guid isPermaLink="true">https://lcz.me/post/18855</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 17 Sep 2026 13:01:48 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 12:35:23 GMT]]></title><description><![CDATA[<p dir="auto">看起来参与热情不高啊，我弟你还是要拿出点简单的方法啊。</p>
]]></description><link>https://lcz.me/post/18851</link><guid isPermaLink="true">https://lcz.me/post/18851</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Thu, 17 Sep 2026 12:35:23 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 11:16:57 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/39018b28-b042-402a-8dc8-233565d4b9ad.png" alt="IMG_20260917_201523_794.png" class=" img-fluid img-markdown" /></p>
<p dir="auto">【投票：zx-bench】双 3090 + NVLink 配置卡，附一组实测数据</p>
<p dir="auto">投票选 zx-bench。理由：仓库开放、题目可哈希复核、维度覆盖比较全。也赞成上面 <a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 的意见——建议在长协议之外附加一组短协议数据（固定输入/输出长度、pp 与 tg 分开报），两种都交，方便跨卡对照。双 3090 按规则计 2 票。</p>
<p dir="auto">既然是晒配置，先把我的双 3090 机器交出来，欢迎拍砖：</p>
<p dir="auto"><strong>硬件</strong></p>
<ul>
<li>2 × RTX 3090 24G（功耗墙 375W × 2），NVLink 桥接（NV4 链路）</li>
<li>CPU：Ultra 9 285K；内存 60G；主板 Z890</li>
<li>系统：Ubuntu，驱动 610.x</li>
</ul>
<p dir="auto"><strong>软件栈</strong></p>
<ul>
<li>引擎：SGLang（main 分支），TP=2</li>
<li>模型：Qwen3.8-27B（AWQ 4bit）+ 多步 MTP / DFlash2 投机解码</li>
<li>KV cache：fp8_e4m3；上下文窗口 262,144 全开</li>
<li>其他：HiCache（write_back）、chunked prefill、投机块 8/8</li>
</ul>
<p dir="auto"><strong>实测（自测口径，输出固定 1024 tokens）</strong></p>
<ul>
<li>单请求：138.8 tok/s</li>
<li>2 并发合计：260.9 tok/s；4 并发合计：243.4 tok/s</li>
<li>长提示（约 5.8K 输入）热缓存后 TTFT 约 70ms</li>
</ul>
<p dir="auto"><strong>NVLink 观测（双卡党稀缺数据）</strong></p>
<ul>
<li>一次约 400 字回答，NVLink 实际搬运约 10.1 GB；空闲时为 0</li>
<li>开机 3 天累计约 9.15 TB；平均利用率约为峰值的 0.06%</li>
<li>结论：对这套配置 NVLink 不是瓶颈；TP2 的收益是把集合通信从 PCIe 挪开，而不是带宽不够用</li>
</ul>
<p dir="auto"><strong>同一对卡跑别的任务</strong>（供对照）</p>
<ul>
<li>图像批处理：每卡一个实例并行，4 张去背级任务 14.7s（约 1.6–2×）</li>
<li>生成类：ComfyUI 官方 MultiGPU CFG Split（cfg&gt;1 工作流）实测 1.71×（24s → 14s）</li>
<li>换用途开销：LLM 服务 ⇄ 媒体任务，去程约 24s、回程约 44s</li>
</ul>
<p dir="auto">硬件实拍：机器不在手边，先附系统实况截图（nvidia-smi + NVLink 拓扑，实机实跑）；实机照片回头补。数据为单机自测、自写脚本，欢迎按统一协议复测交流。</p>
]]></description><link>https://lcz.me/post/18838</link><guid isPermaLink="true">https://lcz.me/post/18838</guid><dc:creator><![CDATA[starryskyknight]]></dc:creator><pubDate>Thu, 17 Sep 2026 11:16:57 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 06:28:40 GMT]]></title><description><![CDATA[<p dir="auto">CBS 已经跑完了，关键发现是：INT8 KV cache 在双 3090 上已经有多个社区实测案例。让我拉几个关键原文做精确对比：</p>
<p dir="auto">从已读取的 CBS 搜索结果和 refs 文件中，我已经掌握了足够的信息来回答这个问题。让我整理一下：</p>
<p dir="auto">关键发现 1：syv-ai/qwen38-27b-rtx3090 项目（1.4k stars）</p>
<p dir="auto">这个仓库专门针对单卡 RTX 3090 做了深度优化，使用了 INT8 激活量化 + FP8 KV cache 的组合：</p>
<ul>
<li>基准测试：vLLM 0.27.1，单卡 3090 @ 250W</li>
<li>并发 64 请求时达到 ~1,094 tok/s（e2e）</li>
<li>使用 int8 activations on the MLP GEMMs（不是全 INT8 KV，而是 MLP 层的 INT8 激活）</li>
<li>上下文容量通过 INT8 缓存翻倍</li>
</ul>
<p dir="auto">关键发现 2：Reddit 用户 iamMess 的 381 tps 记录</p>
<p dir="auto">这个用户在 RTX 3090 上跑出了 381 tps 的单请求成绩：</p>
<ul>
<li>使用 INT8 KV cache（从 69,758 tokens 翻倍到 138,696 tokens）</li>
<li>组合方案：fp8 KV + int8 lm_head/embed_tokens + fp16 recurrent state + int8 activations</li>
<li>需要打 patch 到 vLLM 才能使用 INT8 KV</li>
<li>结合 DFlash2 + lookup-augmented drafting 实现高吞吐</li>
</ul>
<p dir="auto">关键发现 3：yhavinga/vllm-gemma-int8-kv-rtx3090</p>
<p dir="auto">这个 fork 实现了完整的 INT8 KV cache 支持：</p>
<ul>
<li>支持 Gemma 2-27B 在 RTX 3090 上运行 128K 上下文</li>
<li>+87% throughput on memory-bound workloads</li>
<li>使用 per-layer scales 进行 INT8 量化</li>
<li>是目前最完整的 INT8 KV 实现</li>
</ul>
<p dir="auto">关键发现 4：lued/Qwen3.8-27B-INT8-W8A16-MTP</p>
<p dir="auto">HuggingFace 上的模型卡明确说明：</p>
<ul>
<li>专为 Ampere 架构 GPU（RTX 3090）设计</li>
<li>使用 W8A16 量化 + INT8 KV cache</li>
<li>保留了 vision tower、thinking control 等特性</li>
<li>目标是在没有原生 FP8 tensor core 的硬件上获得最佳性能</li>
</ul>
<p dir="auto">结论：</p>
<ol>
<li>INT8 KV cache 在 3090 上确实有实际用例，多个社区用户已经成功部署</li>
<li>主要收益：内存节省约 2×（从 FP8 或 BF16 到 INT8），允许更长的上下文或更高的并发</li>
<li>需要打 patch：上游 vLLM 对 INT8 KV 的支持还在开发中，需要使用 fork 或打 patch</li>
<li>推荐方案：对于双 3090 TP=2 部署 Qwen3.8-27B，可以考虑 W8A8 权重 + INT8 KV cache 的组合</li>
</ol>
<p dir="auto">落地建议：</p>
<ul>
<li>短期：关注 vLLM PR #40391 合入后可直接使用 --kv-cache-dtype int8</li>
<li>中期：使用 yhavinga 的 fork 测试 INT8 PTH</li>
<li>长期：等待上游正式支持后一键切换</li>
</ul>
]]></description><link>https://lcz.me/post/18805</link><guid isPermaLink="true">https://lcz.me/post/18805</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Thu, 17 Sep 2026 06:28:40 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 01:42:43 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/joker_chang" aria-label="Profile: joker_chang">@<bdi>joker_chang</bdi></a> 我倡议的是一个比较全面的 针对本地LLM的综合能力测试，用时估计 在40分钟以上（ 毕竟现在QWEN 3.8写个鹈鹕骑自行车都要30分钟+了） ，尽量一段话，一张图就能展示 自己当前配置的真实情况。</p>
]]></description><link>https://lcz.me/post/18758</link><guid isPermaLink="true">https://lcz.me/post/18758</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Thu, 17 Sep 2026 01:42:43 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 00:56:18 GMT]]></title><description><![CDATA[<p dir="auto">追加一个短任务的日志</p>
<p dir="auto">[34m1512.21.100.458[0m [32mI [0mslot get_availabl: id  0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.960 (&gt; 0.100 thold), f_keep = 0.978<br />
[34m1512.21.102.469[0m [32mI [0mslot launch_slot_: id  0 | task 179867 | processing task, is_child = 0<br />
[34m1512.25.007.360[0m [32mI [0mslot print_timing: id  0 | task 179867 | prompt processing, n_tokens =   2772, progress = 0.99, t =   3.60 s / 769.60 tokens per second<br />
[34m1512.25.854.391[0m [32mI [0mslot print_timing: id  0 | task 179867 | prompt processing, n_tokens =   3284, progress = 1.00, t =   4.13 s / 794.77 tokens per second<br />
[34m1512.29.124.872[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =    157, tg =  51.62 t/s, tg_3s =  51.94 t/s<br />
[34m1512.32.140.736[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =    337, tg =  55.65 t/s, tg_3s =  59.68 t/s<br />
[34m1512.35.187.308[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =    517, tg =  56.80 t/s, tg_3s =  59.08 t/s<br />
[34m1512.38.195.238[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =    697, tg =  57.56 t/s, tg_3s =  59.84 t/s<br />
[34m1512.41.221.698[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =    870, tg =  57.48 t/s, tg_3s =  57.16 t/s<br />
[34m1512.44.257.726[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   1050, tg =  57.78 t/s, tg_3s =  59.29 t/s<br />
[34m1512.47.291.415[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   1230, tg =  58.00 t/s, tg_3s =  59.33 t/s<br />
[34m1512.50.327.965[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   1410, tg =  58.16 t/s, tg_3s =  59.28 t/s<br />
[34m1512.53.372.754[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   1590, tg =  58.27 t/s, tg_3s =  59.12 t/s<br />
[34m1512.56.385.800[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   1770, tg =  58.42 t/s, tg_3s =  59.74 t/s<br />
[34m1512.59.434.598[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   1950, tg =  58.47 t/s, tg_3s =  59.04 t/s<br />
[34m1513.02.481.957[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   2128, tg =  58.47 t/s, tg_3s =  58.41 t/s<br />
[34m1513.05.513.377[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   2301, tg =  58.36 t/s, tg_3s =  57.07 t/s<br />
[34m1513.08.523.772[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   2478, tg =  58.39 t/s, tg_3s =  58.80 t/s<br />
[34m1513.11.541.732[0m [32mI [0mslot print_timing: id  0 | task 179867 | n_gen =   2655, tg =  58.41 t/s, tg_3s =  58.65 t/s<br />
[34m1513.14.620.649[0m [32mI [0mslot print_timing: id  0 | task 179867 | prompt eval time =    4999.56 ms /  3288 tokens (    1.52 ms per token,   657.66 tokens per second)<br />
[34m1513.14.620.655[0m [32mI [0mslot print_timing: id  0 | task 179867 |        eval time =   48518.05 ms /  2830 tokens (   17.15 ms per token,    58.31 tokens per second)<br />
[34m1513.14.620.656[0m [32mI [0mslot print_timing: id  0 | task 179867 |       total time =   53517.61 ms /  6118 tokens<br />
[34m1513.14.620.658[0m [32mI [0mslot print_timing: id  0 | task 179867 |    graphs reused =     168461<br />
[34m1513.14.620.678[0m [32mI [0mslot print_timing: id  0 | task 179867 | draft acceptance = 0.98375 ( 1877 accepted /  1908 generated), mean len =  2.97<br />
[34m1513.14.625.043[0m [32mI [0mslot      release: id  0 | task 179867 | stop processing: n_tokens = 82966, truncated = 0</p>
]]></description><link>https://lcz.me/post/18744</link><guid isPermaLink="true">https://lcz.me/post/18744</guid><dc:creator><![CDATA[joker_chang]]></dc:creator><pubDate>Thu, 17 Sep 2026 00:56:18 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Thu, 17 Sep 2026 00:54:54 GMT]]></title><description><![CDATA[<p dir="auto">我没有看懂楼主要测试什么？<br />
我的系统是：<br />
windows10操作系统、<br />
“Cuda0 ： NVIDIA GeForce RTX 3090 Ti   WDDM（24G显存）” + “Cuda 1 ： NVIDIA GeForce RTX 3060（12G显存）”、<br />
96G物理内存、<br />
Intel Xeon CPU E5-2680v4。<br />
其中，3060跑 Qwen3-VL-8B，3090跑Qwen3.8-27B-q4_k_m</p>
<p dir="auto">楼主是计划用<a href="https://github.com/suncityldp/zx-bench%E8%B7%91%E5%88%86%E5%90%97%EF%BC%9F" rel="nofollow ugc">https://github.com/suncityldp/zx-bench跑分吗？</a><br />
那基准模型准备用哪个？用llama.cpp？</p>
<p dir="auto">【[34m1511.02.675.371[0m [32mI [0mslot get_availabl: id  0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.609 (&gt; 0.100 thold), f_keep = 1.000<br />
[34m1511.02.677.275[0m [32mI [0mslot launch_slot_: id  0 | task 179136 | processing task, is_child = 0<br />
[34m1511.07.156.058[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =   4096, progress = 0.66, t =   3.87 s / 1058.37 tokens per second<br />
[34m1511.09.350.049[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =   6144, progress = 0.69, t =   6.06 s / 1014.27 tokens per second<br />
[34m1511.11.576.393[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =   8192, progress = 0.72, t =   8.27 s / 990.22 tokens per second<br />
[34m1511.13.833.318[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  10240, progress = 0.75, t =  10.52 s / 973.22 tokens per second<br />
[34m1511.16.124.567[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  12288, progress = 0.77, t =  12.80 s / 959.87 tokens per second<br />
[34m1511.18.445.539[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  14336, progress = 0.80, t =  15.11 s / 948.52 tokens per second<br />
[34m1511.20.803.575[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  16384, progress = 0.83, t =  17.46 s / 938.41 tokens per second<br />
[34m1511.23.191.536[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  18432, progress = 0.85, t =  19.84 s / 929.01 tokens per second<br />
[34m1511.25.611.080[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  20480, progress = 0.88, t =  22.25 s / 920.44 tokens per second<br />
[34m1511.28.061.603[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  22528, progress = 0.91, t =  24.69 s / 912.36 tokens per second<br />
[34m1511.30.544.094[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  24576, progress = 0.94, t =  27.17 s / 904.64 tokens per second<br />
[34m1511.33.061.573[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  26624, progress = 0.96, t =  29.67 s / 897.37 tokens per second<br />
[34m1511.35.609.548[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  28672, progress = 0.99, t =  32.21 s / 890.10 tokens per second<br />
[34m1511.35.838.717[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  28839, progress = 0.99, t =  32.96 s / 874.96 tokens per second<br />
[34m1511.36.659.693[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt processing, n_tokens =  29351, progress = 1.00, t =  33.38 s / 879.32 tokens per second<br />
[34m1511.39.916.632[0m [32mI [0mslot print_timing: id  0 | task 179136 | n_gen =    162, tg =  53.38 t/s, tg_3s =  53.70 t/s<br />
[34m1511.39.966.643[0m [32mI [0mslot print_timing: id  0 | task 179136 | prompt eval time =   34222.77 ms / 29355 tokens (    1.17 ms per token,   857.76 tokens per second)<br />
[34m1511.39.966.649[0m [32mI [0mslot print_timing: id  0 | task 179136 |        eval time =    3066.01 ms /   164 tokens (   18.81 ms per token,    53.16 tokens per second)<br />
[34m1511.39.966.688[0m [32mI [0mslot print_timing: id  0 | task 179136 |       total time =   37288.78 ms / 29519 tokens<br />
[34m1511.39.966.691[0m [32mI [0mslot print_timing: id  0 | task 179136 |    graphs reused =     166881<br />
[34m1511.39.966.698[0m [32mI [0mslot print_timing: id  0 | task 179136 | draft acceptance = 0.80159 (  101 accepted /   126 generated), mean len =  2.60<br />
[34m1511.39.970.413[0m [32mI [0mslot      release: id  0 | task 179136 | stop processing: n_tokens = 75261, truncated = 0<br />
[34m1511.42.636.940[0m [32mI [0mslot get_availabl: id  0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.979 (&gt; 0.100 thold), f_keep = 0.999<br />
[34m1511.42.638.860[0m [32mI [0mslot launch_slot_: id  0 | task 179217 | processing task, is_child = 0<br />
[34m1511.48.703.820[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =    165, tg =  54.10 t/s, tg_3s =  54.42 t/s<br />
[34m1511.51.747.295[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =    348, tg =  57.12 t/s, tg_3s =  60.13 t/s<br />
[34m1511.54.754.268[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =    524, tg =  57.59 t/s, tg_3s =  58.53 t/s<br />
[34m1511.57.797.304[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =    704, tg =  57.98 t/s, tg_3s =  59.15 t/s<br />
[34m1512.00.816.266[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =    855, tg =  56.39 t/s, tg_3s =  50.02 t/s<br />
[34m1512.03.855.881[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =   1033, tg =  56.75 t/s, tg_3s =  58.56 t/s<br />
[34m1512.06.878.310[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =   1195, tg =  56.30 t/s, tg_3s =  53.60 t/s<br />
[34m1512.09.892.918[0m [32mI [0mslot print_timing: id  0 | task 179217 | n_gen =   1356, tg =  55.94 t/s, tg_3s =  53.41 t/s】这个数据你觉得有用吗？即时在跑的日志。</p>
<p dir="auto">llama-server的参数：<br />
【<br />
--chat-template-file "%CHAT_TEMPLATE_FILE%" ^<br />
--host 0.0.0.0 ^<br />
--port 3527 ^<br />
--reasoning off ^<br />
--n-gpu-layers -1 ^<br />
--ctx-size 196608 ^<br />
--flash-attn on ^<br />
--cache-type-k q4_0 ^<br />
--cache-type-v q4_0 ^<br />
--spec-type draft-mtp ^<br />
--spec-draft-n-max 2 ^<br />
--spec-draft-n-min 1 ^<br />
--temp 0.7 ^<br />
--parallel 1 ^<br />
--kv-unified ^<br />
--jinja<br />
】</p>
]]></description><link>https://lcz.me/post/18743</link><guid isPermaLink="true">https://lcz.me/post/18743</guid><dc:creator><![CDATA[joker_chang]]></dc:creator><pubDate>Thu, 17 Sep 2026 00:54:54 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Wed, 16 Sep 2026 10:18:58 GMT]]></title><description><![CDATA[<p dir="auto">AMD 显卡飘过</p>
]]></description><link>https://lcz.me/post/18592</link><guid isPermaLink="true">https://lcz.me/post/18592</guid><dc:creator><![CDATA[jatwu]]></dc:creator><pubDate>Wed, 16 Sep 2026 10:18:58 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Wed, 16 Sep 2026 10:12:34 GMT]]></title><description><![CDATA[<p dir="auto">非常好的方式，大家可以积极回帖，有好的我会开分枝，作为长期活动，本帖置顶一周！<br />
建议你改的简单一点，不要投票，选择一个简单的测试方式，然后AI一个脚本就能跑。</p>
]]></description><link>https://lcz.me/post/18589</link><guid isPermaLink="true">https://lcz.me/post/18589</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 16 Sep 2026 10:12:34 GMT</pubDate></item><item><title><![CDATA[Reply to 【倡议】3090 单/双/多卡联合评测帖（2026年9-10月）—— benchmark与配置/模型对比 on Wed, 16 Sep 2026 10:01:58 GMT]]></title><description><![CDATA[<p dir="auto">倡议方向支持，但有个硬伤先修：截止日期写的「下周一 9月7日」已经过去了（今天 9/16），投票窗口要么改期要么重发，不然没人能参加。</p>
<p dir="auto">要让跨机跑分真的可比，建议先把变量表钉死再投票，否则投完套件还是抄不了作业：</p>
<ul>
<li>套件 + 精确版本/commit（只写 zx-bench 不够，它换一版 prompt 集数字就变）；</li>
<li>pp（prefill）和 tg（decode）分开报，标注 prompt 长度与生成长度，单流一个 t/s 没有信息量；</li>
<li>量化格式（Q4KM / AWQ / INT4）、KV cache dtype + 上下文长度、并发数、是否开投机解码（MTP/DFlash）；</li>
<li>硬件附 <code>nvidia-smi topo -m</code>、NVLink 有无及 <code>topo -p2p r</code>、驱动/CUDA、功耗墙；</li>
<li>3 小时/套太长，建议短协议打底（llama-bench 固定 -p/-n/-r）+ 一个真实 agent 任务，两者都交。</li>
</ul>
<p dir="auto">另一个方法学点：附则按「稠密 FP16 算力 ÷ 71」折算票数，对推理场景偏了——decode 是显存带宽瓶颈，不是 FP16 算力瓶颈。3090 是 936 GB/s、3080 20G 约 760 GB/s、7900XTX 960 GB/s、R9700 640 GB/s，按算力折票会把带宽强、算力弱的卡低估。建议权重表同时列带宽，或干脆一卡一票 + 按显存分档（24G / 32G / 48G+）。</p>
<p dir="auto">我可以按这套模板交一份 3080 20G / 7900XTX 的对照数据进来。</p>
]]></description><link>https://lcz.me/post/18587</link><guid isPermaLink="true">https://lcz.me/post/18587</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Wed, 16 Sep 2026 10:01:58 GMT</pubDate></item></channel></rss>