<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[llama-benchy RTX3090x2 PCiX-SLi 跑 pp2048=1267 t/s, tg32=109 token/s]]></title><description><![CDATA[<p dir="auto">使用 Llama-Benchy 对 vLLM tp 2 进行测试. MTP , ctx 128k:<br />
<img src="https://upload.lcz.me/uploads/5a71a356-6f4a-4d26-aea4-40b681a578be.jpeg" alt="a28b047d-d384-4c17-a651-176f79ee94fd-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/8f87c6b6-2ce5-4a4e-9a46-8053f45a36e1.jpeg" alt="93246232-8267-4f7c-b1a1-b39932a88a87-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/b71ea77b-eed7-415a-8b40-53b837499994.jpeg" alt="0ee2d82f-addc-4cee-8cdb-28d3051bbdc8-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">没想到 Qwen3.6 27B Int4 能达到这样的运行速度。还没买 NvLink , 用PCIx 3.0x8</p>
<p dir="auto">vLLM 主要参数<br />
--max-model-len 131072<br />
--tensor-parallel-size 2<br />
--kv-cache-dtype fp8_e5m2<br />
--max-num-batched-tokens 4128<br />
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'</p>
<p dir="auto">大家都用什么Bench测试？</p>
]]></description><link>https://lcz.me/topic/372/llama-benchy-rtx3090x2-pcix-sli-跑-pp2048-1267-t-s-tg32-109-token-s</link><generator>RSS for Node</generator><lastBuildDate>Mon, 27 Jul 2026 00:37:04 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/372.rss" rel="self" type="application/rss+xml"/><pubDate>Sun, 31 May 2026 06:46:13 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to llama-benchy RTX3090x2 PCiX-SLi 跑 pp2048=1267 t/s, tg32=109 token/s on Tue, 14 Jul 2026 09:08:37 GMT]]></title><description><![CDATA[<p dir="auto">你主板和cpu是什么，x99 没有这种速度</p>
]]></description><link>https://lcz.me/post/9872</link><guid isPermaLink="true">https://lcz.me/post/9872</guid><dc:creator><![CDATA[renyi]]></dc:creator><pubDate>Tue, 14 Jul 2026 09:08:37 GMT</pubDate></item><item><title><![CDATA[Reply to llama-benchy RTX3090x2 PCiX-SLi 跑 pp2048=1267 t/s, tg32=109 token/s on Sun, 31 May 2026 13:54:05 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/aresroc" aria-label="Profile: AresROC">@<bdi>AresROC</bdi></a> 差不多， 我单卡3090是60t/s跑mtp</p>
]]></description><link>https://lcz.me/post/4460</link><guid isPermaLink="true">https://lcz.me/post/4460</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Sun, 31 May 2026 13:54:05 GMT</pubDate></item><item><title><![CDATA[Reply to llama-benchy RTX3090x2 PCiX-SLi 跑 pp2048=1267 t/s, tg32=109 token/s on Sun, 31 May 2026 08:14:15 GMT]]></title><description><![CDATA[<p dir="auto">tg 速度不错啊. 加 --depth 参数再看看, 设成 20k 的样子, 模拟下 hermes 首次会话的 token数量, 也看看ttft.<br />
我猜,总体体验应该是不错的.</p>
]]></description><link>https://lcz.me/post/4425</link><guid isPermaLink="true">https://lcz.me/post/4425</guid><dc:creator><![CDATA[laobenxiong]]></dc:creator><pubDate>Sun, 31 May 2026 08:14:15 GMT</pubDate></item></channel></rss>