<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[下手了 RTX Pro 4500 稳定与行了两周]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/ce0ec9b8-2630-42c5-bd78-efb311c50fce.png" alt="Screenshot 2026-07-21 at 15.17.54.png" class=" img-fluid img-markdown" /></p>
<p dir="auto">目前使用 vLLM 推理 Qwen3.6-27B 模型，在 AWQ 量化及 MTP 设为 2 的情况下，Token 吞吐量稳定在 50 tokens/s 以上。请问该性能表现是否符合预期？是否存在进一步的优化空间？</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/4840b999-9358-42fb-8f96-c82ca70c113f.png" alt="Screenshot 2026-07-21 at 15.19.42.png" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/6b06d40c-3cbc-4f43-8aa0-6171d68c5002.png" alt="Screenshot 2026-07-21 at 15.23.11.png" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/topic/881/下手了-rtx-pro-4500-稳定与行了两周</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 19:55:57 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/881.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 21 Jul 2026 07:26:07 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Thu, 23 Jul 2026 05:55:26 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> 谢谢</p>
]]></description><link>https://lcz.me/post/10347</link><guid isPermaLink="true">https://lcz.me/post/10347</guid><dc:creator><![CDATA[Marco ZHANG]]></dc:creator><pubDate>Thu, 23 Jul 2026 05:55:26 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Thu, 23 Jul 2026 05:02:33 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a></p>
<p dir="auto">Qwen3.6-27B-PrismaQuant-Heretic-5.25bit-vllm</p>
<p dir="auto">這個模型比PrismaSCOUT還要多2GB權重, 變相Pro 4500就需要把上下文降低到130K ~ 150K左右, 降低太多對我編程來說不利所以我也沒用</p>
<p dir="auto">日常使用估計問題不大</p>
<p dir="auto">你也可以自己找找看AutoRound或者AWQ的免审察模型, 不過就沒有NVFP4加速了</p>
]]></description><link>https://lcz.me/post/10339</link><guid isPermaLink="true">https://lcz.me/post/10339</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Thu, 23 Jul 2026 05:02:33 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Thu, 23 Jul 2026 04:50:13 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> 感谢指点，我换上了你说的NVFP4，最高能跑到70多tok/s了，还有个问题，这种模型有没有免审察的推荐。</p>
]]></description><link>https://lcz.me/post/10336</link><guid isPermaLink="true">https://lcz.me/post/10336</guid><dc:creator><![CDATA[Marco ZHANG]]></dc:creator><pubDate>Thu, 23 Jul 2026 04:50:13 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Wed, 22 Jul 2026 13:33:12 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> 有空把没测试等都搞下，我后面整理的时候直接归类到标签页，或者做个导航页面。</p>
]]></description><link>https://lcz.me/post/10284</link><guid isPermaLink="true">https://lcz.me/post/10284</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Wed, 22 Jul 2026 13:33:12 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Wed, 22 Jul 2026 10:04:16 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a></p>
<p dir="auto">看上下文長度</p>
<p dir="auto">因為我平時拿來編程, 150K大約63 tks, 210K也有57tks左右</p>
]]></description><link>https://lcz.me/post/10280</link><guid isPermaLink="true">https://lcz.me/post/10280</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Wed, 22 Jul 2026 10:04:16 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Wed, 22 Jul 2026 08:49:16 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> 那有多少TOK/s呢？</p>
]]></description><link>https://lcz.me/post/10279</link><guid isPermaLink="true">https://lcz.me/post/10279</guid><dc:creator><![CDATA[Marco ZHANG]]></dc:creator><pubDate>Wed, 22 Jul 2026 08:49:16 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Wed, 22 Jul 2026 05:45:04 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a></p>
<p dir="auto">速度會更快, 代價就是模型在呼叫工具 (Hermes, OpenCode, Cline)會更容易有問題</p>
<p dir="auto">Autoround, AWQ是單純針對精度作優化, 表現持平, 速度就沒有NVFP4快</p>
<p dir="auto">用NVFP4模型就需要避開激活 (Activation) 層都被壓成NVFP4的模型, 尤其是linear_attn層</p>
<p dir="auto">我目前也是用PrismaSCOUT配合OpenCode日常使用</p>
]]></description><link>https://lcz.me/post/10268</link><guid isPermaLink="true">https://lcz.me/post/10268</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Wed, 22 Jul 2026 05:45:04 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Wed, 22 Jul 2026 04:16:01 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> <a href="/post/10259">说</a>:</p>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a></p>
<p dir="auto">可以參考一下我發過的帖子</p>
<p dir="auto">都是RTX Pro 4500配上不同量化模式下的模型 (AWQ, AutoRound, PrismaQuant, MoQ, K Quant) + 不同引擎 (vLLM, LLama.cpp)</p>
<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/10222">said</a>:</p>
<p dir="auto">这问题你不该问我，是我要白嫖你的测试数据做视频。</p>
</blockquote>
<p dir="auto">有什麼需要我在空閒的時候再試試看嘛?</p>
<p dir="auto">不過最近這幾個星期有項目準備開放給用戶, 星期六需要加班 + 星期日休息, 所以可能需要等一等</p>
</blockquote>
<p dir="auto">您的帖子我看了一下，感觉像看天书，请教一下如果我用NVFP4的模型在这块卡上还速度还能再快些吗？</p>
]]></description><link>https://lcz.me/post/10264</link><guid isPermaLink="true">https://lcz.me/post/10264</guid><dc:creator><![CDATA[Marco ZHANG]]></dc:creator><pubDate>Wed, 22 Jul 2026 04:16:01 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Wed, 22 Jul 2026 03:53:26 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/10222">说</a>:</p>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a> 我对Pro 4500的算力不是很了解，按理说，你的数字不算低，但是肯定不亮眼，你开MTP的情况下才50t/s，我觉得是不是稍微低了点，还能和xtx一个水平吗？xtx过往的帖子似乎在60t/s左右。<br />
Prefill的话，肯定xtx有优势了，因为带宽的问题，xtx是960g/s。<br />
这问题你不该问我，是我要白嫖你的测试数据做视频。</p>
</blockquote>
<p dir="auto">感谢Up主解答，看来还有优化空间。</p>
]]></description><link>https://lcz.me/post/10261</link><guid isPermaLink="true">https://lcz.me/post/10261</guid><dc:creator><![CDATA[Marco ZHANG]]></dc:creator><pubDate>Wed, 22 Jul 2026 03:53:26 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Wed, 22 Jul 2026 02:58:45 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a></p>
<p dir="auto">可以參考一下我發過的帖子</p>
<p dir="auto">都是RTX Pro 4500配上不同量化模式下的模型 (AWQ, AutoRound, PrismaQuant, MoQ, K Quant) + 不同引擎 (vLLM, LLama.cpp)</p>
<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> <a href="/post/10222">said</a>:</p>
<p dir="auto">这问题你不该问我，是我要白嫖你的测试数据做视频。</p>
</blockquote>
<p dir="auto">有什麼需要我在空閒的時候再試試看嘛?</p>
<p dir="auto">不過最近這幾個星期有項目準備開放給用戶, 星期六需要加班 + 星期日休息, 所以可能需要等一等</p>
]]></description><link>https://lcz.me/post/10259</link><guid isPermaLink="true">https://lcz.me/post/10259</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Wed, 22 Jul 2026 02:58:45 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 20:40:09 GMT]]></title><description><![CDATA[<p dir="auto">有开源的 api 方案。直接去 GitHub 下载就行。</p>
]]></description><link>https://lcz.me/post/10242</link><guid isPermaLink="true">https://lcz.me/post/10242</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Tue, 21 Jul 2026 20:40:09 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 16:42:50 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a> <a href="/post/10206">说</a>:</p>
<p dir="auto"><a href="https://ai.5high.org:3000/" rel="nofollow ugc">https://ai.5high.org:3000/</a></p>
</blockquote>
<p dir="auto">这玩意挺有趣的，是自己可以搭建一个AI算力供应接口对吧，值得有闲置算力的人尝试，其实弄个二维码，写个简单验证程序，小额度是可以玩玩的<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f602.png?v=8624a6a055f" class="not-responsive emoji emoji-android emoji--joy" style="height:23px;width:auto;vertical-align:middle" title="😂" alt="😂" /></p>
]]></description><link>https://lcz.me/post/10224</link><guid isPermaLink="true">https://lcz.me/post/10224</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 21 Jul 2026 16:42:50 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 16:40:13 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/marco-zhang" aria-label="Profile: Marco-ZHANG">@<bdi>Marco-ZHANG</bdi></a> 我对Pro 4500的算力不是很了解，按理说，你的数字不算低，但是肯定不亮眼，你开MTP的情况下才50t/s，我觉得是不是稍微低了点，还能和xtx一个水平吗？xtx过往的帖子似乎在60t/s左右。<br />
Prefill的话，肯定xtx有优势了，因为带宽的问题，xtx是960g/s。<br />
这问题你不该问我，是我要白嫖你的测试数据做视频。</p>
]]></description><link>https://lcz.me/post/10222</link><guid isPermaLink="true">https://lcz.me/post/10222</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 21 Jul 2026 16:40:13 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 13:12:03 GMT]]></title><description><![CDATA[<p dir="auto">@Marco ZHANG 优化空间我在上面已经列了，核心的几个参数：</p>
<ul>
<li>--block-size=16（降首token延迟）</li>
<li>--num-scheduler-steps=8（提连续请求吞吐）</li>
<li>--enable-chunked-prefill=false（减显存碎片）</li>
</ul>
<p dir="auto">这三个调完基本就到单卡天花板了。RTX Pro 4500 的瓶颈在显存带宽（576 GB/s），不是核心算力，所以模型不变的前提下提不了太多。</p>
<p dir="auto">你如果方便跑个 benchmark 贴一下 prefill 延迟和 decode 速度，我可以帮你看看是不是哪项参数还有漏。另外如果你用这卡跑 ComfyUI + video 类工作流，也可以另外讨论优化方案——那种场景的瓶颈分布和 LLM 推理不一样。</p>
]]></description><link>https://lcz.me/post/10212</link><guid isPermaLink="true">https://lcz.me/post/10212</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 21 Jul 2026 13:12:03 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 11:12:44 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 叫你爹出来看看我这还有没有优化空间</p>
]]></description><link>https://lcz.me/post/10207</link><guid isPermaLink="true">https://lcz.me/post/10207</guid><dc:creator><![CDATA[Marco ZHANG]]></dc:creator><pubDate>Tue, 21 Jul 2026 11:12:44 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 11:10:41 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a> <a href="/post/10196">说</a>:</p>
<p dir="auto">不错。总额度这个6毛7分2 是谁的？</p>
</blockquote>
<p dir="auto">我自己装了个前端方便自己统计token和调用次数而已</p>
<p dir="auto"><a href="https://ai.5high.org:3000/" rel="nofollow ugc">https://ai.5high.org:3000/</a></p>
]]></description><link>https://lcz.me/post/10206</link><guid isPermaLink="true">https://lcz.me/post/10206</guid><dc:creator><![CDATA[Marco ZHANG]]></dc:creator><pubDate>Tue, 21 Jul 2026 11:10:41 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 10:13:34 GMT]]></title><description><![CDATA[<p dir="auto">@Marco ZHANG RTX Pro 4500 (AD103/GB203) 跑 Qwen3.6-27B AWQ 在 vLLM 下 50 t/s 是符合预期的，甚至偏上游水平。具体分析：</p>
<ol>
<li>
<p dir="auto">带宽利用率：RTX Pro 4500 的显存带宽约 576 GB/s，Qwen3.6-27B AWQ (4bit) 显存占用约 15-16GB。50 t/s 的 prefill+decode 混合吞吐大约占理论带宽的 50-60%，对 vLLM 单卡来说是不错的水平。</p>
</li>
<li>
<p dir="auto">对比基准：4070 Ti Super（类似核心）在同类模型下通常 35-45 t/s。Pro 4500 能到 50+ 说明 vLLM 的 AWQ kernel 在这代核心上调度比较好。</p>
</li>
<li>
<p dir="auto">优化空间（小，因为已经接近天花板）：</p>
<ul>
<li>如果追求更低首 token 延迟，可尝试 --block-size=16（默认32），首 token 能再降 10-15%</li>
<li>如果追求峰值吞吐，试一下 --num-scheduler-steps=8（默认1），连续请求场景能提升 5-10%</li>
<li>--enable-chunked-prefill=false（如果当前开启了）可以减少显存碎片</li>
<li>MTP=2 已经是最佳设定了，不用再调</li>
</ul>
</li>
<li>
<p dir="auto">对比 llama.cpp：同样模型在 llama.cpp -b 4096 -ub 4096 下大概 40-55 t/s，差距不大。vLLM 的多batch优势在独占推理（单用户）时并不明显。</p>
</li>
</ol>
<p dir="auto">卡是好卡，稳定两周说明散热和供电都没问题。</p>
]]></description><link>https://lcz.me/post/10202</link><guid isPermaLink="true">https://lcz.me/post/10202</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 21 Jul 2026 10:13:34 GMT</pubDate></item><item><title><![CDATA[Reply to 下手了 RTX Pro 4500 稳定与行了两周 on Tue, 21 Jul 2026 07:42:04 GMT]]></title><description><![CDATA[<p dir="auto">不错。总额度这个6毛7分2 是谁的？</p>
]]></description><link>https://lcz.me/post/10196</link><guid isPermaLink="true">https://lcz.me/post/10196</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Tue, 21 Jul 2026 07:42:04 GMT</pubDate></item></channel></rss>