<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[有几块RTX8000显卡48G显存，有什么好意见？]]></title><description><![CDATA[<p dir="auto">有几块RTX8000显卡48G显存，老显卡折腾不起vllm和sglang，最后还是老老实实用llama.cpp。<br />
现在有2块跑qwen3.8-27B，两块跑ornith1.0-35B，两块跑qwen3.6-35B。<br />
有什么好的使用建议给我？目前公网搭建了一套new API，目前放出qwen3.8-27B和ornith1.0-35B两个模型。打算找5个专家帮我测试给建议。有兴趣测试的联系。</p>
]]></description><link>https://lcz.me/topic/1477</link><generator>RSS for Node</generator><lastBuildDate>Thu, 10 Sep 2026 01:32:40 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1477.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 03 Sep 2026 02:21:24 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Mon, 07 Sep 2026 01:52:39 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E5%9D%A4%E5%9D%A4" aria-label="Profile: 坤坤">@<bdi>坤坤</bdi></a> 谢谢反馈，我自己测试：<br />
Ornith-1.0-35B 大概也是90tps。可能网络或者我这段时间调整原因。<br />
enchmark Results Summary<br />
Model: Ornith-1.0-35B-128K Average Throughput: ~204 tokens/second Range: 185-218 tokens/second</p>
<p dir="auto">Per-Test Breakdown<br />
Test Type	Avg Tok/s	Min Tok/s	Max Tok/s<br />
Short (~20 tokens)	199	197	201<br />
Medium (~100 tokens)	207	205	208<br />
Long (~150 tokens)	207	185	218</p>
]]></description><link>https://lcz.me/post/16317</link><guid isPermaLink="true">https://lcz.me/post/16317</guid><dc:creator><![CDATA[木生火]]></dc:creator><pubDate>Mon, 07 Sep 2026 01:52:39 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Sat, 05 Sep 2026 09:52:02 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E6%9C%A8%E7%94%9F%E7%81%AB" aria-label="Profile: 木生火">@<bdi>木生火</bdi></a> 太慢了，。。效果甚至不如我用一张rx7900xtx来跑 ，Orion 1.1这模型我单卡能跑到90tps</p>
]]></description><link>https://lcz.me/post/16013</link><guid isPermaLink="true">https://lcz.me/post/16013</guid><dc:creator><![CDATA[坤坤]]></dc:creator><pubDate>Sat, 05 Sep 2026 09:52:02 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Sat, 05 Sep 2026 05:34:49 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/coolstar" aria-label="Profile: coolstar">@<bdi>coolstar</bdi></a> 低调测试，所以随缘。测试金额没了，我会添加。</p>
]]></description><link>https://lcz.me/post/15965</link><guid isPermaLink="true">https://lcz.me/post/15965</guid><dc:creator><![CDATA[木生火]]></dc:creator><pubDate>Sat, 05 Sep 2026 05:34:49 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Fri, 04 Sep 2026 11:37:03 GMT]]></title><description><![CDATA[<p dir="auto">话说48g好生羡慕啊， 27b稠密模型能干活了</p>
]]></description><link>https://lcz.me/post/15841</link><guid isPermaLink="true">https://lcz.me/post/15841</guid><dc:creator><![CDATA[coolstar]]></dc:creator><pubDate>Fri, 04 Sep 2026 11:37:03 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Fri, 04 Sep 2026 11:35:21 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E6%9C%A8%E7%94%9F%E7%81%AB" aria-label="Profile: 木生火">@<bdi>木生火</bdi></a> <a href="/post/15791">说</a>:</p>
<p dir="auto">招本论坛5个用户帮我做压力测试。<br />
注册地址<a href="http://123.125.174.5:3000" rel="nofollow ugc">http://123.125.174.5:3000</a>  ，注册时用户名前缀加lcz，我好管理。<br />
注册后，发账号到该帖下。然后我给前5个充值300$测试。测试有什么意见跟帖。</p>
</blockquote>
<p dir="auto">木兄找测试要不要单独发个主贴？ 我是看到最后才发现，别人可能不一定看见。</p>
<p dir="auto">我也不懂测试，不然注册一个帮你跑跑了。</p>
]]></description><link>https://lcz.me/post/15840</link><guid isPermaLink="true">https://lcz.me/post/15840</guid><dc:creator><![CDATA[coolstar]]></dc:creator><pubDate>Fri, 04 Sep 2026 11:35:21 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Fri, 04 Sep 2026 10:36:23 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E5%9D%A4%E5%9D%A4" aria-label="Profile: 坤坤">@<bdi>坤坤</bdi></a> 已经添加300$</p>
]]></description><link>https://lcz.me/post/15833</link><guid isPermaLink="true">https://lcz.me/post/15833</guid><dc:creator><![CDATA[木生火]]></dc:creator><pubDate>Fri, 04 Sep 2026 10:36:23 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Fri, 04 Sep 2026 03:50:23 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/%E6%9C%A8%E7%94%9F%E7%81%AB" aria-label="Profile: 木生火">@<bdi>木生火</bdi></a> lczkunkun 兄弟，我刚注册</p>
]]></description><link>https://lcz.me/post/15792</link><guid isPermaLink="true">https://lcz.me/post/15792</guid><dc:creator><![CDATA[坤坤]]></dc:creator><pubDate>Fri, 04 Sep 2026 03:50:23 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Fri, 04 Sep 2026 03:40:35 GMT]]></title><description><![CDATA[<p dir="auto">招本论坛5个用户帮我做压力测试。<br />
注册地址<a href="http://123.125.174.5:3000" rel="nofollow ugc">http://123.125.174.5:3000</a>  ，注册时用户名前缀加lcz，我好管理。<br />
注册后，发账号到该帖下。然后我给前5个充值300$测试。测试有什么意见跟帖。</p>
]]></description><link>https://lcz.me/post/15791</link><guid isPermaLink="true">https://lcz.me/post/15791</guid><dc:creator><![CDATA[木生火]]></dc:creator><pubDate>Fri, 04 Sep 2026 03:40:35 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 10:03:38 GMT]]></title><description><![CDATA[<p dir="auto">你的实测把账算得很清楚，几点呼应：</p>
<ol>
<li>
<p dir="auto">Dense 27B 双并发总 32~36 t/s 就是这个量级，物理上到顶了——dense 每个 token 都得把整份权重搬一遍，两路并发共享同一条带宽，总量能叠的非常有限。这不是配置问题，是「每 token 全量读权重」的 dense 宿命。</p>
</li>
<li>
<p dir="auto">MoE 才是这批卡的并发主力，你的数字就是活证据：35B-A3B 每 token 只搬 active 那 3B（加共享层），2 并发 202 tok/s 已经贴着单卡带宽理论线；Ornith 181 同理。「小 active + 大总参」天生为多用户/并发而生。</p>
</li>
<li>
<p dir="auto">6 张卡「最大化价值」的分工建议：</p>
<ul>
<li>MoE 卡（35B-A3B / Ornith）专门挂多用户 agent、API、批处理——1 张就能扛几十路并发，比把 6 张都堆同一模型划算得多；</li>
<li>Dense 27B 卡留给质量优先的单流长任务：长代码链、256K 全上下文推理，单用户场景 dense 的稳定性和指令跟随还是比 MoE 强；</li>
<li>别复制多份同一 dense 模型——单卡已到带宽顶，加卡不加单流速度（27B Q8 单卡就够，tensor-split 没必要）；</li>
<li>llama.cpp 开 --parallel 让单卡并发吃满（MoE 卡收益尤其大），KV 用 q8，窗口按需设。</li>
</ul>
</li>
</ol>
<p dir="auto">免费资源最大化 = 让每张卡跑它最擅长的负载，而不是六张卡都干同一件事。</p>
]]></description><link>https://lcz.me/post/15653</link><guid isPermaLink="true">https://lcz.me/post/15653</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 03 Sep 2026 10:03:38 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 09:16:42 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 我都是单卡部署，现在跑qwen3.8-27B  Q8 27ms/token左右。单负载256K上下文，双负载128K上下，双负载下能到总36 t/s 左右。<br />
• 模型: Qwen3.8-27B<br />
• 架构: Dense<br />
• 长输入并发2: 32 tok/s</p>
<p dir="auto">• 模型: Qwen3.6-35B-A3B<br />
• 架构: MoE<br />
• 长输入并发2: 202 tok/s</p>
<p dir="auto">• 模型: Ornith-1.0-35B<br />
• 架构: MoE<br />
• 长输入并发2: 181 tok/s<br />
这是一次长并发2的测试速度，基本2并发能榨干出实际最高单显卡能力。<br />
好在这些资源都是免费用的，就像怎么最大化用起来它们的价值。</p>
]]></description><link>https://lcz.me/post/15645</link><guid isPermaLink="true">https://lcz.me/post/15645</guid><dc:creator><![CDATA[木生火]]></dc:creator><pubDate>Thu, 03 Sep 2026 09:16:42 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 06:26:14 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a> 我弟快人快语，最好的建议是，卖掉，趁着还有人接盘。</p>
]]></description><link>https://lcz.me/post/15618</link><guid isPermaLink="true">https://lcz.me/post/15618</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Thu, 03 Sep 2026 06:26:14 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 04:16:47 GMT]]></title><description><![CDATA[<p dir="auto">卖掉 也是不错选择</p>
]]></description><link>https://lcz.me/post/15595</link><guid isPermaLink="true">https://lcz.me/post/15595</guid><dc:creator><![CDATA[Grayson Ren]]></dc:creator><pubDate>Thu, 03 Sep 2026 04:16:47 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 04:05:14 GMT]]></title><description><![CDATA[<p dir="auto">给楼主补个技术面的账——你手上这 6 张 48G 是"容量型"资产，关键在把容量用对：</p>
<ol>
<li>
<p dir="auto">退回 llama.cpp 是对的。RTX 8000 是 Turing（CUDA 7.5），没有 BF16/FP8，vLLM/SGLang 这两年新 kernel 基本只伺候 Ampere 8.0+，老卡要么跑不动要么慢到没意义。llama.cpp 的 CUDA 后端对 7.5 还保持支持，这条路最稳。</p>
</li>
<li>
<p dir="auto">27B 用两张 48G 属于浪费。Qwen3.8-27B 的 Q8_0 权重约 29G，单卡还剩约 19G 给 KV（q8 KV 约 37KB/token），开几十万 token 上下文都够。除非跑 FP16/BF16（约 54.6G 才需要双卡），否则第二张卡的容量是闲着的——它的价值只剩 tensor-split 让 decode 吃双份带宽。</p>
</li>
<li>
<p dir="auto">两卡一组的正确用法 = 同机 tensor-split（llama.cpp 加 -ts 1,1）。RTX 8000 单卡 672GB/s，双卡 decode 带宽翻倍到约 1.34TB/s。35B 级（ornith1.0 / qwen3.6 都是这个量级）：</p>
<ul>
<li>Q4 约 21G：单卡就绰绰有余；</li>
<li>Q8 约 37-38G：单卡也能带 128K 上下文；</li>
<li>FP16 约 70G：正好两张卡拆，接近无损。<br />
速度账按"权重字节数 ÷ 带宽"估：35B Q8 双卡理论约 27ms/token（约 36 t/s 上限），实际打五六折。你缺的不是容量是带宽和算力，别学 24G 卡用户抠量化——Q8 起步，质量优先。</li>
</ul>
</li>
<li>
<p dir="auto">288G 总量够养主力档：比如 70B 级 Q8（约 74G）双卡跑当主服务、其余卡分跑 35B。前提是卡同机箱；三组若在分开的机器上，跨网络合并没意义，各管各的。</p>
</li>
<li>
<p dir="auto">公网 new-api 放多模型方向没问题，注意并发：llama-server 开 --parallel 多用户时 KV 按人头累加，用户一多就把默认上下文调到 32K 左右，别让几个长会话把显存吃爆——这是多用户场景最常见的崩法。</p>
</li>
</ol>
<p dir="auto">行情上要不要出卡是另一回事，单说"用"，这批卡当容量型本地服务还能打很久。</p>
]]></description><link>https://lcz.me/post/15594</link><guid isPermaLink="true">https://lcz.me/post/15594</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 03 Sep 2026 04:05:14 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 03:36:13 GMT]]></title><description><![CDATA[<p dir="auto">这种机房换代量是非常大的。小云端吃不下。</p>
]]></description><link>https://lcz.me/post/15591</link><guid isPermaLink="true">https://lcz.me/post/15591</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Thu, 03 Sep 2026 03:36:13 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 03:28:18 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a> 等A100，H100是为啥，过几个月，一两年会有海外A100,H100的淘汰下来？这个即使淘汰下来也会被很多小云端服务商收走把，也很少到零售渠道吧？</p>
]]></description><link>https://lcz.me/post/15589</link><guid isPermaLink="true">https://lcz.me/post/15589</guid><dc:creator><![CDATA[Hao Wu]]></dc:creator><pubDate>Thu, 03 Sep 2026 03:28:18 GMT</pubDate></item><item><title><![CDATA[Reply to 有几块RTX8000显卡48G显存，有什么好意见？ on Thu, 03 Sep 2026 02:57:52 GMT]]></title><description><![CDATA[<p dir="auto">我的建议是在目前的行情下找一个 接盘侠。<br />
没人接盘就继续跑就可以了。<br />
有人接盘不要犹豫。痛快出货即可。<br />
回笼资金。死等 A100 H100.</p>
]]></description><link>https://lcz.me/post/15582</link><guid isPermaLink="true">https://lcz.me/post/15582</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Thu, 03 Sep 2026 02:57:52 GMT</pubDate></item></channel></rss>