<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  )]]></title><description><![CDATA[<p dir="auto">聽說價格很甜甜甜 ～ Atlas 300I Duo 推理卡, 有論壇大神用過 這張卡了嗎？會屬於折騰卡嗎？</p>
<p dir="auto">内存规格	LPDDR4X 96GB或48GB，总带宽408GB/s<br />
功耗	150W</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/de216989-ada8-4cfd-9960-d874278833ee.jpeg" alt="f0af5be5-8efd-4ffb-b359-86b497bae01f-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/e8da9422-0763-4198-b1f4-6b6bbcaae2f7.jpeg" alt="e776bbb5-982a-4d71-b967-7dcfe8369ba9-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/topic/806/atlas-300i-duo-推理卡-lpddr4x-96gb-功耗150w-带宽408gb-s</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 22:16:53 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/806.rss" rel="self" type="application/rss+xml"/><pubDate>Wed, 08 Jul 2026 11:40:18 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Sat, 11 Jul 2026 17:06:16 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a>  算子分散化 哪怕都是int8也行啊 啥都有点 啥都不行 带宽也不行 它实际上不是一颗芯片连接一块统一的96G显存，而是：310P3芯片0 + 独立显存  310P3芯片1 + 独立显存 310P3的INT8优势主要集中在矩阵乘 不能根据INT8 TOPS简单推算大模型的token生成速度 离开官方图模式 性能会断崖式下降 玩不转啊 所以非华为官方版本 我都懒得试了 估计效果也好不到哪里去</p>
]]></description><link>https://lcz.me/post/9757</link><guid isPermaLink="true">https://lcz.me/post/9757</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Sat, 11 Jul 2026 17:06:16 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Sat, 11 Jul 2026 17:05:20 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> 没办法 真的限制太多了</p>
]]></description><link>https://lcz.me/post/9756</link><guid isPermaLink="true">https://lcz.me/post/9756</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Sat, 11 Jul 2026 17:05:20 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Sat, 11 Jul 2026 15:49:32 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bingqin-wang" aria-label="Profile: bingqin-wang">@<bdi>bingqin-wang</bdi></a> <a href="/post/9733">说</a>:</p>
<p dir="auto">Qwen3.6-27B W8A8</p>
</blockquote>
<p dir="auto">沒用過W8A8格式的, 剛查了HugginFace 約Model Size~29 GB (single safetensors)<br />
假如是29GB, 的確生成速度11 tokens 在合理範圍, 就显存带宽约408 GB/s 速度有限</p>
<p dir="auto">感謝分享 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f642.png?v=1376b21da6c" class="not-responsive emoji emoji-android emoji--slightly_smiling_face" style="height:23px;width:auto;vertical-align:middle" title=":)" alt="🙂" /></p>
]]></description><link>https://lcz.me/post/9745</link><guid isPermaLink="true">https://lcz.me/post/9745</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Sat, 11 Jul 2026 15:49:32 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Sat, 11 Jul 2026 15:02:19 GMT]]></title><description><![CDATA[<p dir="auto">天啊, 11tks 這個速度....</p>
]]></description><link>https://lcz.me/post/9735</link><guid isPermaLink="true">https://lcz.me/post/9735</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Sat, 11 Jul 2026 15:02:19 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Sat, 11 Jul 2026 14:57:53 GMT]]></title><description><![CDATA[<p dir="auto">已经租云机试过了<br />
华为 Atlas 300I Duo 96G<br />
内部芯片：2 × Ascend  310P3<br />
总显存：96G<br />
实际为两块相互独立的设备内存<br />
单芯片可用显存约43G<br />
双芯片总显存带宽约408 GB/s<br />
软件路线：<br />
vLLM-Ascend<br />
CANN<br />
Qwen3.6-27B W8A8<br />
Tensor Parallel = 2<br />
官方图模式<br />
官方融合算子<br />
权重常驻显存<br />
KV Cache常驻显存<br />
最终稳定结果：<br />
生成速度：约11.01 tokens/s 最快15tokens/s<br />
语义测试：5/5<br />
双芯片运行：正常<br />
服务接口：稳定<br />
显存占用：每颗芯片约35～36G<br />
这个结果已经不是“模型没配好”或者“只跑了一个未经优化的 Python 示例”。<br />
官方 W8A8、图捕获、融合算子、常驻权重和双芯片调度都已经启用。</p>
]]></description><link>https://lcz.me/post/9733</link><guid isPermaLink="true">https://lcz.me/post/9733</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Sat, 11 Jul 2026 14:57:53 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Wed, 08 Jul 2026 16:18:32 GMT]]></title><description><![CDATA[<p dir="auto">@kos or 好问题，我来给你具体分析一下 Atlas 300I Duo 跑 DeepSeek-V4-Flash 的方案。</p>
<p dir="auto">先说结论：技术上可行但实用性很低，我不推荐这条路线。</p>
<p dir="auto">【系统配置】</p>
<p dir="auto">主板：X99 双路或 X299 平台（需要 PCIe 3.0 x16 插槽）<br />
CPU：E5-2680 v4 或 i9-10900X（主要跑 Linux 系统，算力需求不大）<br />
内存：32GB DDR4 就够了<br />
存储：512GB SSD<br />
GPU（可选）：加一张亮机卡，因为 Ascend 卡不输出显示<br />
系统：Ubuntu 22.04 + CANN 8.0 + MindSpore Lite</p>
<p dir="auto">【价格估算】</p>
<p dir="auto">Atlas 300I Duo：约 800-1500 元（二手/渠道价）<br />
X99 主板 + E5 + 32G DDR4：约 300-500 元<br />
电源 550W：约 200 元<br />
SSD + 机箱：约 200 元<br />
合计：约 1500-2400 元</p>
<p dir="auto">【推理速度预期】</p>
<p dir="auto">Atlas 300I Duo 是推理卡，不是训练卡，核心问题是：</p>
<ol>
<li>
<p dir="auto">生态问题：DeepSeek-V4-Flash 是 vLLM 生态的模型，而 Atlas 300I 只能用 CANN / MindSpore。目前没有现成的 DeepSeek-V4-Flash 对接 CANN 的方案，你需要自己转 ONNX -&gt; Caffe -&gt; MindSpore，工程量巨大。</p>
</li>
<li>
<p dir="auto">带宽瓶颈：408GB/s 的带宽跑 96GB 内存，对于大模型推理来说 decode 速度会很慢。同样跑一个 70B Q4 模型，4090（1008GB/s）能到 40-50 T/s，Atlas 300I 估计只有 10-15 T/s。</p>
</li>
<li>
<p dir="auto">实际速度：即使你成功跑起来，推理速度大概在 3-8 T/s（取决于模型大小和量化精度），这个速度对于交互式对话来说会很慢。</p>
</li>
</ol>
<p dir="auto">【我的建议】</p>
<p dir="auto">如果你想要 96GB 显存跑大模型，更好的选择是：</p>
<ul>
<li>两张二手 3090（22-24K，约 7000-8000 元）— 生态成熟，速度 45-50 T/s</li>
<li>一张 7900 XTX 24GB（约 3500 元）— ROCm 生态，速度也够</li>
<li>如果真的预算极低且愿意折腾，Atlas 300I Duo 可以买来玩玩，但别指望它能稳定跑 DeepSeek-V4-Flash</li>
</ul>
<p dir="auto">简单说：这张卡更适合跑华为自家的 MindSpore 模型，或者作为固定推理任务的专用卡，不适合跑主流通用模型。</p>
]]></description><link>https://lcz.me/post/9500</link><guid isPermaLink="true">https://lcz.me/post/9500</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Wed, 08 Jul 2026 16:18:32 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Wed, 08 Jul 2026 13:41:01 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> 你能折騰出一套 利用Atlas 300I Duo 跑DeepSeek-V4-Flash 的推論主機嗎？請給我系統規格表 和系統建置價格 和此套系統的推論速度</p>
]]></description><link>https://lcz.me/post/9481</link><guid isPermaLink="true">https://lcz.me/post/9481</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Wed, 08 Jul 2026 13:41:01 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Wed, 08 Jul 2026 13:32:11 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/tony-wang" aria-label="Profile: Tony-Wang">@<bdi>Tony-Wang</bdi></a></p>
<p dir="auto">比Nvidia DGX Spark 快一些, 接近RTX5060Ti 16GB 446 GB/s, 150W 功耗低,<br />
假如放一些小模型+ MTP 速度可能還行, 目前還沒想到實用的運作模式</p>
]]></description><link>https://lcz.me/post/9479</link><guid isPermaLink="true">https://lcz.me/post/9479</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Wed, 08 Jul 2026 13:32:11 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Wed, 08 Jul 2026 13:22:34 GMT]]></title><description><![CDATA[<p dir="auto">@kos or 我来帮你分析一下这张卡。</p>
<p dir="auto"><strong>Atlas 300I Duo 的核心参数</strong></p>
<ul>
<li>架构：华为昇腾（Ascend）</li>
<li>显存：96GB LPDDR4X</li>
<li>带宽：408GB/s</li>
<li>功耗：150W</li>
<li>规格：半高半长，被动散热</li>
</ul>
<p dir="auto"><strong>是不是"折腾卡"？—— 是，而且有三个层面的折腾：</strong></p>
<p dir="auto"><strong>1. 软件生态（最大问题）</strong><br />
这张卡用的是华为的 CANN（异构计算架构），不是 CUDA 也不是 ROCm。主流的 AI 框架（llama.cpp、vLLM、ComfyUI、Hermes Agent）默认不支持 Ascend。你需要用 MindSpore Lite 或者华为提供的 ONNX 适配器来转模型，操作门槛比 NVIDIA 卡高很多。</p>
<p dir="auto"><strong>2. 带宽瓶颈</strong><br />
408GB/s 对比 7900 XTX 的 960GB/s 只有 42%，对比 3090 的 936GB/s 更是差了一倍多。对于 LLM 推理来说，decode 速度直接受带宽限制——同样跑 Qwen3.6-27B Q4，3090 能到 45-50 T/s，这张卡估计只能到 20-25 T/s。大显存适合跑大模型，但带宽会卡脖子。</p>
<p dir="auto"><strong>3. 价格与定位</strong><br />
如果价格确实很便宜（比如几百块），那它最大的价值是 96GB 显存跑大模型推理（单 batch，不要求速度）。但如果你需要跑主流框架、需要生态兼容性、需要一定的速度，那还是建议 NVIDIA 或 AMD 卡。</p>
<p dir="auto"><strong>一句话总结</strong>：如果你愿意折腾 CANN 生态、能接受较低的推理速度、主要跑大模型的单 batch 推理，那这张卡有性价比。否则它就是一张"看着很香、用起来很累"的卡片。</p>
]]></description><link>https://lcz.me/post/9478</link><guid isPermaLink="true">https://lcz.me/post/9478</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Wed, 08 Jul 2026 13:22:34 GMT</pubDate></item><item><title><![CDATA[Reply to Atlas 300I Duo 推理卡 (LPDDR4X 96GB / 功耗150W / 带宽408GB/s  ) on Wed, 08 Jul 2026 11:42:10 GMT]]></title><description><![CDATA[<p dir="auto">感觉带宽低了点儿.</p>
]]></description><link>https://lcz.me/post/9470</link><guid isPermaLink="true">https://lcz.me/post/9470</guid><dc:creator><![CDATA[Tony Wang]]></dc:creator><pubDate>Wed, 08 Jul 2026 11:42:10 GMT</pubDate></item></channel></rss>