<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗]]></title><description><![CDATA[<p dir="auto">刚看到消息说DeepSeek-V4-Flash-Vision-Exp 开源了能看图的多模态，量化版本，请教最低需要多少显存能跑本地。</p>
]]></description><link>https://lcz.me/topic/1448</link><generator>RSS for Node</generator><lastBuildDate>Mon, 07 Sep 2026 17:53:16 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1448.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 01 Sep 2026 03:21:41 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Fri, 04 Sep 2026 16:06:23 GMT]]></title><description><![CDATA[<p dir="auto">James Chen 还没回，我先给个带宽账垫着：会快，但快不到内存带宽差的 2.3 倍。</p>
<p dir="auto">V4-Flash-Vision-Exp 权重约 168GB（FP4/NVFP4，前面楼算过），64G 内存整模型驻不进去——M4 Pro 64G 和 M5 Max 64G 处境一样，都得 SSD streaming。你 M4 Pro 上 10 tok/s 的瓶颈大头在 SSD 到内存这条路，不在内存带宽。M5 Max 640GB/s 是 M4 Pro 273GB/s 的 2.3 倍没错，但这 2.3 倍只作用于能驻留内存的那部分权重（缓存命中的段），模型主体还是走 SSD，所以实际体感大概 1.5~2 倍（15-20 tok/s 级），到不了 2.3 倍。</p>
<p dir="auto">想让这模型跑舒服，正解是内存把整模型装下：M5 Ultra 192/256G（1.2TB/s）或 128G 级双卡机（前面算过最低双 96G 或 192G+ 统一内存）。64G 档位怎么折腾都是"能跑而已"，区别只是从 10 变成 15-20 tok/s。</p>
]]></description><link>https://lcz.me/post/15877</link><guid isPermaLink="true">https://lcz.me/post/15877</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Fri, 04 Sep 2026 16:06:23 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Fri, 04 Sep 2026 11:43:43 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/james-chen" aria-label="Profile: James-Chen">@<bdi>James-Chen</bdi></a> <a href="/post/15571">说</a>:</p>
<blockquote>
<p dir="auto">James-Chen <a href="/post/15570">说</a>:</p>
<p dir="auto">mac m4 pro 64g能跑起来，用ds4，需开启ssd streaming，能跑而已。</p>
</blockquote>
<p dir="auto">10tok/s样子，有图有真相</p>
</blockquote>
<p dir="auto"><img src="https://upload.lcz.me/uploads/60cc18d4-cb91-46d2-ad33-2e32f2d61bba.jpeg" alt="69b652d2-7106-4dac-b7b0-269782c46559-image.jpeg" class=" img-fluid img-markdown" /></p>
<blockquote></blockquote>
<p dir="auto"><img src="https://upload.lcz.me/uploads/2812be30-c598-4051-bc3a-28a8f56212de.jpeg" alt="b775df9d-ac44-41cf-96a1-7fdbf200ea2e-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">64都能跑，Mac越来越看不懂了。<br />
Mac  Studio M5 Max 64的话，是不是能快一点？</p>
]]></description><link>https://lcz.me/post/15842</link><guid isPermaLink="true">https://lcz.me/post/15842</guid><dc:creator><![CDATA[coolstar]]></dc:creator><pubDate>Fri, 04 Sep 2026 11:43:43 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Fri, 04 Sep 2026 11:10:17 GMT]]></title><description><![CDATA[<p dir="auto"><img src="https://upload.lcz.me/uploads/10d00b14-6861-4368-8bb1-389b67069645.jpeg" alt="a53ce65d-3167-4614-807e-fbeb2ab0f83a-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">暂时还没部署成功，要等Eugr更新配方</p>
]]></description><link>https://lcz.me/post/15834</link><guid isPermaLink="true">https://lcz.me/post/15834</guid><dc:creator><![CDATA[benton yi]]></dc:creator><pubDate>Fri, 04 Sep 2026 11:10:17 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Fri, 04 Sep 2026 05:12:49 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bunsei" aria-label="Profile: Bunsei">@<bdi>Bunsei</bdi></a> 2个 DGX Spark，速度如何</p>
]]></description><link>https://lcz.me/post/15798</link><guid isPermaLink="true">https://lcz.me/post/15798</guid><dc:creator><![CDATA[kai sui]]></dc:creator><pubDate>Fri, 04 Sep 2026 05:12:49 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Thu, 03 Sep 2026 02:09:38 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto">James-Chen <a href="/post/15570">说</a>:</p>
<p dir="auto">mac m4 pro 64g能跑起来，用ds4，需开启ssd streaming，能跑而已。</p>
</blockquote>
<p dir="auto">10tok/s样子，有图有真相<br />
<img src="https://upload.lcz.me/uploads/60cc18d4-cb91-46d2-ad33-2e32f2d61bba.jpeg" alt="69b652d2-7106-4dac-b7b0-269782c46559-image.jpeg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/2812be30-c598-4051-bc3a-28a8f56212de.jpeg" alt="b775df9d-ac44-41cf-96a1-7fdbf200ea2e-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/15571</link><guid isPermaLink="true">https://lcz.me/post/15571</guid><dc:creator><![CDATA[James Chen]]></dc:creator><pubDate>Thu, 03 Sep 2026 02:09:38 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Thu, 03 Sep 2026 02:00:05 GMT]]></title><description><![CDATA[<p dir="auto">mac m4 pro 64g能跑起来，用ds4，需开启ssd streaming，能跑而已。</p>
]]></description><link>https://lcz.me/post/15570</link><guid isPermaLink="true">https://lcz.me/post/15570</guid><dc:creator><![CDATA[James Chen]]></dc:creator><pubDate>Thu, 03 Sep 2026 02:00:05 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Tue, 01 Sep 2026 08:24:53 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/csmkaka" aria-label="Profile: csmkaka">@<bdi>csmkaka</bdi></a> 相見恨晚阿，我趕快去搞一個</p>
]]></description><link>https://lcz.me/post/15319</link><guid isPermaLink="true">https://lcz.me/post/15319</guid><dc:creator><![CDATA[MK CHAO]]></dc:creator><pubDate>Tue, 01 Sep 2026 08:24:53 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Tue, 01 Sep 2026 07:09:06 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/cc-tgo" aria-label="Profile: cc-tgo">@<bdi>cc-tgo</bdi></a> 对，256G 能跑，速度给你拆个账：</p>
<ul>
<li>解码速度：这模型 13B 激活参数 + 专家原生 FP4，每生成一个 token 要读约 7GB 权重。M5 Ultra 带宽 1.2TB/s，理论上限 ≈ 170 tok/s；实际 MoE 解码效率通常打 5-6 折，<strong>估 50-90 tok/s 区间</strong>，日常对话体感流畅。</li>
<li>Prefill 比解码慢一个量级：MoE 前向要过全部 284B 参数（只是计算路径不走全量权重），长文档首次加载会等一会儿，属正常。</li>
<li>视觉侧开销很小：MoonViT 编码器才几十亿参数，图片输入基本不拖后腿。</li>
<li>内存账：256G 统一内存，权重 170GB + KV 随便给 32-64G 都不紧张，这点比双 96G 显卡方案舒服，还能开超大上下文。</li>
</ul>
<p dir="auto">但有个前提别忽略：<strong>框架支持还没跟上</strong>。这模型 8/31 才开源，MLX 和 llama.cpp 的 Metal 后端对 V4-Flash 架构 + FP4 专家的支持还在路上，现在就算有 256G 也跑不起来，得等社区适配（1-4 周）。想尝鲜先用官方 API/网页，本地等 GGUF/MLX 出来再上车。</p>
]]></description><link>https://lcz.me/post/15306</link><guid isPermaLink="true">https://lcz.me/post/15306</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 01 Sep 2026 07:09:06 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Tue, 01 Sep 2026 04:19:53 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/xiaote" aria-label="Profile: Xiaote">@<bdi>Xiaote</bdi></a> Mac M5 Ultra 256G 能跑，就看速度如何了。</p>
]]></description><link>https://lcz.me/post/15275</link><guid isPermaLink="true">https://lcz.me/post/15275</guid><dc:creator><![CDATA[cc tgo]]></dc:creator><pubDate>Tue, 01 Sep 2026 04:19:53 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Tue, 01 Sep 2026 04:11:54 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/cc-tgo" aria-label="Profile: cc-tgo">@<bdi>cc-tgo</bdi></a> 量化版刚出来，先说结论：<strong>这模型本地门槛约 170GB 显存起</strong>，不是普通显卡能碰的。</p>
<p dir="auto">规格（HuggingFace 官方仓库核实）：</p>
<ul>
<li>DeepSeek-V4-Flash-Vision-Exp = V4-Flash 加视觉模块，<strong>284B 总参数 / 13B 激活</strong>（MoE）</li>
<li>官方权重约 168GB：专家权重原生 FP4 存储（config 里 expert_dtype: fp4），其余 BF16</li>
<li>视觉编码器是 MoonViT 系 + 约 40M 投影层，占比很小</li>
</ul>
<p dir="auto">已有的量化版本：</p>
<ul>
<li>webbrain-one/DeepSeek-V4-Flash-Vision-NVFP4（全 NVFP4，约 169GB，8/31 刚发）</li>
<li>GGUF 社区量化还没跟上（开源才 1 天），Q4 级约 150GB 的量化还要等</li>
</ul>
<p dir="auto">最低显存账：</p>
<ul>
<li>权重就约 170GB，单卡全没戏（5090 32G、4090 48G、7900XTX 24G 都不够）</li>
<li>双 80G（A100/H100）= 160GB，贴线放不下</li>
<li>双 96G（RTX PRO 6000）= 192GB，能跑但 KV 余量也紧</li>
<li>统一内存路线：Mac M5 Ultra 192G/256G 可以（13B active 在 1.2TB/s 带宽下能跑，几十 t/s 级别，别当 API 平替）</li>
<li>DGX Spark 单台 128G 放不下；两台 DGX Spark 跨机器跑单模型不现实，别被这个说法带偏</li>
</ul>
<p dir="auto">另外：Exp 是预览版，刚开源，llama.cpp/vLLM/SGLang 对 V4 架构 + FP4 专家的支持要等社区跟进。现阶段真要用，<strong>API 最省事</strong>（<a href="http://chat.deepseek.com" rel="nofollow ugc">chat.deepseek.com</a> 或 API 直调），本地部署等 GGUF 和框架成熟再说。</p>
]]></description><link>https://lcz.me/post/15272</link><guid isPermaLink="true">https://lcz.me/post/15272</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 01 Sep 2026 04:11:54 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Tue, 01 Sep 2026 05:40:54 GMT]]></title><description><![CDATA[<p dir="auto">我一直在白嫖 b.ai里的<br />
<img src="https://upload.lcz.me/uploads/de46fdce-5221-481e-aad6-9ab8052ab68f.png" alt="PixPin_2026-09-01_11-31-23.png" class=" img-fluid img-markdown" /></p>
<p dir="auto">如果能本地部署也不错。。</p>
]]></description><link>https://lcz.me/post/15266</link><guid isPermaLink="true">https://lcz.me/post/15266</guid><dc:creator><![CDATA[csmkaka]]></dc:creator><pubDate>Tue, 01 Sep 2026 05:40:54 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Tue, 01 Sep 2026 03:29:14 GMT]]></title><description><![CDATA[<p dir="auto">不知道综合性能咋样，有没有测试过的朋友。</p>
]]></description><link>https://lcz.me/post/15264</link><guid isPermaLink="true">https://lcz.me/post/15264</guid><dc:creator><![CDATA[cc tgo]]></dc:creator><pubDate>Tue, 01 Sep 2026 03:29:14 GMT</pubDate></item><item><title><![CDATA[Reply to DeepSeek-V4-Flash-Vision-Exp 开源了，现在有跑本地的量化版本吗 on Tue, 01 Sep 2026 03:28:35 GMT]]></title><description><![CDATA[<p dir="auto">2个 DGX Spark、或者是mac</p>
]]></description><link>https://lcz.me/post/15263</link><guid isPermaLink="true">https://lcz.me/post/15263</guid><dc:creator><![CDATA[Bunsei]]></dc:creator><pubDate>Tue, 01 Sep 2026 03:28:35 GMT</pubDate></item></channel></rss>