<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想]]></title><description><![CDATA[<p dir="auto">Deepseek V4 Flash 0731 Q4量化版，终于本地跑起来了。速度当然惨不忍睹，也就10-11tok/s——但回想来时路,还是有点感慨的。这台 6 年的老电脑,缝缝补补能到这一步,挺满足了。（<strong>最新情况， 到 15tok/s 了， 看 17楼帖子</strong>）</p>
<p dir="auto"><strong>1. 本来的样子</strong></p>
<p dir="auto">CPU        : AMD Ryzen 9 3950X<br />
内存       : 32 GB DDR4<br />
显卡 : ASUS RTX 2060 SUPER 8G<br />
主板 : ASUS X570 TUF Gaming</p>
<p dir="auto">这种配置当年装出来其实已经过剩,日常写代码、打游戏绰绰有余,完全没必要折腾。</p>
<p dir="auto"><strong>2. 沉寂的那些年</strong></p>
<p dir="auto">前两年有次朋友给了块 3070 8G,我顺手把 2060S 换掉了——机器对我来说还是"够用就行",从来没想着升级配置。</p>
<p dir="auto">存储倒是买过几根——存储最便宜那阵子加了一条 32G,凑成 64G;再买了 2T、4T 的 SSD。那会儿 2T 才 500 多、4T 也就 1000 左右,买的时候还嫌贵。现在想想,幸亏那时候的自己——没在便宜的时候囤货这件事上犹豫,后面才知道什么叫真贵。</p>
<p dir="auto"><strong>3. 去年:开始关注大模型</strong></p>
<p dir="auto">DeepSeek 出圈那会儿,我开始正儿八经研究大模型。早期本地能跑的那些小模型,玩了一圈觉得意思不大， 后面发现Qwen3系列模型不错， 3070 能跑起来像样的东西了,才看见希望。</p>
<p dir="auto">心里想升级,行动上又拖了下来。</p>
<p dir="auto"><strong>4. 今年:开始加速</strong></p>
<p dir="auto">年初开始玩龙虾,进一步意识到本地模型的价值。Qwen3.5 一波小尺寸又好用的模型出来后,我下定决心——加一块 3090。</p>
<p dir="auto">拿回来一跑,真的豁然开朗:很多以前跑不动的大模型都能跑起来了。那一刻甚至有点飘,觉得自己已经"无所不能"——直到碰到 DeepSeek。才发现 DeepSeek 是一道真正的鸿沟。</p>
<p dir="auto">网上搜了无数教程,基本只有两类:要么 M3 Ultra 大内存一体机,要么好几张 RTX 6000 Pro。看着别人的配置,真有一种望洋兴叹的无力感。</p>
<p dir="auto">后来 Qwen3.8 Flash Next 出来,3070 拖了它大腿,于是直接换成了 3080 20G,内存加到 96G。</p>
<p dir="auto">升级完也就爽了两星期——立刻发现两个新问题:</p>
<ul>
<li>3080+3090 不对称, 两张卡基本各跑各的</li>
<li>96G 内存单双通道混插,带宽掉得很厉害,影响所有 batch / offload 类的参数</li>
</ul>
<p dir="auto">这些问题虽然烦,至少 Qwen3.8 Flash Next 是跑起来了。DeepSeek V4 Flash Q4 也第一次看到了"希望"——虽然最后还是 OOM。</p>
<p dir="auto">于是决定:加到 128G 内存,3080 换成 3090。顺便整顿了机箱风道,装了几个 F9 R120 风扇,真的吹得很爽。</p>
<p dir="auto"><strong>5. 现在的样子</strong></p>
<p dir="auto">CPU        : AMD Ryzen 9 3950X<br />
内存       : 128 GB DDR4 3200<br />
显卡       : ASUS RTX 3090 24G × 2<br />
主板       : ASUS X570 TUF Gaming(原地没动)</p>
<p dir="auto">前前后后所有升级加一起,不到 1.6W RMB。<br />
除了能跑大模型， 我喜欢的游戏，现在也能用3090支持 DLSS5 和多帧生成了~老游戏焕发第二春，也挺开心的，一不小心就玩了一整天。</p>
<p dir="auto"><strong>6. 关于主板</strong></p>
<p dir="auto">但要说说这块主板。</p>
<p dir="auto">搞了那么多升级,最后才发现——这块主板只支持 PCIE 4.0 16X + 4X,跑不了 8X + 8X,也不能 P2P。<br />
真是有点如鲠在喉。</p>
<p dir="auto">下一步,只能换主板。但至少 DeepSeek V4 Flash 0731 在这台机器上真的跑起来了——也算完成了当初那个小小的愿望吧，不知道有没有大神能在指导指导，让我跑的更快点。</p>
<p dir="auto"><strong>后记</strong></p>
<p dir="auto">这机器,陪我 6 年，稀里糊涂各种很拮据的升级，虽然没钱，但继续升级只是因为不愿意放弃。</p>
<p dir="auto">下一步折腾什么?谁知道呢。折腾无止境,后面继续折腾。</p>
]]></description><link>https://lcz.me/topic/1836</link><generator>RSS for Node</generator><lastBuildDate>Fri, 25 Sep 2026 20:36:03 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1836.rss" rel="self" type="application/rss+xml"/><pubDate>Sun, 20 Sep 2026 06:13:56 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Fri, 25 Sep 2026 02:29:33 GMT]]></title><description><![CDATA[<p dir="auto">我也是DDR43500,我跑QWEN3.8 FLASH NEXT初始速度是25TS，可惜我只有80G RAM，跑不了DS V4了</p>
]]></description><link>https://lcz.me/post/20609</link><guid isPermaLink="true">https://lcz.me/post/20609</guid><dc:creator><![CDATA[vosrock]]></dc:creator><pubDate>Fri, 25 Sep 2026 02:29:33 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Fri, 25 Sep 2026 02:01:08 GMT]]></title><description><![CDATA[<p dir="auto">牛呀！我赶紧去抱大腿。我只有32G显存，不知道能不能抱到住</p>
]]></description><link>https://lcz.me/post/20600</link><guid isPermaLink="true">https://lcz.me/post/20600</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Fri, 25 Sep 2026 02:01:08 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Fri, 25 Sep 2026 00:24:07 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> 现在 GLM5.3-flash 也能跑到 19 tok/s 了。 <a href="https://lcz.me/topic/1929">https://lcz.me/topic/1929</a></p>
]]></description><link>https://lcz.me/post/20579</link><guid isPermaLink="true">https://lcz.me/post/20579</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Fri, 25 Sep 2026 00:24:07 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Wed, 23 Sep 2026 03:30:43 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/johnnybegood" aria-label="Profile: johnnybegood">@<bdi>johnnybegood</bdi></a> 老机器新生命 真棒！</p>
]]></description><link>https://lcz.me/post/20205</link><guid isPermaLink="true">https://lcz.me/post/20205</guid><dc:creator><![CDATA[David Wang]]></dc:creator><pubDate>Wed, 23 Sep 2026 03:30:43 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Wed, 23 Sep 2026 02:34:13 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> 有两个问题， 据说exllamav3只支持nvidia， 不知道你是nvidia的卡么？  第二， offload 到内存是肯定支持的， 要不我也跑不起来啊， 内存里面我放了 30层 moe呢。</p>
]]></description><link>https://lcz.me/post/20185</link><guid isPermaLink="true">https://lcz.me/post/20185</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Wed, 23 Sep 2026 02:34:13 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Wed, 23 Sep 2026 02:30:55 GMT]]></title><description><![CDATA[<p dir="auto">试了，120 GB 装不进 我的机器 32 GB显存+128G内存，exl3 不支持 offload。你的48G显存立功了。另：祝贺，这些速度都到了17了，其实可用了，不着急的任务，或者晚上睡觉前丢给它让它跑。</p>
]]></description><link>https://lcz.me/post/20183</link><guid isPermaLink="true">https://lcz.me/post/20183</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Wed, 23 Sep 2026 02:30:55 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Wed, 23 Sep 2026 02:18:32 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> 刚刚又试了一下， 短暂地将CPU超频了 30%，  跑起来了， 速度还真涨了接近 20% ...</p>
<p dir="auto">完整对比：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>threads</th>
<th>avg decode</th>
<th>avg e2e</th>
<th>备注</th>
</tr>
</thead>
<tbody>
<tr>
<td>16</td>
<td>15.73 t/s</td>
<td>12.47</td>
<td>物理核满载</td>
</tr>
<tr>
<td>28</td>
<td>17.44 t/s</td>
<td>14.77</td>
<td>14 物理核 × 2 SMT 部分加速</td>
</tr>
<tr>
<td>32</td>
<td>7.71 t/s</td>
<td>3.94</td>
<td>全 SMT 兄弟抢资源，掉速 50%</td>
</tr>
</tbody>
</table>
<p dir="auto">结论：threads=28 是最优（14 物理核 × 2 SMT，让部分兄弟核帮忙分摊访存延迟，但不全开）。</p>
]]></description><link>https://lcz.me/post/20176</link><guid isPermaLink="true">https://lcz.me/post/20176</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Wed, 23 Sep 2026 02:18:32 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Wed, 23 Sep 2026 01:58:43 GMT]]></title><description><![CDATA[<p dir="auto">厉害啊兄弟，我也试试</p>
]]></description><link>https://lcz.me/post/20172</link><guid isPermaLink="true">https://lcz.me/post/20172</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Wed, 23 Sep 2026 01:58:43 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Wed, 23 Sep 2026 01:57:53 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> 我又研究了一下， 跑得时候我发现cpu满载， 看来跟cpu速度也有关， 我是 3950X， 如果cpu是更猛的 285K， 9950X 等等， 估计速度还会进一步提升。 现在看来MTP是用 CPU在跑的。</p>
]]></description><link>https://lcz.me/post/20171</link><guid isPermaLink="true">https://lcz.me/post/20171</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Wed, 23 Sep 2026 01:57:53 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Wed, 23 Sep 2026 01:28:02 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a>  到 15t/s 了！  用的这个模型  <a href="https://huggingface.co/turboderp/DeepSeek-V4-Flash-Vision-Exp-exl3/tree/3.04bpw" rel="nofollow ugc">https://huggingface.co/turboderp/DeepSeek-V4-Flash-Vision-Exp-exl3/tree/3.04bpw</a> ， 特地部署了 exllamaV3 + TabbyAPI ， 跑起来确实快了不少，聊起天来肉眼上还是能接受的比较舒服了， 另外试了一下coding， 应该是比 15tok/s 更快一些，可能超过20-25了。</p>
]]></description><link>https://lcz.me/post/20152</link><guid isPermaLink="true">https://lcz.me/post/20152</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Wed, 23 Sep 2026 01:28:02 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Tue, 22 Sep 2026 01:57:22 GMT]]></title><description><![CDATA[<p dir="auto">听你这么一说，我查了下，我的是 4条 Patriot 32GB Viper Steel DDR4 3600 MHz UDIMM Memory。我一直还以为我的是3200 MHz,哈哈。想当年，这4根内存的价格跟现在四根DDR5的价格真实天差地别。</p>
]]></description><link>https://lcz.me/post/19897</link><guid isPermaLink="true">https://lcz.me/post/19897</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Tue, 22 Sep 2026 01:57:22 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Tue, 22 Sep 2026 01:13:58 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> 不敢啊， 两条3200超频到 3600 还敢用用， 四条怕不稳定啊。。。。什么品牌型号的内存？ 我是海盗船.</p>
<p dir="auto">更新一下： 调了一下内存时序， 超频了一下， 也用了 3600， 但是发现并没有什么帮助， decode也没有更快。 看来内存不是瓶颈。 另外再问一下， 你的 prefill 速度是多少？</p>
]]></description><link>https://lcz.me/post/19882</link><guid isPermaLink="true">https://lcz.me/post/19882</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Tue, 22 Sep 2026 01:13:58 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Mon, 21 Sep 2026 23:23:20 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/johnnybegood" aria-label="Profile: johnnybegood">@<bdi>johnnybegood</bdi></a> 对，有个什么XMP，打开就是可以超到3600</p>
]]></description><link>https://lcz.me/post/19867</link><guid isPermaLink="true">https://lcz.me/post/19867</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Mon, 21 Sep 2026 23:23:20 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Mon, 21 Sep 2026 22:20:10 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> 那看来差不多就这个速度了， 当个知识数据库聊聊天应该没啥问题， 干活儿估计就比较痛苦了。另外你加了Dspark预测编码了么？ 按你说的你现在应该是baseline速度啊， 加了预测说不定能到15</p>
]]></description><link>https://lcz.me/post/19861</link><guid isPermaLink="true">https://lcz.me/post/19861</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Mon, 21 Sep 2026 22:20:10 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Mon, 21 Sep 2026 22:16:40 GMT]]></title><description><![CDATA[<blockquote>
<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> <a href="/post/19666">说</a>:</p>
<p dir="auto">内存和你一模一样 128 GB DDR4 3200</p>
</blockquote>
<p dir="auto">不是3200么？ 测出3600？ 超频用了？</p>
]]></description><link>https://lcz.me/post/19860</link><guid isPermaLink="true">https://lcz.me/post/19860</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Mon, 21 Sep 2026 22:16:40 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Mon, 21 Sep 2026 20:49:40 GMT]]></title><description><![CDATA[<p dir="auto">我刚测过了，i7-12700K + DDR4-3600 双通道: 测试数据如下</p>
<p dir="auto">1 线程:   9.3 GB/s<br />
2 线程:  23.3 GB/s<br />
4 线程:  36.0 GB/s<br />
8 线程:  53.4 GB/s  ← 峰值</p>
<p dir="auto">AI计算过程</p>
<p dir="auto">DSF 10B active params @ ~4bit量化 ≈ 5 GB/token<br />
53.4 GB/s 除以  5GB/token → 11 tok/s 左右 (理论上限)</p>
<p dir="auto">我实测最高到12 tok/s</p>
]]></description><link>https://lcz.me/post/19855</link><guid isPermaLink="true">https://lcz.me/post/19855</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Mon, 21 Sep 2026 20:49:40 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Mon, 21 Sep 2026 04:37:41 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a> 你让ai编个程序实测一下内存带宽。 应该不到50G/s  ，反正我的主板插满四根内存后， 给我降速到了 30G多， 另外瓶颈这个事， 别说30G/s , 其实 20G/s 也基本够， 我最大瓶颈是第二个PCIE 插槽最多只能 X4 ,无法 X8 更不能 x16, 否则速度我觉得应该能翻倍</p>
]]></description><link>https://lcz.me/post/19667</link><guid isPermaLink="true">https://lcz.me/post/19667</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Mon, 21 Sep 2026 04:37:41 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Mon, 21 Sep 2026 04:29:30 GMT]]></title><description><![CDATA[<p dir="auto">我是32G显存N卡，内存和你一模一样 128 GB DDR4 3200，也只能跑出11~12tok/s。问了AI，发现瓶颈在内存交换速度。我的只有53G，差太远了，用这个瓶颈AI算给我说也就是12tok/s</p>
]]></description><link>https://lcz.me/post/19666</link><guid isPermaLink="true">https://lcz.me/post/19666</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Mon, 21 Sep 2026 04:29:30 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Mon, 21 Sep 2026 04:22:33 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/ben-lee" aria-label="Profile: Ben-Lee">@<bdi>Ben-Lee</bdi></a>  我也写了， 非作业， 因为感觉弄得不是特别满意， 只是能跑起来而已。</p>
<p dir="auto">1、用的是 llama.cpp , dspark 加速<br />
2、Q4 用的就是unsloth的原版 UD Q4 KL 模型， 166G<br />
3、几乎占满了， 都是22G多， 多少层没仔细看，因为draft也弄到显存了，PLE表肯定是在内存， 内存几乎用满了，128G一度达到98%<br />
4、10-11tok/s 对应 1024 context, 又跑了几次长程， 基本就是7-9左右了。</p>
<p dir="auto">我也想知道有没有别的方法能继续加速， 能超过15我就很满足了。</p>
]]></description><link>https://lcz.me/post/19663</link><guid isPermaLink="true">https://lcz.me/post/19663</guid><dc:creator><![CDATA[johnnybegood]]></dc:creator><pubDate>Mon, 21 Sep 2026 04:22:33 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Sun, 20 Sep 2026 19:12:31 GMT]]></title><description><![CDATA[<p dir="auto">老哥，恭喜跑起来！想请教几个部署细节：</p>
<ol>
<li>用的是 llama.cpp、DwarfStar 还是其他推理引擎？</li>
<li>Q4 模型具体是哪个 GGUF 版本？文件有多大？</li>
<li>双 3090 分别占用了多少显存？CPU Offload 了多少层？有没有用 --n-cpu-moe？</li>
<li>10–11 tok/s 对应多长的 Context？</li>
</ol>
<p dir="auto">我也想研究这条路线。</p>
]]></description><link>https://lcz.me/post/19607</link><guid isPermaLink="true">https://lcz.me/post/19607</guid><dc:creator><![CDATA[Ben Lee]]></dc:creator><pubDate>Sun, 20 Sep 2026 19:12:31 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Sun, 20 Sep 2026 13:02:37 GMT]]></title><description><![CDATA[<p dir="auto">折腾路线很真实，你这台的下一步其实挺明确，别先急着换主板。</p>
<p dir="auto"><strong>1. 先量再换</strong>：用 <code>llama-bench</code> 把 pp（prefill）和 tg（decode）分开跑，确认 DeepSeek V4 Flash Q4 现在卡在哪一段。10–11 t/s 对这类模型，多半是<strong>权重没全进显存、专家层在 CPU 上跑</strong>，而不是算力不够。</p>
<p dir="auto"><strong>2. 48GB 显存怎么放</strong>：如果是 MoE，用 <code>--n-cpu-moe N</code> 只把一部分专家放 CPU（从多到少试），attention/dense 层留 GPU；纯 dense 大模型 48GB 塞不下就别硬塞，Q4 也塞不下。同时开 <code>--flash-attn</code>，KV 用 <code>--cache-type-k q8_0 --cache-type-v q8_0</code>，能省大约一半 KV 显存，给权重腾地方。</p>
<p dir="auto"><strong>3. 内存先别掉链子</strong>：128G DDR4-3200 在 3950X 上插满 4 条通常会掉到 2933/2666，先 <code>dmidecode -t memory | grep -i "Configured Memory Speed"</code> 看实际频率。掉频的话，把内存稳在 3200（比换主板便宜得多）比换板更值。</p>
<p dir="auto"><strong>4. X570 的 x16+x4 是真瓶颈</strong>：第二条物理 x16 槽只有 CPU x4 走线，双卡各跑各的、做不了 TP。williamlouis 说的 <strong>BIOS 拆 x8/x8</strong> 是首选——先更新到最新 BIOS，再看 Advanced → Onboard Devices Configuration → PCIEX16_1 Bifurcation，改成 X8/X8；但 TUF X570 是否物理支持要看你具体板型和 BIOS，拆不了再谈换板。<strong>别为 P2P 折腾</strong>——X570 同样有 ACS 挡着，双 3090 想要低延迟 all-reduce 只有 NVLink（3-slot 桥，约 ¥2300），换 X99 或 Gen4 板子的收益都不如它；要换就按"先 NVLink、后主板"排。</p>
<p dir="auto"><strong>5. 发热</strong>：220W/卡 + 前下进后上出 + 限功耗方向对，3090 显存 85°C 可接受但别再往上，核心热点温度用 <code>nvidia-smi -q -d TEMPERATURE</code> 看。</p>
<p dir="auto">一句话：先把 V4 Flash 的 pp/tg 测出来、把专家 offload 和 KV 量化调对，再决定换不换板；真要花钱，NVLink 优先。</p>
]]></description><link>https://lcz.me/post/19544</link><guid isPermaLink="true">https://lcz.me/post/19544</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sun, 20 Sep 2026 13:02:37 GMT</pubDate></item><item><title><![CDATA[Reply to 感慨，Deepseek V4 Flash 终于本地跑起来了，非作业，纯感想 on Sun, 20 Sep 2026 10:20:08 GMT]]></title><description><![CDATA[<p dir="auto">对于喜欢折腾的兄弟都赞一个</p>
]]></description><link>https://lcz.me/post/19536</link><guid isPermaLink="true">https://lcz.me/post/19536</guid><dc:creator><![CDATA[xiaopbro]]></dc:creator><pubDate>Sun, 20 Sep 2026 10:20:08 GMT</pubDate></item></channel></rss>