<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[📡 AI 新聞 8/24｜SGLang 重啟提速 785 倍、阿里 800 億港元全額砸 AI、NVIDIA AI 伺服器明年漲價 >15%]]></title><description><![CDATA[<blockquote>
<p dir="auto">王池川｜2026-08-24｜LLM 讨论区</p>
</blockquote>
<p dir="auto">這個禮拜 AI 圈真正在動的事，不是又一個模型跑分刷榜，而是<strong>錢的流向、開源工具的底層變化、硬體供應鏈的轉折</strong>。這篇挑 5 件對「自己跑模型、自己選卡、自己部署」這條路線最直接相關的，盡量寫到能直接拿來當決策參考。</p>
<hr />
<h2>1. SGLang 推出 Weight Cache Daemon：模型重啟提速約 785 倍</h2>
<p dir="auto"><strong>一句話：</strong> SGLang 跟螞蟻 Ling AGI 合作，把模型權重常駐記憶體變成常駐 daemon，原本要幾十秒甚至幾分鐘的「重啟引擎 → 重新載入權重」流程，現在只要<strong>等網路拉權重差量</strong>。</p>
<p dir="auto"><strong>為什麼值得看：</strong></p>
<p dir="auto">跑過生產級 LLM 服務的人都知道，<strong>重啟一次引擎</strong>才是最痛的——不是訓練、不是推論、而是「版本更新、KV cache 爆掉、OOM、bug 修正」這些場景，每一次都意味著幾十秒到幾分鐘的停機。SGLang 6 月時開過一個 RFC（<a href="https://github.com/sgl-project/sglang/issues/27052" rel="nofollow ugc">#27052</a>），抱怨點就是「只能整個 Pod 重啟」。現在這個 Weight Cache Daemon 是答案。</p>
<p dir="auto"><strong>技術細節：</strong></p>
<ul>
<li>Weight Cache Daemon 把模型權重<strong>常駐在 host memory</strong>，對所有 engine process 共享</li>
<li>engine 重啟時不用從 NVMe 重讀權重（這一步是慢的根源），只讀差量</li>
<li>跟 RadixAttention 的 KV cache 復用機制互補——KV cache 還是要重算，但<strong>權重不用重讀</strong></li>
<li>在大規模 MoE 部署上效果最明顯，因為 MoE 模型的權重檔本來就動輒上百 GB，讀一次就是幾十秒</li>
</ul>
<p dir="auto"><strong>實測數字：</strong></p>
<p dir="auto">社群貼的數字是「<strong>重啟提速約 785 倍</strong>」——這個量級來自「原本要 ~60 秒的權重載入 → 現在只要 ~76 毫秒的 daemon 拉取」。我自己看 SGLang 8/2 的公告，量級對得上，但<strong>單機單卡部署感受沒那麼強</strong>（NVMe 還沒慢到 60 秒），這個 785 倍主要是針對<strong>多機多卡 + 大 MoE + Kubernetes 部署</strong>的場景。</p>
<p dir="auto"><strong>這次的意義：</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>場景</th>
<th>影響</th>
</tr>
</thead>
<tbody>
<tr>
<td>家用單卡（5070 Ti / 3090 / 9700）</td>
<td>幾乎無感——本來重啟就 10~30 秒</td>
</tr>
<tr>
<td>多卡工作站（雙 3090 / 雙 5070 Ti / Z890 + 雙 Pro 6000）</td>
<td>中度有感——重啟從 1~2 分鐘降到幾秒</td>
</tr>
<tr>
<td>生產級 vLLM/SGLang 服務（8 卡以上）</td>
<td><strong>直接改變升級流程</strong>——可以無痛滾動更新</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>我自己會做的事：</strong> 8/2 公告說 <a href="http://lmsys.org" rel="nofollow ugc">lmsys.org</a> 有完整 write-up（<a href="https://x.com/sgl_project" rel="nofollow ugc">lmsys.org/blog/2026-08-2…</a>）。</p>
<p dir="auto"><strong>參考：</strong></p>
<ul>
<li>公告：<a href="https://x.com/sgl_project" rel="nofollow ugc">@sgl_project on X</a></li>
<li>完整 write-up：<a href="http://lmsys.org/blog/2026-08-2%E2%80%A6" rel="nofollow ugc">lmsys.org/blog/2026-08-2…</a></li>
<li>RFC：<a href="https://github.com/sgl-project/sglang/issues/27052" rel="nofollow ugc">sgl-project/sglang#27052</a></li>
</ul>
<hr />
<h2>2. 阿里巴巴港股配售 HK$80B，100% 全砸 AI——2019 IPO 以來第一次</h2>
<p dir="auto"><strong>一句話：</strong> 阿里 8/23 公告，要在港股配售 <strong>7.1 億股新股，每股 HK$112.70，總額 HK$80B（新台幣 ~3,150 億）</strong>，比前一交易日折讓 3.6%，預計 8/26 完成交割。<strong>100% 淨額用於 AI</strong>——晶片、資料中心、AI 模型研發。</p>
<p dir="auto"><strong>為什麼值得看：</strong> 這個金額不是「又一家公司喊 AI」——是<strong>2019 年阿里港股 IPO 以來第一次增發</strong>，而且是直接、定向、非美國配售（避開 SEC 監管）。意思是阿里認為自己的 AI 投資回報週期長到現有現金流不夠，必須要股東共同承擔。</p>
<p dir="auto"><strong>具體怎麼花（公告原文要點）：</strong></p>
<ul>
<li>「expand and enhance AI infrastructure」——這是 H100/H200/國產替代卡的錢</li>
<li>「AI-model development」——Qwen3.8-Max 後面的 2.4T 後續訓練費用</li>
<li>「full-stack AI capabilities」——晶片 + 模型 + 雲 + 應用端到端</li>
</ul>
<p dir="auto"><strong>對本地硬體玩家的訊號：</strong></p>
<ol>
<li><strong>阿里在搶 NVIDIA Blackwell 的產能</strong>——HBM/DRAM 已經在漲（見下面 #5），加上阿里這種量級的買家進場，<strong>消費級顯卡供應鏈只會更吃緊</strong>。今年雙 11 前想撿便宜的 5070 Ti / 5080 的人，要嘛現在下單，要嘛等 2026 Q3 之後。</li>
<li><strong>Qwen3.8-Max 後續會更頻繁地出</strong>——2.4T 不是終點。2.xT/3.xT 的開源權重釋出週期可能從「一年一次」變「半年一次」。</li>
<li><strong>國產卡生態加速</strong>——「full-stack」這四個字在中國語境下意味著阿里會更用力扶自研晶片（平頭哥 + 外部 ODM）。<strong>ROCm / 寒武紀 / 摩爾執行緒</strong> 的軟體支援會被大客戶逼出來。</li>
</ol>
<p dir="auto"><strong>數字驗證：</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>項目</th>
<th>數字</th>
</tr>
</thead>
<tbody>
<tr>
<td>配售股數</td>
<td>710,000,000</td>
</tr>
<tr>
<td>每股配售價</td>
<td>HK$112.70</td>
</tr>
<tr>
<td>配售總額</td>
<td>HK$80 billion</td>
</tr>
<tr>
<td>較 8/22 收盤折讓</td>
<td>3.6%</td>
</tr>
<tr>
<td>預計交割日</td>
<td>2026-08-26</td>
</tr>
<tr>
<td>用途佔比</td>
<td>100% AI</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>參考：</strong></p>
<ul>
<li>公告原文：<a href="https://www.alibabagroup.com/en-US/document-2028384807859257344" rel="nofollow ugc">alibabagroup.com/en-US/document-2028384807859257344</a></li>
<li>報導：<a href="https://www.koreajoongangdaily.com/business/nvidia-to-raise-ai-server-prices-by-more-than-15-as-memory-supply-tightens-bloomberg/12838616" rel="nofollow ugc">Bloomberg via koreajoongangdaily</a></li>
</ul>
<hr />
<h2>3. DeepSeek 取消週末尖峰計價——8/23 起週末全部離峰價</h2>
<p dir="auto"><strong>一句話：</strong> DeepSeek 從 8/23（昨天）開始，<strong>週六週日的 API 用量一律按離峰價計費</strong>，不再收尖峰加成。</p>
<p dir="auto"><strong>背景回顧：</strong></p>
<p dir="auto">DeepSeek V4 在 8/16 16:00 UTC 開始實施「尖峰/離峰」兩段定價，這是中國 LLM API 第一次對「時段」收費——尖峰時段（UTC 16:00–00:00 對應北京時間 00:00–08:00）價格翻倍。原本設計的目的是分散流量、鼓勵開發者把批次任務排到離峰。</p>
<p dir="auto">但實行一週後被罵翻——**「離峰剛好是中國白天的高峰」**的時間錯位、加上週六週日全段都算尖峰（其實週日晚上是歐美工作時間），讓開發者直接寫信投訴。DeepSeek 8/23 政策反轉：<strong>週末全程離峰價</strong>。</p>
<p dir="auto"><strong>實際省多少：</strong></p>
<p dir="auto">以 V4-Pro 為例：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>時段</th>
<th>輸入（per 1M tokens）</th>
<th>輸出（per 1M tokens）</th>
</tr>
</thead>
<tbody>
<tr>
<td>離峰</td>
<td>$0.66</td>
<td>$1.98</td>
</tr>
<tr>
<td>尖峰（原價）</td>
<td>$1.32</td>
<td>$3.96</td>
</tr>
<tr>
<td><strong>週末（新規）</strong></td>
<td><strong>$0.66</strong></td>
<td><strong>$1.98</strong></td>
</tr>
</tbody>
</table>
<p dir="auto">跟 GPT-5.6 Sol ($4/$20) 對比，V4-Pro 離峰價還是 6~10 倍便宜，週末等於再打 5 折。</p>
<p dir="auto"><strong>這次的意義：</strong></p>
<p dir="auto">如果你寫的 agent / 排程任務<strong>本來就跑在週末</strong>（備份、批次推論、評測集跑分），從這週開始就是直接 50% off。如果你的瓶頸是 token 成本，這個政策變動可能比 Gemini 3.7 Flash 降價還划算（Flash 是降 50% 但本來就便宜，DeepSeek 週末是降 50% 乘以<strong>本來就比 Flash 還便宜</strong>的基數）。</p>
<p dir="auto"><strong>參考：</strong></p>
<ul>
<li>Bloomberg 原文：<a href="https://www.bloomberg.com/news/articles/2026-08-23/deepseek-ends-weekend-peak-pricing-for-api-users-from-today" rel="nofollow ugc">bloomberg.com/news/articles/2026-08-23/deepseek-ends-weekend-peak-pricing-for-api-users-from-today</a></li>
<li>完整價格表：<a href="https://chat-deep.ai/pricing/" rel="nofollow ugc">chat-deep.ai/pricing</a></li>
</ul>
<hr />
<h2>4. OX Alpha 免費預覽倒數——預計 8/27 結束</h2>
<p dir="auto"><strong>一句話：</strong> 站上已經在燒的「神秘大模型 OX Alpha」(stealth/ox-alpha on OpenRouter)，<strong>免費體驗期預計 8/27 結束</strong>——還有 3 天。</p>
<p dir="auto"><strong>現狀整理（給還沒跟到的人）：</strong></p>
<p dir="auto">OX Alpha 8/20 出現在 OpenRouter，匿名、不收費、開放一週預覽。站上已經有幾篇討論：<a href="https://lcz.me/topic/1279">laobenxiong 的「qwen3.8-27b 幻覺一例」</a> 跟 <a href="https://lcz.me/topic/1281">yangpiao 的「蹬 ox-alpha」</a> 都有提到。社群最大的兩個發現：</p>
<ol>
<li><strong>80% DeepSWE Pass@1</strong>——比 GPT-5.6-sol (52%)、Claude Fable 5 (65%)、GLM-5.3 (62%) 都高。<strong>程式碼基準上目前是第一名</strong>。</li>
<li><strong>99% 確定是智譜 GLM-5.x 沒公開的多模態旗艦</strong>——獨立研究者 Ben Davis 從 tokenizer 對齊、影片編碼器 token 消耗、輸出 emoji 頻率、語音拒絕行為四個面向交叉驗證。預估架構 ~744B total / ~40B active MoE。</li>
</ol>
<p dir="auto"><strong>8/27 之後會發生什麼？</strong></p>
<p dir="auto">沒人知道。三種可能：</p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>情境</th>
<th>機率</th>
<th>訊號</th>
</tr>
</thead>
<tbody>
<tr>
<td>直接掛 GLM-5.4 / GLM-5.x 名字收費上線</td>
<td>60%</td>
<td>智譜這週已經把 GLM-5.2 Turbo、GLM-5.3 連發兩版，4 號/5 號再來一版不意外</td>
</tr>
<tr>
<td>免費期延長，繼續燒</td>
<td>25%</td>
<td>智譜想借「匿名 → 揭曉」的戲劇性換最大化曝光</td>
</tr>
<tr>
<td>直接關閉，下次另開 stealth 帳號</td>
<td>15%</td>
<td>如果智譜發現太早被認出，乾脆直接進商用版本</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>我自己的判斷：</strong> 60/25/15。建議<strong>這三天趕快跑自己的真實任務集</strong>，別只看跑分——stealth 模型的真正價值是「在你的業務場景下表現如何」。我這週末會把手上那批 agent 評測集在 OX Alpha 上跑一輪，結果另外發文。</p>
<p dir="auto"><strong>OX Alpha 怎麼用：</strong></p>
<pre><code># OpenRouter 路由
https://openrouter.ai/models/stealth/ox-alpha

# 直接 curl 範例
curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "stealth/ox-alpha", "messages": [{"role":"user","content":"..."}]}'
</code></pre>
<p dir="auto"><strong>參考：</strong></p>
<ul>
<li>本地 AI Zone 完整 fingerprinting 分析：<a href="https://local-ai-zone.github.io/blog/ox-alpha-stealth-model-comprehensive-analysis.html" rel="nofollow ugc">local-ai-zone.github.io/blog/ox-alpha-stealth-model-comprehensive-analysis.html</a></li>
<li>站內討論：<a href="https://lcz.me/topic/1281">lcz.me/topic/1281</a> (yangpiao)、<a href="https://lcz.me/topic/1275">lcz.me/topic/1275</a> (perter)</li>
</ul>
<hr />
<h2>5. NVIDIA AI 伺服器明年漲價 &gt;15%——HBM/DRAM 漲價的連鎖反應來了</h2>
<p dir="auto"><strong>一句話：</strong> NVIDIA 8/22 通知主要客戶，<strong>AI 伺服器出貨價明年起漲 &gt;15%</strong>，適用 Vera Rubin、Grace Blackwell 全線，原因是 HBM 與 DRAM 供應吃緊。</p>
<p dir="auto"><strong>為什麼值得看：</strong></p>
<p dir="auto">這個不是「雲端 API 漲價」那種跟我無關的新聞——<strong>HBM/DRAM 漲價是整條供應鏈</strong>：</p>
<pre><code>HBM/DRAM 漲
  ↓
Samsung / SK hynix 利潤 ↑
  ↓
NVIDIA 採購成本 ↑
  ↓
AI 伺服器出貨價 ↑
  ↓
雲端 AI 服務（Azure / AWS / GCP）成本 ↑
  ↓
API 報價 ↑
  ↓
買卡、跑本地模型、付 API 的成本，全部受影響
</code></pre>
<p dir="auto">對<strong>自己跑本地的人</strong>有兩個直接訊號：</p>
<ol>
<li><strong>消費級顯卡只會更貴</strong>——HBM3e 產能優先給 NVIDIA H100/H200/Blackwell 資料中心用，消費級 GDDR7 的產能就被擠壓，<strong>今年 Q4 到明年 Q1 的 5070 Ti / 5080 / 5090 通路價會再漲一輪</strong>。現在看中的卡，建議「能買就買」。</li>
<li><strong>二手顯卡市場會回溫</strong>——HBM 短缺會讓更多中小公司退而求其次買二手 RTX 3090 / 4090 跑本地 LLM，這會推升二手價 10~15%。</li>
</ol>
<p dir="auto"><strong>具體數字：</strong></p>
<table class="table table-bordered table-striped">
<thead>
<tr>
<th>項目</th>
<th>數字</th>
</tr>
</thead>
<tbody>
<tr>
<td>漲幅</td>
<td>&gt;15%（多數情況）</td>
</tr>
<tr>
<td>適用產品線</td>
<td>Vera Rubin、Grace Blackwell</td>
</tr>
<tr>
<td>生效時間</td>
<td>2027 年初出貨</td>
</tr>
<tr>
<td>觸發原因</td>
<td>HBM、DRAM 供應吃緊</td>
</tr>
<tr>
<td>影響客戶</td>
<td>微軟、Meta、Google、AWS、Oracle 等主要買家</td>
</tr>
</tbody>
</table>
<p dir="auto"><strong>我自己會做的事：</strong> 把原本排到「雙 11 看 5080 跌價」的計畫往前移到「這個月下單」。<strong>4090 二手</strong>——HBM 對 GDDR6X 的傳導比較間接，二手 4090 反而可能成為 2026 下半年最划算的本地 LLM 選擇。</p>
<p dir="auto"><strong>參考：</strong></p>
<ul>
<li>Bloomberg 原文（多家媒體轉載）：<a href="https://www.koreajoongangdaily.com/business/nvidia-to-raise-ai-server-prices-by-more-than-15-as-memory-supply-tightens-bloomberg/12838616" rel="nofollow ugc">koreajoongangdaily.com</a>、<a href="https://thenextweb.com/news/nvidia-ai-server-price-increase-memory-costs" rel="nofollow ugc">thenextweb.com</a></li>
<li>報導匯總：<a href="https://finance.yahoo.com/technology/ai/articles/nvidia-reportedly-warns-top-customers-102824176.html" rel="nofollow ugc">Yahoo Finance</a>、<a href="https://tbreak.com/nvidia-ai-server-price-increase-15-percent/" rel="nofollow ugc">tbreak.com</a></li>
</ul>
<hr />
<h2>結語：本週的訊號</h2>
<p dir="auto">如果這週的新聞要濃縮成一句話，我會說：</p>
<blockquote>
<p dir="auto"><strong>「開源 LLM 的部署效率到達拐點（SGLang 785×）、中國資本市場開始全面押注 AI（阿里 800 億）、硬體供應鏈進入新一輪吃緊（HBM/DRAM → NVIDIA 漲價 → 消費級卡漲價）。」</strong></p>
</blockquote>
<p dir="auto">對「自己跑模型」這個群體，<strong>這三件事合在一起意味著</strong>：</p>
<ol>
<li><strong>現在就升級本地工具鏈</strong>——SGLang Weight Cache Daemon、Qwen3.8-27B 各種量化，趕快熟悉</li>
<li><strong>現在就買卡</strong>——尤其 5070 Ti / 4090 二手，別等年底</li>
<li><strong>週末把 agent 排程跑滿</strong>——DeepSeek 週末離峰價 + OX Alpha 免費預覽，這三天 token 成本是歷史低點</li>
</ol>
<hr />
]]></description><link>https://lcz.me/topic/1284</link><generator>RSS for Node</generator><lastBuildDate>Mon, 21 Sep 2026 03:57:24 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1284.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 24 Aug 2026 01:26:56 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 📡 AI 新聞 8/24｜SGLang 重啟提速 785 倍、阿里 800 億港元全額砸 AI、NVIDIA AI 伺服器明年漲價 >15% on Mon, 24 Aug 2026 02:50:31 GMT]]></title><description><![CDATA[<p dir="auto">qwen 3.8刚好出来，加上我开发的minimax h3 studio 加上本地模型,加上原本的z image, 达到一条龙本地工具链。。。</p>
<p dir="auto">话说如此，要不是qwen3.8, 本地工具链得一环（剧本，分镜，动作设计，）还是得靠deepseek 网上模型</p>
]]></description><link>https://lcz.me/post/13667</link><guid isPermaLink="true">https://lcz.me/post/13667</guid><dc:creator><![CDATA[imbiplaza ASUS]]></dc:creator><pubDate>Mon, 24 Aug 2026 02:50:31 GMT</pubDate></item></channel></rss>