<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见]]></title><description><![CDATA[<p dir="auto">大家好，第一次在论坛发帖，有点紧张，有什么不对的地方请大家多包涵。</p>
<p dir="auto">我一直对LLM推理很感兴趣，自己动手写了一个极简的推理引擎。因为对数值计算有点执念，我走了一条比较偏门的路子——尽可能把浮点运算都换成整数和定点数（像Q8.8/Q16这种），想看看在纯CPU上能跑成什么样。</p>
<p dir="auto">折腾了一段时间，目前的情况是这样的：</p>
<p dir="auto">速度：在我的台式机CPU上，2B大小的模型能稳定跑到 15+ token/s，最低也在10以上</p>
<p dir="auto">确定性：因为全是整数运算，我做了Trace Hash校验，结果是一致的</p>
<p dir="auto">语义：简单对比过，和浮点版本的输出差异不大，没有明显崩坏</p>
<p dir="auto">我知道这个速度不算惊艳，毕竟只是个2B小模型，而且CPU跑大模型本来就不现实。但我自己觉得这条路好像可以继续往下走——接下来想试试27B的模型，还有昇腾那边的NPU，也在关注Tenstorrent，感觉他们的分层计算思路和我的方向有点契合。</p>
<p dir="auto">我发这个帖子，主要是想请大家帮我看看：</p>
<p dir="auto">这种“全部用整数算”的方向，你们觉得有实际落地的价值吗？还是说只是我个人的玩具？</p>
<p dir="auto">如果我想往NPU上移植，有没有什么是你们觉得最该提前注意的坑？</p>
<p dir="auto">论坛里有没有同样在搞推理引擎的朋友，想交个朋友，互相学习一下</p>
<p dir="auto">代码还很糙，暂时没好意思发出来，等再打磨一下我再开源吧，目前有个验证版的开源<a href="https://github.com/superalp1985/DCA-Inference-Engine%E3%80%82%E8%B0%A2%E8%B0%A2%E5%A4%A7%E5%AE%B6%E7%9C%8B%E5%AE%8C%EF%BC%8C%E5%B8%8C%E6%9C%9B%E6%B2%A1%E5%8D%A0%E7%94%A8%E4%BD%A0%E4%BB%AC%E5%A4%AA%E5%A4%9A%E6%97%B6%E9%97%B4%E3%80%82" rel="nofollow ugc">https://github.com/superalp1985/DCA-Inference-Engine。谢谢大家看完，希望没占用你们太多时间。</a></p>
]]></description><link>https://lcz.me/topic/800/刚把推理引擎跑到15-token-s-新人发帖想听听大家的意见</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 20:02:14 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/800.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 07 Jul 2026 16:06:22 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Sun, 19 Jul 2026 11:56:33 GMT]]></title><description><![CDATA[<p dir="auto">那个卡应该是8卡叠加才效果好，也是再国产的无奈。业余玩家玩的话，单卡感觉有点像dgx spark或者amd 395。看被显存大，实则跑得慢，除非你能忍受晚上让它自己跑， 用时长换质量</p>
]]></description><link>https://lcz.me/post/10115</link><guid isPermaLink="true">https://lcz.me/post/10115</guid><dc:creator><![CDATA[stxpnet]]></dc:creator><pubDate>Sun, 19 Jul 2026 11:56:33 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Thu, 09 Jul 2026 02:06:26 GMT]]></title><description><![CDATA[<p dir="auto">不求能追上CUDA 只要大概能跑 没有过于忍不了的延迟就行 特别是国内单位 现在都必须全国产 所以这问题早晚得解决</p>
]]></description><link>https://lcz.me/post/9515</link><guid isPermaLink="true">https://lcz.me/post/9515</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Thu, 09 Jul 2026 02:06:26 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Thu, 09 Jul 2026 02:04:06 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> CANN社区他们内部开发人员都吐槽 别提了 但是没钱买显卡 这是硬伤</p>
]]></description><link>https://lcz.me/post/9514</link><guid isPermaLink="true">https://lcz.me/post/9514</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Thu, 09 Jul 2026 02:04:06 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 14:25:22 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a></p>
<p dir="auto">CANN架構要追上CUDA估計還要起碼五年以上的開發跟成熟吧...</p>
<p dir="auto">不過CUDA可不會停下來等CANN</p>
]]></description><link>https://lcz.me/post/9494</link><guid isPermaLink="true">https://lcz.me/post/9494</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Wed, 08 Jul 2026 14:25:22 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 14:23:19 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/bingqin-wang" aria-label="Profile: bingqin-wang">@<bdi>bingqin-wang</bdi></a></p>
<p dir="auto">因爲我沒有做過内核維護跟開發, 所以沒辦法作太多評論</p>
<p dir="auto">不過我以前有見過有人討論過純整數運算的LLM引擎期刊, 名字好像叫I-LLM</p>
<p dir="auto">Integer-Only Inference, 如果沒記錯的話用的是INT4以及Bit Packing的INT6</p>
]]></description><link>https://lcz.me/post/9493</link><guid isPermaLink="true">https://lcz.me/post/9493</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Wed, 08 Jul 2026 14:23:19 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 11:41:53 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> 这卡目前来看 除了显存大 貌似就都是骂 没看到社区有夸优点的</p>
]]></description><link>https://lcz.me/post/9469</link><guid isPermaLink="true">https://lcz.me/post/9469</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Wed, 08 Jul 2026 11:41:53 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 11:35:55 GMT]]></title><description><![CDATA[<p dir="auto">我今天才在研究這張卡 Atlas 300I Duo 推理卡 不知道群裡是否有人使用過？</p>
<p dir="auto">LPDDR4X 96GB或48GB，总带宽408GB/s</p>
<p dir="auto"><img src="https://upload.lcz.me/uploads/0cb17d35-71f3-4ce0-9d3f-512730641327.jpeg" alt="7c339a5c-81d6-41e8-8924-446cc7540b06-image.jpeg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/de60f1b0-5f6f-4137-aa28-5e4e54a86d1c.jpeg" alt="85456b9f-6bb6-4736-ab32-ba04704a7e71-image.jpeg" class=" img-fluid img-markdown" /></p>
]]></description><link>https://lcz.me/post/9467</link><guid isPermaLink="true">https://lcz.me/post/9467</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Wed, 08 Jul 2026 11:35:55 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 11:27:25 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kos-or" aria-label="Profile: kos-or">@<bdi>kos-or</bdi></a> 其实说白了就是穷的 我就想现在不是升腾的加速卡便宜 看看能不能想办法把不到一万块的96g显存用起来 不求n卡的速度 起码能用就是胜利 一分钱一分货 能用就行</p>
]]></description><link>https://lcz.me/post/9466</link><guid isPermaLink="true">https://lcz.me/post/9466</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Wed, 08 Jul 2026 11:27:25 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 11:08:28 GMT]]></title><description><![CDATA[<p dir="auto">不懂, 但是只要能你做得比其他引擎快 就成了; 建議和其他引擎benchmark 在相同條件下 對照一下推理速度和輸出品質 等你的好消息 <img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f642.png?v=8624a6a055f" class="not-responsive emoji emoji-android emoji--slightly_smiling_face" style="height:23px;width:auto;vertical-align:middle" title=":)" alt="🙂" /></p>
]]></description><link>https://lcz.me/post/9464</link><guid isPermaLink="true">https://lcz.me/post/9464</guid><dc:creator><![CDATA[kos or]]></dc:creator><pubDate>Wed, 08 Jul 2026 11:08:28 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 07:59:17 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a> 果然是专家 确实是非浮点运算 但是比您说的这种更极端 我从数学底层把所有乘除和微分计算都换成了有限加法表达的公式 也就是全面离散数学化 非线性计算用了一部分查表 我的目标是NPU 利用CANN算子做优化 这样有俩个好处1.可以最大避免带宽影响大力出奇迹 2.所有结果计算可审计 相同的seed得到的答案完全一致 比较适合高确定性场景</p>
]]></description><link>https://lcz.me/post/9441</link><guid isPermaLink="true">https://lcz.me/post/9441</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Wed, 08 Jul 2026 07:59:17 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Wed, 08 Jul 2026 01:17:36 GMT]]></title><description><![CDATA[<p dir="auto">整數運算? 是指非FP形式的GEMM嗎? INT8的W8A8或者INT4的W4A4?</p>
<p dir="auto">如果是用來適配舊顯卡/NPU這個思路不是不行</p>
<p dir="auto">但是目前的新顯卡大方向是FP低精度運算, 可能聚焦在NPU比較好? 目前大多數NPU的用法也是NPU Prefill + GPU Decode</p>
<p dir="auto">AMD XDNA系列跟Intel的NPU系列</p>
<p dir="auto">Intel的爛命名, 沒有架構名字, 只叫NPU然後加個數字, Arrow Lake (Ultra 200)用NPU 3720, Panther Lake (Ultra 300)用NPU 5</p>
]]></description><link>https://lcz.me/post/9416</link><guid isPermaLink="true">https://lcz.me/post/9416</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Wed, 08 Jul 2026 01:17:36 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Tue, 07 Jul 2026 18:14:18 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 感谢老大 你的视频对我帮助也很大 少踩了不少坑 你太谦虚了</p>
]]></description><link>https://lcz.me/post/9406</link><guid isPermaLink="true">https://lcz.me/post/9406</guid><dc:creator><![CDATA[bingqin wang]]></dc:creator><pubDate>Tue, 07 Jul 2026 18:14:18 GMT</pubDate></item><item><title><![CDATA[Reply to 刚把推理引擎跑到15 token/s，新人发帖想听听大家的意见 on Tue, 07 Jul 2026 17:02:41 GMT]]></title><description><![CDATA[<p dir="auto">这里又不是麻省理工，发个帖子有什么紧张的。你发的这个玩意一般人看不懂，技术大牛们可以来给你品品，超出了我的理解范围。</p>
]]></description><link>https://lcz.me/post/9403</link><guid isPermaLink="true">https://lcz.me/post/9403</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 07 Jul 2026 17:02:41 GMT</pubDate></item></channel></rss>