<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Aimax 395的定制qwen 3.8 flash，prefill突破1000了，decode 40-50]]></title><description><![CDATA[<p dir="auto">地址：<a href="https://github.com/peonist-ai/halogen-flash-server" rel="nofollow ugc">https://github.com/peonist-ai/halogen-flash-server</a><br />
按照作者说法，是硬件匹配模型的1对1特殊优化，所以能这么牛逼<br />
拉取docker安装试了，确实能达到他说的速度，并且质量我没看出有啥大的衰减<br />
本来觉得aimax395就一台鸡肋，这个模型让它成为现在性价比最高的选择了</p>
]]></description><link>https://lcz.me/topic/1609</link><generator>RSS for Node</generator><lastBuildDate>Thu, 10 Sep 2026 17:27:13 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1609.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 10 Sep 2026 13:12:24 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Aimax 395的定制qwen 3.8 flash，prefill突破1000了，decode 40-50 on Thu, 10 Sep 2026 16:36:16 GMT]]></title><description><![CDATA[<p dir="auto">qwen3.8 flash很慢，不是prefill的问题，是它内部思维链的问题</p>
]]></description><link>https://lcz.me/post/17177</link><guid isPermaLink="true">https://lcz.me/post/17177</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Thu, 10 Sep 2026 16:36:16 GMT</pubDate></item></channel></rss>