<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[5090 + vLLM + Qwen3.6-27B 成功分享]]></title><description><![CDATA[[[topic:post-is-deleted]]]]></description><link>https://lcz.me/topic/289/5090-vllm-qwen3.6-27b-成功分享</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 22:16:56 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/289.rss" rel="self" type="application/rss+xml"/><pubDate>Sun, 24 May 2026 09:08:39 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Tue, 26 May 2026 03:19:26 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 刚刚问了问AI，如果DFlash合并以后，大显存还是有优势。3080 40g又香了<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f924.png?v=1376b21da6c" class="not-responsive emoji emoji-android emoji--drooling_face" style="height:23px;width:auto;vertical-align:middle" title=":drooling_face:" alt="🤤" /></p>
]]></description><link>https://lcz.me/post/3716</link><guid isPermaLink="true">https://lcz.me/post/3716</guid><dc:creator><![CDATA[rock shi]]></dc:creator><pubDate>Tue, 26 May 2026 03:19:26 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Mon, 25 May 2026 08:58:57 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/im17me" aria-label="Profile: im17me">@<bdi>im17me</bdi></a> 大显存还是有好处，一个27b展开128k上下文基本上就32g了</p>
]]></description><link>https://lcz.me/post/3568</link><guid isPermaLink="true">https://lcz.me/post/3568</guid><dc:creator><![CDATA[rock shi]]></dc:creator><pubDate>Mon, 25 May 2026 08:58:57 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Mon, 25 May 2026 07:07:08 GMT]]></title><description><![CDATA[<p dir="auto">实测效果，模型Qwen3.6-27B-AWQ-INT4，5090还没有双卡3090跑的快。双卡3090稳定在60token/s</p>
]]></description><link>https://lcz.me/post/3550</link><guid isPermaLink="true">https://lcz.me/post/3550</guid><dc:creator><![CDATA[im17me]]></dc:creator><pubDate>Mon, 25 May 2026 07:07:08 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Mon, 25 May 2026 03:54:50 GMT]]></title><description><![CDATA[<p dir="auto">@Trypt-Wang llama.cpp比vllm简单10倍哈哈哈哈，DeepSeek基本上自己就能搞定。一两个小时差不多，还得算上下载模型的时间</p>
]]></description><link>https://lcz.me/post/3514</link><guid isPermaLink="true">https://lcz.me/post/3514</guid><dc:creator><![CDATA[rock shi]]></dc:creator><pubDate>Mon, 25 May 2026 03:54:50 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Sun, 24 May 2026 18:56:06 GMT]]></title><description><![CDATA[<p dir="auto">要搞到110，否则就浪费了这个卡</p>
]]></description><link>https://lcz.me/post/3448</link><guid isPermaLink="true">https://lcz.me/post/3448</guid><dc:creator><![CDATA[Hank Wang]]></dc:creator><pubDate>Sun, 24 May 2026 18:56:06 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Sun, 24 May 2026 13:46:17 GMT]]></title><description><![CDATA[<p dir="auto">@Trypt-Wang 不行试试llama.cpp，3080跑27b开了MTP差不多40-60tokens/s，hermes跑任务差不多50左右。可以先给hermes接入DeepSeek，让DeepSeek帮忙部署，我是花了两三块钱搞定的。</p>
]]></description><link>https://lcz.me/post/3410</link><guid isPermaLink="true">https://lcz.me/post/3410</guid><dc:creator><![CDATA[rock shi]]></dc:creator><pubDate>Sun, 24 May 2026 13:46:17 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Sun, 24 May 2026 13:24:14 GMT]]></title><description><![CDATA[<p dir="auto">那只有搜下vllm下其他人怎么mtp优化了，我最近都没跑本地llm，在用deepseek<img src="https://lcz.me/assets/plugins/nodebb-plugin-emoji/emoji/android/1f605.png?v=1376b21da6c" class="not-responsive emoji emoji-android emoji--sweat_smile" style="height:23px;width:auto;vertical-align:middle" title="😅" alt="😅" /></p>
]]></description><link>https://lcz.me/post/3406</link><guid isPermaLink="true">https://lcz.me/post/3406</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sun, 24 May 2026 13:24:14 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Sun, 24 May 2026 11:21:56 GMT]]></title><description><![CDATA[<p dir="auto">@Trypt-Wang 不要尝试dflash，还不够成熟，未来看它。你找apex mtp的模型看看，问下AI，论坛有AMD的帖子你去参考下，它们跑什么你就跑什么。</p>
]]></description><link>https://lcz.me/post/3397</link><guid isPermaLink="true">https://lcz.me/post/3397</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sun, 24 May 2026 11:21:56 GMT</pubDate></item><item><title><![CDATA[Reply to 5090 + vLLM + Qwen3.6-27B 成功分享 on Sun, 24 May 2026 11:17:39 GMT]]></title><description><![CDATA[<p dir="auto">非常不错，这下子图文并茂，挺好的，5090跑差不多就行了，不要着迷于优化，先把hermes安排点小任务，跑起来，优化有空了慢慢来，主要是MTP，Dflash，Turboquant，MTP目前比较成熟了。先跑。你不用折腾NVFP4，它和Q4量化没啥区别，它在Deepseek这样的原生模型中有优势，区别也不是很大。5090比较值得你折腾的就MTP。然后是turoquant。</p>
]]></description><link>https://lcz.me/post/3393</link><guid isPermaLink="true">https://lcz.me/post/3393</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sun, 24 May 2026 11:17:39 GMT</pubDate></item></channel></rss>