<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090]]></title><description><![CDATA[<p dir="auto">想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090的问题</p>
<p dir="auto">我知道这个问题有点奇怪，就是关于资源调用这部份，有没有可能设定优先顺序或是更优化的安排让 显卡 榨干价值？</p>
<p dir="auto">这问题我问过 hermes (deepseek) 主要是建议我 两个服务都使用 各半的 VRAM 虽然说听起来很美好，但我总觉得怪怪的，牺牲上下文长度不说，感觉生成出来的效果也会不好。</p>
<p dir="auto">想请教看看各位大神有没有其他解法可以参考 谢谢</p>
]]></description><link>https://lcz.me/topic/872/想请问各位大神关于vram共用的问题-comfyui-llamacpp-想共用一张3090</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 20:02:27 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/872.rss" rel="self" type="application/rss+xml"/><pubDate>Sun, 19 Jul 2026 07:59:52 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Thu, 23 Jul 2026 10:18:10 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/botio-kuo" aria-label="Profile: botio-kuo">@<bdi>botio-kuo</bdi></a>  你直接告诉 hermes 。我需要配置一个辅助的deepseek-falsh模型。apikey是****。 让他自己改 config.yaml 。<br />
当然，你先的备份好<br />
当然，你干脆把config.yaml 贴给网页 的deepseek 。让他把deppsek改为辅助模型。 它可能考虑比你 的本地模型更加细致。<br />
总之都能改好，<br />
万一没弄好。备份 的覆盖一下，重启，再来。直到搞好。<br />
我的大概也是这么搞好的。<br />
全程你只管测试</p>
]]></description><link>https://lcz.me/post/10372</link><guid isPermaLink="true">https://lcz.me/post/10372</guid><dc:creator><![CDATA[wwcd2016]]></dc:creator><pubDate>Thu, 23 Jul 2026 10:18:10 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 16:45:18 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/q-maria" aria-label="Profile: Q-maria">@<bdi>Q-maria</bdi></a> 那个 comfyui 在LINUX跑 可以用其他主机 (mac也可以 我目前就这样）安装 comfyui desktop 当作GUI去使用，就可以观察实际状况，但配置还是可以丢给 hermes 用 SSH 处理</p>
<p dir="auto">上面 <a class="plugin-mentions-user plugin-mentions-a" href="/user/wwcd2016" aria-label="Profile: wwcd2016">@<bdi>wwcd2016</bdi></a> 大神说的应用 我刚好都有使用 obsidian cli (当作 LLM WIKI），确实 comfyui 使用的频率比较低，我是比较好奇自动切换的部分，可能要让 hermes 写个 SKILL 试试 但确实是个好思路 感谢</p>
]]></description><link>https://lcz.me/post/10126</link><guid isPermaLink="true">https://lcz.me/post/10126</guid><dc:creator><![CDATA[Botio Kuo]]></dc:creator><pubDate>Sun, 19 Jul 2026 16:45:18 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 16:38:55 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 这个想法好像有戏，我也在考虑上面大神说的方式 多一张AI PRO R9700卡，主要目前硬体配置来说电源W数有点紧，不然搞个PVE去直通两张卡不同工作这件事情是蛮简单的，因为我本来就有闲置的 AMD 1950x CPU 之前是跑k8s 多NODE，但都是吃电怪兽。。。</p>
<p dir="auto">另外一个现实的问题是，目前没有用这些设备创立营收的话，AI PRO R9700 也是不小的开销</p>
<p dir="auto">不过谢谢各位大神的意见</p>
]]></description><link>https://lcz.me/post/10125</link><guid isPermaLink="true">https://lcz.me/post/10125</guid><dc:creator><![CDATA[Botio Kuo]]></dc:creator><pubDate>Sun, 19 Jul 2026 16:38:55 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 15:03:13 GMT]]></title><description><![CDATA[<p dir="auto">27b模型载入就没啥空间了，不可能同时跑。FLux LTX这些模型经常十几个G，不可能同时运行。让Hermes自己切换就好了。它知道怎么管理，自己理解。但是你的Hermes最好是在线模型，如果用本地的，模型卸载了，你的hermes就嗝屁了。或者你让Hermes写一个脚本，自动切换。</p>
]]></description><link>https://lcz.me/post/10121</link><guid isPermaLink="true">https://lcz.me/post/10121</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Sun, 19 Jul 2026 15:03:13 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 14:58:28 GMT]]></title><description><![CDATA[<p dir="auto">我就是7900xtx 单卡 一边跑LLM和VOXCPM时没障碍。但只要你要跑comfyui时Hermes就会提示要终止llm或者voxcpm。我现在也考虑在弄一张3090主要comfyui在Linux上跑起来还是不太协调</p>
]]></description><link>https://lcz.me/post/10120</link><guid isPermaLink="true">https://lcz.me/post/10120</guid><dc:creator><![CDATA[Q maria]]></dc:creator><pubDate>Sun, 19 Jul 2026 14:58:28 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 13:18:46 GMT]]></title><description><![CDATA[<p dir="auto">我跟你一个情况。3090 单卡，默认启动就是llama 。驱动了 hermes。 obsidian。 wiki。 把事情干完了<br />
晚上没事想玩玩comfyui ，就对hermes 说，打开comfyui就好了。会帮你配置好的。</p>
<p dir="auto">关键配置是：你把deepseek作为辅助模型就好了。主模型随便关，随便切换。<br />
我自己有至少5个模型常用的，也是让hermes 随意切换。让它做好脚本保存好。</p>
<p dir="auto">总之：你的业务如果不冲突：就是说大批量调用llama 的时候也必须用 comfyui。<br />
一张卡挺好用的。</p>
]]></description><link>https://lcz.me/post/10117</link><guid isPermaLink="true">https://lcz.me/post/10117</guid><dc:creator><![CDATA[wwcd2016]]></dc:creator><pubDate>Sun, 19 Jul 2026 13:18:46 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 10:28:24 GMT]]></title><description><![CDATA[<p dir="auto">24G VRAM 同时跑 ComfyUI（生图）和 llama.cpp（Qwen 36B）确实不太现实。Qwen 3.6-27B 即使 Q4_K_M 量化也要约 18G VRAM，ComfyUI 跑 Flux/SDXL 也要 8-12G，加起来远超 24G。</p>
<p dir="auto">几个实际方案（按推荐顺序）：</p>
<ol>
<li>
<p dir="auto"><strong>分时复用（最实际）</strong>：不同时运行，哪个在用就加载哪个。ComfyUI 关掉后再开 llama.cpp，或者反过来。现代 GPU 加载模型很快（10-20 秒），切换成本可以接受</p>
</li>
<li>
<p dir="auto"><strong>llama.cpp CPU offload</strong>：如果一定要同时运行，让 llama.cpp 只 offload 少量层到 GPU（--ngl 16，约占 4-5G），大部分推理走 CPU。缺点是推理速度会慢 5-10 倍，但至少可以同时用</p>
</li>
<li>
<p dir="auto"><strong>ComfyUI --lowvram 模式</strong>：能大幅降低 VRAM 占用到 4-6G，但生图速度会慢 2-3 倍。配合方案 2 可以同时跑，但两边都慢</p>
</li>
<li>
<p dir="auto"><strong>vLLM / Ollama 限制 GPU 内存</strong>：用 <code>--gpu-memory-utilization 0.5</code> 或 <code>OLLAMA_NUM_PARALLEL=1</code> 限制 LLM 只用一半 VRAM，另一半留给 ComfyUI。代价是上下文长度受限</p>
</li>
<li>
<p dir="auto"><strong>最佳长期方案</strong>：再加一张卡。预算有限的话，收一张二手 3060 12G（约 1000 元）专门跑 llama.cpp / vLLM，3090 专注 ComfyUI。两张卡分工明确，互不干扰，是最省心的方案</p>
</li>
</ol>
<p dir="auto">总结：一张 3090 同时跑两个重型任务不现实。分时复用最实际，两张卡最省心。可以先从方案 1 或 2 开始，等体验到位后再考虑加卡。</p>
]]></description><link>https://lcz.me/post/10106</link><guid isPermaLink="true">https://lcz.me/post/10106</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Sun, 19 Jul 2026 10:28:24 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 09:17:36 GMT]]></title><description><![CDATA[<p dir="auto">同时用一般比较困难，24G跑单一模式也就刚好。多任务并行建议加卡。比如说一张7900xtx专门用来跑LLM，3090留来跑comfyui<br />
<a href="https://lcz.me/topic/637/qwen3.6-27b-%E5%85%AD%E5%A4%A7%E5%90%AF%E5%8A%A8%E6%A8%A1%E5%BC%8F%E8%AF%A6%E8%A7%A3-%E6%80%A7%E8%83%BD-%E5%8F%82%E6%95%B0%E4%B8%8E%E5%9C%BA%E6%99%AF">https://lcz.me/topic/637/qwen3.6-27b-六大启动模式详解-性能-参数与场景</a><br />
我这个贴有涉及多卡分工，你可以参考下</p>
]]></description><link>https://lcz.me/post/10100</link><guid isPermaLink="true">https://lcz.me/post/10100</guid><dc:creator><![CDATA[abaalei]]></dc:creator><pubDate>Sun, 19 Jul 2026 09:17:36 GMT</pubDate></item><item><title><![CDATA[Reply to 想请问各位大神关于VRAM共用的问题，comfyUI + llamacpp 想共用一张3090 on Sun, 19 Jul 2026 08:02:05 GMT]]></title><description><![CDATA[<p dir="auto">补充说明一下： llamacpp 主要是跑 QWEN 36，本身是内网给其他电脑的hermes做调用，如果说要给多台电脑同时调用肯定不太实际</p>
]]></description><link>https://lcz.me/post/10099</link><guid isPermaLink="true">https://lcz.me/post/10099</guid><dc:creator><![CDATA[Botio Kuo]]></dc:creator><pubDate>Sun, 19 Jul 2026 08:02:05 GMT</pubDate></item></channel></rss>