好消息是Nvidia Pro 4500 Blackwell,說不定還有其他卡,也能用
jlist
-
RTX 5090 Qwen3.8-27B dsh 全套實測:自寫推理引擎 NInfer × DeepSeek Harness × 開源編碼 agent 任務 —— 跑 13 小時真實編程任務的數據、6 個坑、以及「並發不是越大越快」 -
RTX 5090 Qwen3.8-27B dsh 全套實測:自寫推理引擎 NInfer × DeepSeek Harness × 開源編碼 agent 任務 —— 跑 13 小時真實編程任務的數據、6 個坑、以及「並發不是越大越快」NInfer有Windows版,可以直接用, 説不定比WSL快一點?
-
單張R9700 AI PRO 32G + VLLM + amd/Qwen3.8-27B-Quark-AWQ-MXFP4 -
主要玩comfyui跑minimax-h3,想把手上的7900XTX换成R9700可行吗? -
主要玩comfyui跑minimax-h3,想把手上的7900XTX换成R9700可行吗? -
求配置 4090 还是 RTX PRO 5000 谁有实际数据 -
求配置 4090 还是 RTX PRO 5000 谁有实际数据 -
Mac 低配篇M4系列16-32G。大众基本都是这个版本的用户。敬请参考!@williamlouis 同意,做控制可用,也有些鸡肋,35B还不是很可靠,27B好一些,很慢。
-
Mac 低配篇M4系列16-32G。大众基本都是这个版本的用户。敬请参考! -
对 M5 MAX 跑本地大模型有点失望 -
对 M5 MAX 跑本地大模型有点失望@Tony-Wang 有道理。我没有写过,可以想象写东西知识多会有帮助。做研究是否需要大模型?
-
对 M5 MAX 跑本地大模型有点失望感觉unified memory小主机的sweet spot在64GB,跑35B MOE+Hermes比较适合,虽然不算smart,基本可用。27B驱动Hermes就很慢不可用。其余24-32GB留作系统内存,接eGPU也够用。DGX Spark/Strix Halo 128GB RAM的,可以装进大model因为速度太慢也不实用,也就鸡肋了。
-
求助下大家,打算买个R9700,想问下几年前的电脑CPU和显卡能继续用上吗@Xiaote 謝謝。不過不知道哪個是對的, Gemini如此說:
By default, ComfyUI does not load full model weights into system RAM as a complete duplicate before copying them to VRAM. Instead, it uses PyTorch's memory mapping (mmap) or meta-devices to inspect and stage weights directly, dynamically streaming or moving only the necessary components into VRAM as execution demands.How ComfyUI Manages MemoryMemory Mapping (mmap): For formats like safetensors, ComfyUI maps files via pointers rather than performing deep copies into system RAM.Dynamic Loading: Individual model parts are paged or transferred to the GPU device context selectively during processing.Startup Flags: Behavior changes depending on arguments like --highvram (which keeps models resident on the GPU) or --gpu-only (which attempts to bypass system memory staging entirely).
-
分享:4090/48G, R9700/32G, AI Max 395 (8060S) 跑大语言模型的实测数据 -
分享:4090/48G, R9700/32G, AI Max 395 (8060S) 跑大语言模型的实测数据@Xiaote 不好意思,我前面打错了。改过model.context_length也无效。这个是hermes显示设定值:
$ hermes config get model.context_length
32000Hermes可以启动,不回答任何问题,直接显示前面截图里面的消息。
Model qwen3.5-4b-mtp@q6_k_xl has a context window of 32,000 tokens, which is below the minimum 64,000 required by Hermes Agent. Choose a model with at least 64K context. If your server reports a window smaller than the model's true window, set model.context_length in config.yaml to the real value (this must be at least 64K).
-
求助下大家,打算买个R9700,想问下几年前的电脑CPU和显卡能继续用上吗 -
分享:4090/48G, R9700/32G, AI Max 395 (8060S) 跑大语言模型的实测数据 -
小白請教入門級的全新硬體 -
分享:4090/48G, R9700/32G, AI Max 395 (8060S) 跑大语言模型的实测数据 -
分享:4090/48G, R9700/32G, AI Max 395 (8060S) 跑大语言模型的实测数据
