跳转至内容
  • 版块
  • 最新
  • 标签
  • 热门
  • 用户
  • 群组
皮肤
  • 浅色
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • 深色
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • 默认(LCZ-Blue)
  • 不使用皮肤
  • LCZ-Green
  • LCZ-Blue
  • LCZ-Black
折叠
品牌标识

抡锤者

首页 版块 标签 硬件 AI 广场
胜威马胜

ahsayhdoabnklah

@胜威马
取消关注 关注
关于
帖子
4
主题
1
分享
0
群组
0
粉丝
0
关注
0

帖子

最新 最佳 有争议的

  • 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G,本地部署 Qwen3.8-27B 怎么吐字最快?
    胜威马胜 胜威马

    @asd2667 试了一下,现在是不能开mtp的……速度太慢了……Found it. The truth:

    • Warm TTFT is actually fine: your unique 37-token prompt just processed at 29.6 tok/s (1.25s).
    • The 2.5 tok/s you hit was the first request after the restart — one-time CUDA-graph + MTP draft compilation. That's a cold-start cost, not steady-state.
      But I also found MTP is hurting, not helping on this 16GB setup:
    • Decode: 12.8 tok/s with MTP vs 18.9 tok/s without it earlier (the draft compute overhead exceeds the speculative gain at these speeds; draft acceptance mean was only ~2.3).
    • MTP adds the cold-start compile overhead you felt.
      So per the official page, MTP was meant for the bigger M-model on larger VRAM. On your XS-Pro/16GB it's net-negative. I'll disable MTP, keep the rest of the recommended config (q8/q4 KV, flash-attn, 64k, checkpoints, sampling), and pre-warm so you never hit the cold start:
    AI硬件 rtx4060 qwen-27b intel

  • 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G,本地部署 Qwen3.8-27B 怎么吐字最快?
    胜威马胜 胜威马

    @williamlouis 说:

    只换显卡的话 还是 7900XTX 24G 。速度提升最快。
    考虑后续开发建议 R9700 32G。

    有点纠结,7900xt好像涨价了。性价比感觉都不如9700了。想着要不要等一等,而且换了这个还要换电源,现在的电源是750w的

    AI硬件 rtx4060 qwen-27b intel

  • 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G,本地部署 Qwen3.8-27B 怎么吐字最快?
    胜威马胜 胜威马

    @asd2667 说:

    zerodigest/Qwen3.8-27B-Uncensored-YMQ-MTP-GGUF

    谢谢,我今天晚上试试

    AI硬件 rtx4060 qwen-27b intel

  • 微星 B660M 迫击炮 + 4060Ti 16G + i7-12700K + DDR4 32G,本地部署 Qwen3.8-27B 怎么吐字最快?
    胜威马胜 胜威马

    配置如下:

    • 主板:微星 B660M 迫击炮(第一根 x16 槽是 CPU 的 PCIe 4.0 x16,第二根 x16 长槽是芯片组的 PCIe 3.0 x4)
    • 显卡:RTX 4060Ti 16G 单卡
    • CPU:i7-12700K
    • 内存:DDR4 32G

    想本地跑 Qwen3.8-27B(稠密 27B),请教几个问题:

    1. 单卡 16G 显然装不下 27B 权重(Q4_K_XL 就 16G+)+ KV 缓存,必然 offload 到内存。这种单机单卡 + 32G DDR4 的 offload 方案,实测吐字大概能到多少 t/s?是不是基本就是“慢成乌龟”?

    2. 想提速最划算的路子是不是再淘一张 16G 卡(第二张 4060Ti 16G 或 5060Ti 16G)组双卡 32G,走 llama.cpp 张量并行(-sm tensor -ts 1,1)+ MTP?B660M 副槽只有 PCIe 3.0 x4,跑 27B 推理带宽够不够用?

    3. 量化与参数:27B 上 Q4_K_M / UD-Q4_K_XL 就行了吧?MTP 的 spec-draft-n-max 在双 16G 卡上设 1 还是 2 比较稳(怕 OOM)?上下文开到 64K 是不是就只能放弃 KV 量化?

    论坛里看过双 3060(24G) 跑 43~50 t/s、双 5060Ti 65 t/s 的帖子,想确认下我这板子加第二张卡是否是最优解,还是单卡硬扛有更聪明的办法。谢谢各位大佬!

    AI硬件 rtx4060 qwen-27b intel
  • 登录

  • 登录或注册以进行搜索。
  • 第一个帖子
    最后一个帖子
0
  • 版块
  • 最新
  • 标签
  • 热门
  • 用户
  • 群组