跳转至内容
  • 版块
  • 最新
  • 标签
  • 热门
  • 用户
  • 群组
皮肤
  • 浅色
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • 深色
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • 默认(LCZ-Blue)
  • 不使用皮肤
  • LCZ-Green
  • LCZ-Blue
折叠
品牌标识

抡锤者

  1. 主页
  2. 版块
  3. LLM讨论区
  4. Qwen3.8 27B Q5_K_M + 7900xtx + DeepSeek Harness实战平均 52 t/s

Qwen3.8 27B Q5_K_M + 7900xtx + DeepSeek Harness实战平均 52 t/s

已定时 已固定 已锁定 已移动 LLM讨论区
qwen-27b7900xtxdsharness
36 帖子 14 发布者 1.1k 浏览 4 关注中
  • 从旧到新
  • 从新到旧
  • 最多赞同
回复
  • 在新帖中回复
登录后回复
此主题已被删除。只有拥有主题管理权限的用户可以查看。
  • X xl

    @exllm 并行改为1后,预填充性能有接近5~10倍的提升(不同的任务变化很大),但生成token的性能提升不到10%。我也将“--override-tensor blk.\d+.ffn_.*_exps.=CPU \”注释掉进行过测试,性能反而有略微下降,可能跟我的硬件和跑的任务也有一定关系。

    terryT 离线
    terryT 离线
    terry
    超级版主
    编写于 最后由 编辑
    #19

    @xl 不是巧合,就是应该注释掉

    油管:https://www.youtube.com/@抡锤者

    X 1 条回复 最后回复
    0
    • terryT terry

      @xl 不是巧合,就是应该注释掉

      X 离线
      X 离线
      xl
      编写于 最后由 编辑
      #20

      @terry 不是,在我的电脑上,注释掉后性能反而有略微下降(预填充变化不大,生成性能有略微降低)。只测一次的结果可能有偏差,等有空的时候再多测几次。

      terryT 1 条回复 最后回复
      0
      • X xl

        @terry 不是,在我的电脑上,注释掉后性能反而有略微下降(预填充变化不大,生成性能有略微降低)。只测一次的结果可能有偏差,等有空的时候再多测几次。

        terryT 离线
        terryT 离线
        terry
        超级版主
        编写于 最后由 terry 编辑
        #21

        @xl 总之参数大体是对的,我是复制链接给AI自己配置的。它把这一行直接干掉了,效率上和楼主的差不多。

        油管:https://www.youtube.com/@抡锤者

        1 条回复 最后回复
        0
        • Wang WeiW Wang Wei

          抄了作业 只有20-30 t/s 不过我是在win下的 wsl2挂着跑 显存有溢出 换原生Ubuntu太折腾了

          Wang WeiW 离线
          Wang WeiW 离线
          Wang Wei
          编写于 最后由 编辑
          #22

          Wang-Wei 说:

          抄了作业 只有20-30 t/s 不过我是在win下的 wsl2挂着跑 显存有溢出 换原生Ubuntu太折腾了

          我把np从2改到1就可以了,显存不溢出,q5 128k上下文 55t/s 换成q4可以到70 一个slot够我用了

          1 条回复 最后回复
          0
          • E exllm

            软硬件配置:

            cpu: AMD Ryzen 5 5600
            RAM: 32GB 3200
            GPU: 7900xtx
            OS: Ubuntu 26.04
            软件 : llama.cpp + vulkan
            版本:

               version: 9737 (67e9fd3b7)
               built with GNU 15.2.0 for Linux x86_64
            

            测试:

            1. llama.cpp 自带webui 创建打飞机游戏
              提示词:
            "Create a complete, fully functional 2D space shooter arcade game inside a single HTML file using HTML5 Canvas, CSS, and vanilla JavaScript.        
            **Requirements:**        
                    
            1.  **Canvas & Layout:** Set up a centered 800x600 black canvas with a retro arcade UI showing the current score and remaining lives at the top.        
                        
            2.  **Player Ship:** Draw a distinct player ship at the bottom center. Allow smooth movement left and right using the Arrow Keys or A/D keys, constrained within the canvas bounds.        
                        
            3.  **Shooting Mechanics:** Pressing the Spacebar should fire a laser projectile upward from the player's position. Implement a brief cooldown between shots.        
                        
            4.  **Enemies:** Spawn alien ships or asteroids at random X coordinates along the top, moving downward at varying speeds.        
                        
            5.  **Collision Detection:** Write precise collision logic for bullets hitting enemies (destroying both and adding points) and enemies hitting the player or bottom screen (losing a life).        
                        
            6.  **Game Loop & States:** Include smooth requestAnimationFrame logic, a 'Game Over' screen, and a 'Restart' button functionality. Clean, well-commented code only."
            

            日志

            1.47.020.662 I slot print_timing: id  1 | task 0 | n_decoded =   5409, tg =  89.47 t/s, tg_3s =  71.71 t/s
            1.47.476.673 I slot print_timing: id  1 | task 0 | prompt eval time =     519.21 ms /   262 tokens (    1.98 ms per token,   504.61 tokens per second)
            1.47.476.677 I slot print_timing: id  1 | task 0 |        eval time =   60909.77 ms /  5440 tokens (   11.20 ms per token,    89.31 tokens per second)
            1.47.476.677 I slot print_timing: id  1 | task 0 |       total time =   61428.98 ms /  5702 tokens
            1.47.476.681 I slot print_timing: id  1 | task 0 |    graphs reused =       1480
            1.47.476.684 I slot print_timing: id  1 | task 0 | draft acceptance = 0.87497 ( 3940 accepted /  4503 generated), mean acceptance length =  3.62, acceptance rate per position = (0.943, 0.875, 0.806)
            
            
            1. 配合 Deepseek Harness 修改QT + C++ 项目 SimulIDE, 使其能在apple silicon上运行,并修改Qt 文本框输入bug

            llama.cpp 日志中的一段

            32.32.489.805 I slot print_timing: id  1 | task 10430 | prompt eval time =   17284.15 ms /  4098 tokens (    4.22 ms per token,   237.10 tokens per second)
            32.32.489.807 I slot print_timing: id  1 | task 10430 |        eval time =    5461.48 ms /   277 tokens (   19.72 ms per token,    50.72 tokens per second)
            32.32.489.808 I slot print_timing: id  1 | task 10430 |       total time =   22745.63 ms /  4375 tokens
            32.32.489.808 I slot print_timing: id  1 | task 10430 |    graphs reused =       9929
            32.32.489.811 I slot print_timing: id  1 | task 10430 | draft acceptance = 0.89333 (  201 accepted /   225 generated), mean acceptance length =  3.68, acceptance rate per position = (0.960, 0.893, 0.827)
            
            

            DSH 摘要:

             63歩ILLM 13m13S・工具週用 2m2s|首token 平均3.8s・52 tok/s|緩存命中98%1輸入 4.4M tok・輸出 29.2K tok
            

            llama.cpp运行参数:

            -t 10 
            -b 512 
            -ub 256 
            --spec-draft-n-max 3 
            --fit off 
            --no-context-shift 
            --metrics 
            --kv-unified 
            --jinja 
            --cache-type-k q8_0 
            --cache-type-v q8_0 
            -fa on 
            --spec-type draft-mtp 
            --ctx-size 131072 
            --parallel 2 
            -ngl -1 
            --host 0.0.0.0 
            --port 8080 
            --chat-template-kwargs {"enable_thinking": true, "preserve_think": false, "reasoning_effort": "medium"} 
            --no-mmap 
            --temp 0.6 
            --top-p 0.5 
            --top-k 15 
            --repeat-penalty 1.0 
            --override-tensor blk\.\d+\.ffn_.*_exps\.=CPU 
            --alias qwen3.8-27b-q5 
            
            懒人烘培懒 离线
            懒人烘培懒 离线
            懒人烘培
            编写于 最后由 编辑
            #23

            @exllm @terry
            根据楼主的测试了一下
            cpu: AMD Ryzen 5 7500F
            RAM:DDR5 32GB 6000
            GPU: 7900XTX
            OS: Ubuntu 24.04
            软件 : llama.cpp + vulkan

            先说结果:
            d9100061-f4ce-41c6-a642-6c786631721d-image.jpeg
            llama.cpp运行参数,spec-draft-n-max 3 速度最好:

            ~/llama_vulkan/bin/llama-server \
            -m ~/models/Qwen3.8-27B-Q5_K_M.gguf \
            -t 10 \
            -b 512 \
            -ub 256 \
            --spec-draft-n-max 3 \
            --fit off \
            --no-context-shift \
            --metrics \
            --kv-unified \
            --jinja \
            --cache-type-k q8_0 \
            --cache-type-v q8_0 \
            -fa on \
            --spec-type draft-mtp \
            --ctx-size 131072 \
            --parallel 1 \
            -ngl -1 \
            --host 0.0.0.0 \
            --port 8080 \
            --chat-template-kwargs '{"enable_thinking":true,"preserve_think":false,"reasoning_effort":"medium"}' \
            --no-mmap \
            --temp 0.6 \
            --top-p 0.5 \
            --top-k 15 \
            --repeat-penalty 1.0 \
            --alias qwen3.8-27b-q5
            

            llama.cpp运行参数,spec-draft-n-max 2 :
            a0024f56-4593-4dff-b3e1-f28a779fa278-image.jpeg

            llama.cpp运行参数,spec-draft-n-max 4 :
            990f1775-f067-4cb1-a475-497ac29df862-image.jpeg

            总结:

            参数 Prompt速度 生成速度 平均延迟 MTP命中率
            n=2 288.73 t/s 68.14 t/s 17.436 s 87.62%
            n=3 286.60 t/s 71.74 t/s 16.772 s 82.50%
            n=4 249.53 t/s 47.96 t/s 24.911 s 74.99%

            这台电脑n=3最好。
            注:修改这个参数是MTP speculative decoding(推测解码)一次最多提前“猜”多少个 token。

            1 条回复 最后回复
            1
            • 6 离线
              6 离线
              6cccccc
              编写于 最后由 编辑
              #24

              为啥我的启动参数是这样:
              -m "$SELECTED_MODEL" -t 10 -b 512 -ub 256
              --spec-draft-n-max 3 --fit off --no-context-shift
              --metrics --kv-unified --jinja
              --cache-type-k q8_0 --cache-type-v q8_0
              -fa on --spec-type draft-mtp
              --ctx-size 131072 --parallel 1 -ngl -1
              --host 0.0.0.0 --port $LLAMA_PORT
              --chat-template-kwargs {"enable_thinking": true, "preserve_think": false, "reasoning_effort": "medium"}
              --no-mmap --temp 0.6 --top-p 0.5 --top-k 15 --repeat-penalty 1.0
              --alias qwen3.8-27b-q5
              仍然速度很低,只有12左右
              配置如下
              cpu: AMD Ryzen 7 2700 Eight-Core Processor (8核16线程, 3.2GHz)
              RAM: 64GB DDR4
              GPU: AMD Radeon RX 7900 XTX 24GB (RADV NAVI31)
              OS: Ubuntu 26.04 LTS (内核 7.0.0-29-generic)
              软件 : llama.cpp Vulkan b10486

              懒人烘培懒 E 3 条回复 最后回复
              0
              • 6 6cccccc

                为啥我的启动参数是这样:
                -m "$SELECTED_MODEL" -t 10 -b 512 -ub 256
                --spec-draft-n-max 3 --fit off --no-context-shift
                --metrics --kv-unified --jinja
                --cache-type-k q8_0 --cache-type-v q8_0
                -fa on --spec-type draft-mtp
                --ctx-size 131072 --parallel 1 -ngl -1
                --host 0.0.0.0 --port $LLAMA_PORT
                --chat-template-kwargs {"enable_thinking": true, "preserve_think": false, "reasoning_effort": "medium"}
                --no-mmap --temp 0.6 --top-p 0.5 --top-k 15 --repeat-penalty 1.0
                --alias qwen3.8-27b-q5
                仍然速度很低,只有12左右
                配置如下
                cpu: AMD Ryzen 7 2700 Eight-Core Processor (8核16线程, 3.2GHz)
                RAM: 64GB DDR4
                GPU: AMD Radeon RX 7900 XTX 24GB (RADV NAVI31)
                OS: Ubuntu 26.04 LTS (内核 7.0.0-29-generic)
                软件 : llama.cpp Vulkan b10486

                懒人烘培懒 离线
                懒人烘培懒 离线
                懒人烘培
                编写于 最后由 编辑
                #25

                @6cccccc
                试试我那参数呢,不应该只有12

                1 条回复 最后回复
                0
                • 6 6cccccc

                  为啥我的启动参数是这样:
                  -m "$SELECTED_MODEL" -t 10 -b 512 -ub 256
                  --spec-draft-n-max 3 --fit off --no-context-shift
                  --metrics --kv-unified --jinja
                  --cache-type-k q8_0 --cache-type-v q8_0
                  -fa on --spec-type draft-mtp
                  --ctx-size 131072 --parallel 1 -ngl -1
                  --host 0.0.0.0 --port $LLAMA_PORT
                  --chat-template-kwargs {"enable_thinking": true, "preserve_think": false, "reasoning_effort": "medium"}
                  --no-mmap --temp 0.6 --top-p 0.5 --top-k 15 --repeat-penalty 1.0
                  --alias qwen3.8-27b-q5
                  仍然速度很低,只有12左右
                  配置如下
                  cpu: AMD Ryzen 7 2700 Eight-Core Processor (8核16线程, 3.2GHz)
                  RAM: 64GB DDR4
                  GPU: AMD Radeon RX 7900 XTX 24GB (RADV NAVI31)
                  OS: Ubuntu 26.04 LTS (内核 7.0.0-29-generic)
                  软件 : llama.cpp Vulkan b10486

                  E 离线
                  E 离线
                  exllm
                  德高望重
                  编写于 最后由 编辑
                  #26

                  @6cccccc 你BIOS开above 4G decoding了吗?

                  terryT 6 2 条回复 最后回复
                  0
                  • E exllm

                    @6cccccc 你BIOS开above 4G decoding了吗?

                    terryT 离线
                    terryT 离线
                    terry
                    超级版主
                    编写于 最后由 编辑
                    #27

                    @exllm 开不开影响都不大,我实测没感觉出来区别,但是有大神测了有区别。

                    油管:https://www.youtube.com/@抡锤者

                    1 条回复 最后回复
                    0
                    • 6 6cccccc

                      为啥我的启动参数是这样:
                      -m "$SELECTED_MODEL" -t 10 -b 512 -ub 256
                      --spec-draft-n-max 3 --fit off --no-context-shift
                      --metrics --kv-unified --jinja
                      --cache-type-k q8_0 --cache-type-v q8_0
                      -fa on --spec-type draft-mtp
                      --ctx-size 131072 --parallel 1 -ngl -1
                      --host 0.0.0.0 --port $LLAMA_PORT
                      --chat-template-kwargs {"enable_thinking": true, "preserve_think": false, "reasoning_effort": "medium"}
                      --no-mmap --temp 0.6 --top-p 0.5 --top-k 15 --repeat-penalty 1.0
                      --alias qwen3.8-27b-q5
                      仍然速度很低,只有12左右
                      配置如下
                      cpu: AMD Ryzen 7 2700 Eight-Core Processor (8核16线程, 3.2GHz)
                      RAM: 64GB DDR4
                      GPU: AMD Radeon RX 7900 XTX 24GB (RADV NAVI31)
                      OS: Ubuntu 26.04 LTS (内核 7.0.0-29-generic)
                      软件 : llama.cpp Vulkan b10486

                      懒人烘培懒 离线
                      懒人烘培懒 离线
                      懒人烘培
                      编写于 最后由 编辑
                      #28

                      @6cccccc
                      刚才实际测试了,如果开启Thinking了,速度也只有16,把Thinking Off后,速度我就恢复了70了。注意

                      6 1 条回复 最后回复
                      0
                      • Michael GillM 离线
                        Michael GillM 离线
                        Michael Gill
                        编写于 最后由 编辑
                        #29

                        @exllm @terry 感谢二位,以为本地部署没啥希望,刚好是7900xtx,简单测试

                        fca927c4-1f8f-4a08-b70c-feeb467a8ba1-image.jpeg

                        9383fec9-16fa-4792-b1dd-d25cc6dcaad5-image.jpeg

                        1 条回复 最后回复
                        0
                        • farmer nodeF 离线
                          farmer nodeF 离线
                          farmer node
                          编写于 最后由 编辑
                          #30
                          此主題已被删除!
                          1 条回复 最后回复
                          0
                          • E exllm

                            @6cccccc 你BIOS开above 4G decoding了吗?

                            6 离线
                            6 离线
                            6cccccc
                            编写于 最后由 编辑
                            #31

                            @exllm 我一会试试

                            1 条回复 最后回复
                            0
                            • 懒人烘培懒 懒人烘培

                              @6cccccc
                              刚才实际测试了,如果开启Thinking了,速度也只有16,把Thinking Off后,速度我就恢复了70了。注意

                              6 离线
                              6 离线
                              6cccccc
                              编写于 最后由 编辑
                              #32

                              @懒人烘培 嗯,我一会试试,好像开了后,mtp命中率很低

                              1 条回复 最后回复
                              0
                              • L 离线
                                L 离线
                                laobenxiong
                                德高望重 劳动模范
                                编写于 最后由 laobenxiong 编辑
                                #33

                                感谢楼主. 7900xtx 能跑,不错. 明天试试dsh. 今天dsh+v4pro花了100多, 肉疼(不过确实强, 长链自主调试FPGA, 综合+ila probe+烧写+测试+uart输出分析一条龙, 我基本可以不用管, 花token就行).
                                802712c0-6935-4c19-b583-ca3851ade37f-image.jpeg

                                bruin@lmde7 ~ $ ./run-3.8-q5.sh
                                0.00.035.069 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead.
                                0.00.035.097 W DEPRECATED: --mmap and --no-mmap are deprecated. use --load-mode mmap instead
                                0.00.039.005 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
                                0.00.039.491 W srv  llama_server: -----------------
                                0.00.039.495 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
                                0.00.039.495 W srv  llama_server: this can be a security risk (cross-origin attacks)
                                0.00.039.495 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
                                0.00.039.495 W srv  llama_server: -----------------
                                0.00.040.765 I srv    load_model: loading model '/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q5_K_S.gguf'
                                0.08.836.391 I cmn          init: llama threadpool init, n_threads = 10
                                0.09.247.607 I common_speculative_init_result: creating MTP draft context against the target model '/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q5_K_S.gguf'
                                0.09.318.357 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 131072, kv_unified = 'true'
                                0.09.457.688 I srv          init: chat template supports preserving reasoning, consider enabling it via --reasoning-preserve
                                0.09.457.746 I srv  llama_server: model loaded
                                0.09.457.750 I srv  llama_server: listening on http://0.0.0.0:8000
                                0.31.502.907 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
                                0.31.503.144 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
                                0.33.674.917 I slot print_timing: id  0 | task 0 | prompt eval time =    1607.08 ms /   333 tokens (    4.83 ms per token,   207.21 tokens per second)
                                0.33.674.922 I slot print_timing: id  0 | task 0 |        eval time =     564.47 ms /    34 tokens (   17.11 ms per token,    58.46 tokens per second)
                                0.33.674.923 I slot print_timing: id  0 | task 0 |       total time =    2171.56 ms /   367 tokens
                                0.33.674.928 I slot print_timing: id  0 | task 0 |    graphs reused =         11
                                0.33.674.930 I slot print_timing: id  0 | task 0 | draft acceptance = 0.72727 (   24 accepted /    33 generated), mean len =  3.18
                                0.33.674.999 I slot      release: id  0 | task 0 | stop processing: n_tokens = 368, truncated = 0
                                0.39.250.971 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.963 (> 0.100 thold), f_keep = 1.000
                                0.39.251.208 I slot launch_slot_: id  0 | task 16 | processing task, is_child = 0
                                0.40.592.295 I slot print_timing: id  0 | task 16 | prompt eval time =     368.90 ms /    14 tokens (   26.35 ms per token,    37.95 tokens per second)
                                0.40.592.303 I slot print_timing: id  0 | task 16 |        eval time =     972.01 ms /    67 tokens (   14.73 ms per token,    67.90 tokens per second)
                                0.40.592.303 I slot print_timing: id  0 | task 16 |       total time =    1340.91 ms /    81 tokens
                                0.40.592.305 I slot print_timing: id  0 | task 16 |    graphs reused =         31
                                0.40.592.307 I slot print_timing: id  0 | task 16 | draft acceptance = 0.71429 (   45 accepted /    63 generated), mean len =  3.14
                                0.40.592.367 I slot      release: id  0 | task 16 | stop processing: n_tokens = 448, truncated = 0
                                

                                llama.cpp 我用的最新 b10485; 完全抄作业:

                                bruin@lmde7 ~ $ cat run-3.8-q5.sh
                                #!/bin/bash
                                
                                # ref: https://lcz.me/topic/1157/qwen3.8-27b-q5_k_m-7900xtx-deepseek-harness%E5%AE%9E%E6%88%98%E5%B9%B3%E5%9D%87-52-t-s
                                
                                LLAMA_SERVER=/home/bruin/llama-server-vulkan-b10485
                                MAIN_MODEL="/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q5_K_S.gguf"
                                MTMD_MODEL="/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/mmproj-F16.gguf"
                                #--mmproj "${MTMD_MODEL}" \
                                
                                ${LLAMA_SERVER} \
                                --device Vulkan0 \
                                --model "${MAIN_MODEL}" \
                                -t 10 \
                                -b 512 \
                                -ub 256 \
                                --spec-draft-n-max 3 \
                                --fit off \
                                --no-context-shift \
                                --metrics \
                                --kv-unified \
                                --jinja \
                                --cache-type-k q8_0 \
                                --cache-type-v q8_0 \
                                -fa on \
                                --spec-type draft-mtp \
                                --ctx-size 131072 \
                                --parallel 1 \
                                -ngl -1 \
                                --host 0.0.0.0 \
                                --port 8000 \
                                --chat-template-kwargs '{"enable_thinking": true, "preserve_think": false, "reasoning_effort": "medium"}' \
                                --no-mmap \
                                --temp 0.6 \
                                --top-p 0.5 \
                                --top-k 15 \
                                --repeat-penalty 1.0 \
                                --override-tensor blk\.\d+\.ffn_.*_exps\.=CPU \
                                --alias qwen3.8-27b-q5
                                

                                b7d3fd10-965f-4fff-86c9-7e0b456f7dd6-image.jpeg

                                6 1 条回复 最后回复
                                1
                                • L laobenxiong

                                  感谢楼主. 7900xtx 能跑,不错. 明天试试dsh. 今天dsh+v4pro花了100多, 肉疼(不过确实强, 长链自主调试FPGA, 综合+ila probe+烧写+测试+uart输出分析一条龙, 我基本可以不用管, 花token就行).
                                  802712c0-6935-4c19-b583-ca3851ade37f-image.jpeg

                                  bruin@lmde7 ~ $ ./run-3.8-q5.sh
                                  0.00.035.069 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead.
                                  0.00.035.097 W DEPRECATED: --mmap and --no-mmap are deprecated. use --load-mode mmap instead
                                  0.00.039.005 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
                                  0.00.039.491 W srv  llama_server: -----------------
                                  0.00.039.495 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
                                  0.00.039.495 W srv  llama_server: this can be a security risk (cross-origin attacks)
                                  0.00.039.495 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
                                  0.00.039.495 W srv  llama_server: -----------------
                                  0.00.040.765 I srv    load_model: loading model '/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q5_K_S.gguf'
                                  0.08.836.391 I cmn          init: llama threadpool init, n_threads = 10
                                  0.09.247.607 I common_speculative_init_result: creating MTP draft context against the target model '/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q5_K_S.gguf'
                                  0.09.318.357 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 131072, kv_unified = 'true'
                                  0.09.457.688 I srv          init: chat template supports preserving reasoning, consider enabling it via --reasoning-preserve
                                  0.09.457.746 I srv  llama_server: model loaded
                                  0.09.457.750 I srv  llama_server: listening on http://0.0.0.0:8000
                                  0.31.502.907 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
                                  0.31.503.144 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
                                  0.33.674.917 I slot print_timing: id  0 | task 0 | prompt eval time =    1607.08 ms /   333 tokens (    4.83 ms per token,   207.21 tokens per second)
                                  0.33.674.922 I slot print_timing: id  0 | task 0 |        eval time =     564.47 ms /    34 tokens (   17.11 ms per token,    58.46 tokens per second)
                                  0.33.674.923 I slot print_timing: id  0 | task 0 |       total time =    2171.56 ms /   367 tokens
                                  0.33.674.928 I slot print_timing: id  0 | task 0 |    graphs reused =         11
                                  0.33.674.930 I slot print_timing: id  0 | task 0 | draft acceptance = 0.72727 (   24 accepted /    33 generated), mean len =  3.18
                                  0.33.674.999 I slot      release: id  0 | task 0 | stop processing: n_tokens = 368, truncated = 0
                                  0.39.250.971 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.963 (> 0.100 thold), f_keep = 1.000
                                  0.39.251.208 I slot launch_slot_: id  0 | task 16 | processing task, is_child = 0
                                  0.40.592.295 I slot print_timing: id  0 | task 16 | prompt eval time =     368.90 ms /    14 tokens (   26.35 ms per token,    37.95 tokens per second)
                                  0.40.592.303 I slot print_timing: id  0 | task 16 |        eval time =     972.01 ms /    67 tokens (   14.73 ms per token,    67.90 tokens per second)
                                  0.40.592.303 I slot print_timing: id  0 | task 16 |       total time =    1340.91 ms /    81 tokens
                                  0.40.592.305 I slot print_timing: id  0 | task 16 |    graphs reused =         31
                                  0.40.592.307 I slot print_timing: id  0 | task 16 | draft acceptance = 0.71429 (   45 accepted /    63 generated), mean len =  3.14
                                  0.40.592.367 I slot      release: id  0 | task 16 | stop processing: n_tokens = 448, truncated = 0
                                  

                                  llama.cpp 我用的最新 b10485; 完全抄作业:

                                  bruin@lmde7 ~ $ cat run-3.8-q5.sh
                                  #!/bin/bash
                                  
                                  # ref: https://lcz.me/topic/1157/qwen3.8-27b-q5_k_m-7900xtx-deepseek-harness%E5%AE%9E%E6%88%98%E5%B9%B3%E5%9D%87-52-t-s
                                  
                                  LLAMA_SERVER=/home/bruin/llama-server-vulkan-b10485
                                  MAIN_MODEL="/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q5_K_S.gguf"
                                  MTMD_MODEL="/opt/gguf-models/unsloth/Qwen3.8-27B-GGUF/mmproj-F16.gguf"
                                  #--mmproj "${MTMD_MODEL}" \
                                  
                                  ${LLAMA_SERVER} \
                                  --device Vulkan0 \
                                  --model "${MAIN_MODEL}" \
                                  -t 10 \
                                  -b 512 \
                                  -ub 256 \
                                  --spec-draft-n-max 3 \
                                  --fit off \
                                  --no-context-shift \
                                  --metrics \
                                  --kv-unified \
                                  --jinja \
                                  --cache-type-k q8_0 \
                                  --cache-type-v q8_0 \
                                  -fa on \
                                  --spec-type draft-mtp \
                                  --ctx-size 131072 \
                                  --parallel 1 \
                                  -ngl -1 \
                                  --host 0.0.0.0 \
                                  --port 8000 \
                                  --chat-template-kwargs '{"enable_thinking": true, "preserve_think": false, "reasoning_effort": "medium"}' \
                                  --no-mmap \
                                  --temp 0.6 \
                                  --top-p 0.5 \
                                  --top-k 15 \
                                  --repeat-penalty 1.0 \
                                  --override-tensor blk\.\d+\.ffn_.*_exps\.=CPU \
                                  --alias qwen3.8-27b-q5
                                  

                                  b7d3fd10-965f-4fff-86c9-7e0b456f7dd6-image.jpeg

                                  6 离线
                                  6 离线
                                  6cccccc
                                  编写于 最后由 编辑
                                  #34

                                  @laobenxiong 确实,用了你这个q5ks的模型,我的速度终于正常了,也不用关thinking,不知道为什么其他人的q5km可以

                                  1 条回复 最后回复
                                  0
                                  • X 离线
                                    X 离线
                                    xl
                                    编写于 最后由 编辑
                                    #35

                                    关于--parallel 2 的测试

                                    软硬件环境:
                                    cpu: AMD Ryzen 7 9700X
                                    RAM:DDR5 64GB 6000
                                    GPU:R9700 32G
                                    OS: Ubuntu 24.04
                                    软件 : llama.cpp + vulkan

                                    结论:将并行设为1或者2,大部分情况下,性能变化属于正常波动,没有明显差别。

                                    但是,即使有32G显存,有些任务还是会报以下错误:
                                    26.48.401.777 E init_batch: failed to prepare attention ubatches
                                    26.48.401.778 W decode: failed to find a memory slot for batch of size 16
                                    26.48.401.779 W srv decode: failed to find free space in the KV cache, retrying with smaller batch size, off = 190, n_batch = 8, ret = 1
                                    26.48.403.275 E init_batch: failed to prepare attention ubatches
                                    26.48.403.277 W decode: failed to find a memory slot for batch of size 8
                                    26.48.403.277 W srv decode: failed to find free space in the KV cache, retrying with smaller batch size, off = 190, n_batch = 4, ret = 1
                                    虽然使用了 --cache-type-k/v q8_0 做 KV 量化来节省显存,但是 128K 的超长上下文,加上 --parallel 2,对 KV Cache 的要求依然极其恐怖。模型权重稳稳地占了 17GB。但当开启并发,且 Prompt 超长时,给 128K 上下文预留的 KV Cache 空间,在高峰期(特别是并发请求时)瞬间耗尽了,导致没有空位给下一个 Token 的生成。
                                    好消息是,llama.cpp不能为KV Cache分配空间,并不会导致崩溃或任务中断(我没有碰到过),只会导致llama.cpp降级重试,严重的情况下会感觉明显卡顿而已,不用过于担心。
                                    所以,用7900XTX的朋友,把并发设为1可能是更好的选择。

                                    1 条回复 最后回复
                                    0
                                    • ,系统 取消固定了此主题
                                    • E 离线
                                      E 离线
                                      Enigma
                                      编写于 最后由 编辑
                                      #36

                                      感谢分享,学习中。。。

                                      1 条回复 最后回复
                                      0

                                      你好!看起来您对这段对话很感兴趣,但您还没有一个账号。

                                      厌倦了每次访问都刷到同样的帖子?您注册账号后,您每次返回时都能精准定位到您上次浏览的位置,并可选择接收新回复通知(通过邮件或推送通知)。您还能收藏书签、为帖子顶,向社区成员表达您的欣赏。

                                      有了你的建议,这篇帖子会更精彩哦 💗

                                      注册 登录
                                      回复
                                      • 在新帖中回复
                                      登录后回复
                                      • 从旧到新
                                      • 从新到旧
                                      • 最多赞同


                                      • 登录

                                      • 登录或注册以进行搜索。
                                      • 第一个帖子
                                        最后一个帖子
                                      0
                                      • 版块
                                      • 最新
                                      • 标签
                                      • 热门
                                      • 用户
                                      • 群组