<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[AMD AI Pro R9700 LLM调教]]></title><description><![CDATA[<p dir="auto">在座的各位大神，有没有关于在R9700上运行千问 3.6 27B 的时候各种框架时候的速度对比？<br />
拿它配合 Hermes 和其他的 Agent 使用的时候，用llamacpp Vulkan版本，PP 的速度还是太惨了，卡到不太想用。有没有大神试验过 vLLM 和 SGlang，跑起来速度怎么样？我在论坛里面找了一圈，也没有找到相关的材料，谢了！</p>
]]></description><link>https://lcz.me/topic/848/amd-ai-pro-r9700-llm调教</link><generator>RSS for Node</generator><lastBuildDate>Sun, 26 Jul 2026 20:02:06 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/848.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 14 Jul 2026 06:12:10 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Thu, 23 Jul 2026 03:30:04 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/luke-mao" aria-label="Profile: Luke-Mao">@<bdi>Luke-Mao</bdi></a><br />
其实也没什么，你就让Hermes帮你重新编译llamacpp就可以了。编译的过程当中开启WMMA，之后是用系统原生服务跑还是docker跑都可以。效果没差。</p>
]]></description><link>https://lcz.me/post/10325</link><guid isPermaLink="true">https://lcz.me/post/10325</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Thu, 23 Jul 2026 03:30:04 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Thu, 23 Jul 2026 03:29:02 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/linkdesu" aria-label="Profile: linkdesu">@<bdi>linkdesu</bdi></a><br />
谢谢，我之前就是看的这个，受到的启发，现在已经基本上稳定了。</p>
]]></description><link>https://lcz.me/post/10324</link><guid isPermaLink="true">https://lcz.me/post/10324</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Thu, 23 Jul 2026 03:29:02 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Wed, 22 Jul 2026 23:05:13 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> 能否详细讲讲这个wmma部署，谢谢！</p>
]]></description><link>https://lcz.me/post/10304</link><guid isPermaLink="true">https://lcz.me/post/10304</guid><dc:creator><![CDATA[Luke Mao]]></dc:creator><pubDate>Wed, 22 Jul 2026 23:05:13 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Sat, 18 Jul 2026 05:47:50 GMT]]></title><description><![CDATA[<p dir="auto"><a href="https://kyuz0.github.io/amd-r9700-ai-toolboxes/index.html" rel="nofollow ugc">https://kyuz0.github.io/amd-r9700-ai-toolboxes/index.html</a> 这里有个双 9700 在不同驱动下的 benchmark ，希望也能有帮助</p>
]]></description><link>https://lcz.me/post/10061</link><guid isPermaLink="true">https://lcz.me/post/10061</guid><dc:creator><![CDATA[linkdesu]]></dc:creator><pubDate>Sat, 18 Jul 2026 05:47:50 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Fri, 17 Jul 2026 11:26:26 GMT]]></title><description><![CDATA[<p dir="auto">@倭寇国を滅ぼす<br />
哦 不是，pp 是 prompt processing，就是那个prefill所谓输入的预处理。可以理解为是输入token处理速度。</p>
]]></description><link>https://lcz.me/post/10029</link><guid isPermaLink="true">https://lcz.me/post/10029</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Fri, 17 Jul 2026 11:26:26 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Fri, 17 Jul 2026 09:28:45 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> 我意思是你指的PP具体是什么，是那个PP512 1024 2048吗？</p>
]]></description><link>https://lcz.me/post/10025</link><guid isPermaLink="true">https://lcz.me/post/10025</guid><dc:creator><![CDATA[用户名违规]]></dc:creator><pubDate>Fri, 17 Jul 2026 09:28:45 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Fri, 17 Jul 2026 09:19:36 GMT]]></title><description><![CDATA[<p dir="auto">@倭寇国を滅ぼす<br />
用agent啊，做个测试脚本，每次服务拉起来了跑一遍脚本就有了。</p>
]]></description><link>https://lcz.me/post/10022</link><guid isPermaLink="true">https://lcz.me/post/10022</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Fri, 17 Jul 2026 09:19:36 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Fri, 17 Jul 2026 09:18:52 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/iamvirus" aria-label="Profile: iamvirus">@<bdi>iamvirus</bdi></a><br />
我的是R9700说是有个什么vllm不支持，就很恶心，速度比llamacpp vulkan都不如。开启wmma后用MoE直接起飞！</p>
]]></description><link>https://lcz.me/post/10021</link><guid isPermaLink="true">https://lcz.me/post/10021</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Fri, 17 Jul 2026 09:18:52 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Fri, 17 Jul 2026 08:46:46 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> PP怎么测试，sglang ，prefill每次16K或者8K，prefill beth在1000-1400t/s，说的是这个吗？</p>
]]></description><link>https://lcz.me/post/10018</link><guid isPermaLink="true">https://lcz.me/post/10018</guid><dc:creator><![CDATA[用户名违规]]></dc:creator><pubDate>Fri, 17 Jul 2026 08:46:46 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Thu, 16 Jul 2026 03:03:32 GMT]]></title><description><![CDATA[<p dir="auto">再买一张组tp=2 的vllm体验会很丝滑</p>
]]></description><link>https://lcz.me/post/9970</link><guid isPermaLink="true">https://lcz.me/post/9970</guid><dc:creator><![CDATA[iamvirus]]></dc:creator><pubDate>Thu, 16 Jul 2026 03:03:32 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Wed, 15 Jul 2026 23:44:04 GMT]]></title><description><![CDATA[<p dir="auto">@倭寇国を滅ぼす<br />
使用来讲的话，TP25 或者以上就挺不错的了。</p>
<p dir="auto">但是因为 Agent 里应用的那个系统提示词啊，和来回工具调用的那个上下文很长，所以 PP 才是最最关键的。那个如果 PP 慢了的话，等消息简直等到天荒地老。</p>
]]></description><link>https://lcz.me/post/9961</link><guid isPermaLink="true">https://lcz.me/post/9961</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Wed, 15 Jul 2026 23:44:04 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Wed, 15 Jul 2026 11:40:07 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> PP不知道150K上下文速度能打50t/s，200k的上下文时候大概也是在50t/s,这个有mtp。如果不开MTP稳定100K内31t/s，超过150k 27t/s的样子。速度不能和云API对比</p>
]]></description><link>https://lcz.me/post/9944</link><guid isPermaLink="true">https://lcz.me/post/9944</guid><dc:creator><![CDATA[用户名违规]]></dc:creator><pubDate>Wed, 15 Jul 2026 11:40:07 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Wed, 15 Jul 2026 07:21:46 GMT]]></title><description><![CDATA[<p dir="auto">@倭寇国を滅ぼす<br />
你这个跑起来pp能到多少？就是尤其输入比较大的时候，比如说16K32k这种，这种比较接近agent使用时候的真实场景。  有点奇怪，我做研究的时候都说SGlang是三个主要框架里面速度最慢的，求个分享。</p>
]]></description><link>https://lcz.me/post/9923</link><guid isPermaLink="true">https://lcz.me/post/9923</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Wed, 15 Jul 2026 07:21:46 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Wed, 15 Jul 2026 07:18:53 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/566656661" aria-label="Profile: 566656661">@<bdi>566656661</bdi></a>  我在GitHub上下载了一个类似的里面下载量最高的一个，最后实际测试下来v l l m在r9700上性能不如llamacpp vulkan版本。<br />
最后还是老老实实用vulkan吧，毕竟跑35b q5的模型，开三路，pp都能到2300，tp能到50~55。体验已经相当不错了，如果要是能把这个pp再提高到4000左右就比较爽了，agent应用太吃pp了。不过好像也就5090能做到这个水平，但是这破玩意儿现在是真的贵。</p>
]]></description><link>https://lcz.me/post/9922</link><guid isPermaLink="true">https://lcz.me/post/9922</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Wed, 15 Jul 2026 07:18:53 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Tue, 14 Jul 2026 18:01:16 GMT]]></title><description><![CDATA[<p dir="auto">@倭寇国を滅ぼす SG-Lang是最终解决方案，主要是Prefill速度提升很多，说实话，只有这个框架有一些本地部署的意义。Qwen3.6和DeepSeek V4 Flash差距不大，各有优劣，它反正不蠢，文档写好了是完全够用的，工具属性也足，驱动Hermes挺不错的。</p>
]]></description><link>https://lcz.me/post/9890</link><guid isPermaLink="true">https://lcz.me/post/9890</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Tue, 14 Jul 2026 18:01:16 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Tue, 14 Jul 2026 14:33:02 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a></p>
<p dir="auto">vLLM的ROCm我記得是有支持AITER加速 (AMD自家Tensor Core, 應該是透過分拆Kernel來增加Decode速度)</p>
<p dir="auto">可以找找看有沒有人特意出vLLM 基於AITER版本的Docker Image, aml731我記得也有出, 就是不知道有沒有更新</p>
]]></description><link>https://lcz.me/post/9888</link><guid isPermaLink="true">https://lcz.me/post/9888</guid><dc:creator><![CDATA[566656661]]></dc:creator><pubDate>Tue, 14 Jul 2026 14:33:02 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Tue, 14 Jul 2026 13:18:47 GMT]]></title><description><![CDATA[<p dir="auto">这个是我调了一周的sglang配置。<br />
主要目的是单线程，没有MTP，速度30t/s,有500K上下文。<br />
上了MTP，速度有60t/s，有300k上下文。最主要的是QWEN3.6-27B-FP8也不聪明。。但是又觉得有时候比云API聪明。<br />
主要是我的soul。md 明确了，所有问题优先社区找到答案。。处理问题的能力提升了一些。</p>
<pre><code>#!/bin/bash
# SGLang 0.5.15 — sglang serve (官方推荐方式)
# Qwen3.6-27B-MTP | RTX 4090D 48G
set -euo pipefail

# 路径与环境
readonly SGLANG_VENV="/mnt/dataM4/venv_sglang"
readonly MODEL_PATH="/mnt/dataM4/models/Qwen3.6-27B-AEON-Ultimate-Uncensored-FP8-MTP"

export HF_HOME="/root/.cache/huggingface"
export TRITON_CACHE_DIR="/root/sglang_cache/triton/cache"
export TORCHINDUCTOR_CACHE_DIR="/root/sglang_cache/torch_compile"
export FLASH_ATTENTION_CUTE_CACHE_DIR="/root/sglang_cache/flash_attn_cute_cache"

export CUDA_HOME=/usr/local/cuda-12.8
export PATH="$CUDA_HOME/bin:$PATH"
export LD_LIBRARY_PATH="$CUDA_HOME/lib64"
export LD_LIBRARY_PATH="/mnt/dataM4/venv_sglang/lib/python3.11/site-packages/nvidia/cusparselt/lib:$LD_LIBRARY_PATH"
export CUDAHOSTCXX=/usr/bin/gcc-12
export CC=/usr/bin/gcc-12
export CXX=/usr/bin/g++-12

export PYTORCH_ALLOC_CONF="expandable_segments:True"
export CUDA_DEVICE_ORDER="PCI_BUS_ID"
export CUDA_VISIBLE_DEVICES="0"
export FLASHINFER_DISABLE_VERSION_CHECK=1
export FLASH_ATTENTION_CUTE_DSL_CACHE_ENABLED=1

source "${SGLANG_VENV}/bin/activate"

SGLANG_ARGS=(
  --model-path "${MODEL_PATH}"
  --host 0.0.0.0
  --port 30000
  --trust-remote-code
  --tp-size 1
  --attention-backend flashinfer
  --mamba-backend flashinfer
  --kv-cache-dtype fp8_e4m3
  --page-size 16
  --mem-fraction-static 0.93
  --max-running-requests 1
  --prefill-max-requests 1
  --schedule-policy lpm
  --schedule-conservativeness 1.0
  --mamba-radix-cache-strategy extra_buffer
  --mamba-ssm-dtype bfloat16
  --mamba-full-memory-ratio 0.08
  --speculative-algorithm nextn
  --speculative-num-steps 2
  --speculative-eagle-topk 1
  --speculative-num-draft-tokens 2
  --tool-call-parser qwen3_coder
  --reasoning-parser qwen3
  --preferred-sampling-params '{"temperature": 0.7, "repetition_penalty": 1.2, "frequency_penalty": 0.5, "presence_penalty": 0.3}'
  --enable-dynamic-chunking
  --chunked-prefill-size 16384
  --radix-eviction-policy lru
  --enable-metrics
  --enable-streaming-session
  --enable-request-time-stats-logging
  --log-requests
  --log-requests-level 0
  --decode-log-interval 60
)

# 启动
readonly SGLOG="/dev/shm/sglang/sglang.log"
mkdir -p "$(dirname "$SGLOG")"

# 日志轮转
if ! pgrep -f sglang_log_rotate &gt; /dev/null 2&gt;&amp;1; then
    nohup bash /root/sglang_log_rotate.sh &gt;/dev/null 2&gt;&amp;1 &amp;
fi

exec sglang serve "${SGLANG_ARGS[@]}" &gt;&gt;"$SGLOG" 2&gt;&amp;1

</code></pre>
]]></description><link>https://lcz.me/post/9885</link><guid isPermaLink="true">https://lcz.me/post/9885</guid><dc:creator><![CDATA[用户名违规]]></dc:creator><pubDate>Tue, 14 Jul 2026 13:18:47 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Tue, 14 Jul 2026 13:14:39 GMT]]></title><description><![CDATA[<p dir="auto">京东自营买的的这个卡，到手一天后，自己退货了。奇葩的是，32g，是真的到高不下的。。最后换了48g。</p>
]]></description><link>https://lcz.me/post/9884</link><guid isPermaLink="true">https://lcz.me/post/9884</guid><dc:creator><![CDATA[用户名违规]]></dc:creator><pubDate>Tue, 14 Jul 2026 13:14:39 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Tue, 14 Jul 2026 12:12:41 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/williamlouis" aria-label="Profile: williamlouis">@<bdi>williamlouis</bdi></a>  谢谢我找找看</p>
]]></description><link>https://lcz.me/post/9883</link><guid isPermaLink="true">https://lcz.me/post/9883</guid><dc:creator><![CDATA[fcme]]></dc:creator><pubDate>Tue, 14 Jul 2026 12:12:41 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Tue, 14 Jul 2026 09:17:07 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> 有一个双 R9700 32G的帖子。配置 SGlang 成功。你再找找</p>
]]></description><link>https://lcz.me/post/9876</link><guid isPermaLink="true">https://lcz.me/post/9876</guid><dc:creator><![CDATA[williamlouis]]></dc:creator><pubDate>Tue, 14 Jul 2026 09:17:07 GMT</pubDate></item><item><title><![CDATA[Reply to AMD AI Pro R9700 LLM调教 on Tue, 14 Jul 2026 07:16:36 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/fcme" aria-label="Profile: fcme">@<bdi>fcme</bdi></a> 关于 R9700 上跑 Qwen 3.6 27B 的框架速度对比，分享一些实测经验：</p>
<p dir="auto"><strong>框架选择顺序：ROCm HIP &gt; Vulkan</strong></p>
<p dir="auto">llama.cpp 的 Vulkan 后端在 PP（Prompt Processing）阶段确实慢，主要原因是 Vulkan 的算子融合不如 ROCm/HIP 充分。建议优先上 ROCm：</p>
<ol>
<li>
<p dir="auto"><strong>llama.cpp + ROCm/HIP</strong> — 这是目前 R9700 上最快的方案。PP 速度比 Vulkan 快约 2-3 倍，TG 也能有 20-30+ T/S。需要 ROCm 6.3+（支持 gfx1200/gfx1201）。</p>
<ul>
<li>编译：<code>cmake -DCMAKE_BUILD_TYPE=Release -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1200 ..</code></li>
<li>ROCm 安装：可以用 amdgpu-install 或者直接装 rocminfo + rocm-hip-sdk</li>
</ul>
</li>
<li>
<p dir="auto"><strong>vLLM + ROCm</strong> — 技术上可行但配置门槛高。vLLM 对 AMD 的支持在持续改进中，但 R9700 (RDNA4) 的 ROCm 支持还在早期。需要自行编译 rocm 版本的 vLLM，且 flash attention 在 RDNA4 上可能有兼容性问题。好处是支持 continuous batching，适合多用户场景。</p>
</li>
<li>
<p dir="auto"><strong>SGlang + AMD</strong> — SGlang 目前对 AMD 的支持比 vLLM 更有限，主要针对 MI 系列计算卡开发，RDNA4 上不建议尝试，坑比较多。</p>
</li>
</ol>
<p dir="auto"><strong>配合 Hermes Agent 的建议：</strong></p>
<ul>
<li>用 llama.cpp + HIP 后端 + <code>--no-kv-offload</code> 控制显存</li>
<li>Q4_K_M 量化下 Qwen 3.6 27B 约 16-18GB 显存，R9700 32GB 留 14GB+ 给 KV Cache，256K 上下文没问题</li>
<li>启动后用 <code>--host 0.0.0.0 --port 8080</code> 暴露 OpenAI 兼容 API，Hermes 里配置 <code>openai</code> provider 指向本地地址即可</li>
</ul>
<p dir="auto"><strong>ROCm 安装注意事项：</strong></p>
<ul>
<li>R9700 需要 ROCm 6.3+ 才有 gfx1200 支持</li>
<li>建议用 Ubuntu 24.04 或 Arch Linux，驱动支持最好</li>
<li>装完后 <code>rocminfo</code> 确认能识别显卡</li>
</ul>
<p dir="auto">如果不想折腾 ROCm，Vulkan 后端也能用，但建议把 <code>-ub</code> （ublock）关掉，context size 设小一点（128K），PP 会快一些。</p>
]]></description><link>https://lcz.me/post/9869</link><guid isPermaLink="true">https://lcz.me/post/9869</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Tue, 14 Jul 2026 07:16:36 GMT</pubDate></item></channel></rss>