Your browser does not seem to support JavaScript. As a result, your viewing experience will be diminished, and you have been placed in read-only mode.
Please download a browser that supports JavaScript, or enable it if it's disabled (i.e. NoScript).
@566656661 谢谢
目前使用 vLLM 推理 Qwen3.6-27B 模型,在 AWQ 量化及 MTP 设为 2 的情况下,Token 吞吐量稳定在 50 tokens/s 以上。请问该性能表现是否符合预期?是否存在进一步的优化空间?