12 KiB
12 KiB
x86 Rust Align Optimization Plan
当前结论
- 当前工作区的 vendored Rust backend 已经收敛到更接近 upstream
huanglizhuo/QwenASR的 Linux/x86_64 路径:releaseRUSTFLAGS="-C target-cpu=native"BLAS/OpenBLAS- x86_64 默认
BF16decode
- 保留的有意偏离只有两类:
SharedQwenModel/ shared model cache- 中性的
ffifeature(macos-ffi仅作为兼容别名保留)
upstream 参考
- upstream repo:
https://github.com/huanglizhuo/QwenASR - inspected commit:
4e85a19b05f034e106a345d279c68f50df718ab8
本机环境
- CPU:
Intel Core i5-13600KF - visible CPUs:
14 - memory: user reported
G.SKILL DDR5-6400
已验证 benchmark
音频:
/opt/qwen3-asr/temp/test_assets/podcast_demo_2min_16k.wav- duration:
120s
decode 路径对照(runtime concurrency = 4)
INT8 decode
- total:
173.42s - asr:
54.69s - align:
118.74s - rtf:
1.4452
来源:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_int8.json
BF16 decode
- total:
125.87s - asr:
46.40s - align:
79.47s - rtf:
1.0489
来源:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_bf16.json
结论:
- 在当前这台 x86_64 机器上,
BF16 decode明显优于INT8 decode - 因此 x86_64 默认 decode 路径应保持
BF16
收敛后的默认路径 benchmark
来源:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_default_after_converge.json
结果:
runtime concurrency = 4- total:
131.77s - asr:
51.54s - align:
80.23s - rtf:
1.0981
- total:
runtime concurrency = 14- total:
128.25s - asr:
57.12s - align:
71.13s - rtf:
1.0688
- total:
结论:
- 当前主瓶颈仍然在
align align_sec明显大于或接近asr_secruntime concurrency的最优值并不稳定,说明问题不在单纯线程数,而在具体阶段的访问模式和 kernel 行为
为什么不继续默认走 INT8
- upstream README 对 Linux/x86_64 的主路径描述是
BLAS + AVX2/FMA - 当前本机实测中,
INT8 decode明显慢于BF16 decode - 说明这台机器上的主瓶颈不只是权重带宽,更多是:
- x86_64 上 INT8 kernel 的有效带宽利用率
- cache / 数据布局
- 实现成熟度差异
当前仍需保留的偏离
1. Shared model cache
文件:
/opt/qwen3-asr/vendor/qwenasr/crates/qwen-asr/src/context.rs
目的:
- 多 runtime / 多 worker 场景下复用只读模型权重
- 避免每个 runtime 重复 mmap / 持有整套权重
2. ffi feature
文件:
/opt/qwen3-asr/vendor/qwenasr/crates/qwen-asr/Cargo.toml/opt/qwen3-asr/vendor/qwenasr/crates/qwen-asr/src/lib.rs/opt/qwen3-asr/Dockerfile.cpu
目的:
- 让 Linux CPU 集成不再依赖命名不准确的
macos-ffi - 同时保留兼容别名,避免已有脚本立即失效
align 热点拆解计划
Phase 1: 阶段级 profiling
状态:已完成
目标:
- 先确认
align的主耗时究竟在哪一段
位置:
/opt/qwen3-asr/vendor/qwenasr/crates/qwen-asr/src/align.rs/opt/qwen3-asr/vendor/qwenasr/crates/qwen-asr/src/decoder.rs
需要拆出的阶段:
mel_spectrogramencoder.forwardinput_embeds builddecoder_prefill_logitstimestamp argmax extractfix_timestamps
验收:
- 2 分钟样本上输出稳定的阶段级耗时表
实际结果:
- 产物:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_align_profile.json/opt/qwen3-asr/temp/test_logs/qwen_rust_runtime_concurrency_2min_align_profile.stderr
- 结论:
align的主热点明确落在decoder_prefill_logitsfinal rms_norm和lm_head projection不是主矛盾
Phase 2: decoder_prefill_logits 内部分解
状态:已完成
如果 decoder_prefill_logits 是主热点,则继续拆分:
decoder_prefill- final
rms_norm lm_head projection
目的:
- 判断到底是 decoder prefill 慢,还是最后分类头 projection 慢
实际结果:
- 产物:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_align_breakdown.json/opt/qwen3-asr/temp/test_logs/qwen_rust_runtime_concurrency_2min_align_breakdown.stderr
- 结论:
- 真正的大头是
decoder_prefill - 长段(
seq_len=1801)时,attention_ms占decoder_prefill的绝大部分 - 关键样本:
decoder_prefill total_ms=79280.78attention_ms=73769.91qkv_ms=1336.33gate_up_ms=1866.68down_proj_ms=955.94
- 真正的大头是
Phase 2.5: 失败尝试记录
状态:已完成并回退
尝试:
- 针对 x86_64 预先物化 prefill 用 F32 权重,避免每次
align反复做BF16 -> F32
结果:
- 长段
decoder_prefill没有稳定收益,反而出现回归 - 这条路径已经回退,不保留在主线代码里
结论:
- 当前瓶颈不是简单的 BF16 权重转换
- 更直接的问题是 multi-token causal attention 的算法路径
Phase 3: 对热点段做针对性优化
状态:第一轮已完成
根据 profiling 结果,按优先级选一个方向:
-
如果热点在
input_embeds build- 复用固定 prefix/suffix embeddings
- 减少逐 token 小块 copy
- 降低每段 align 的重复构造开销
-
如果热点在
decoder_prefill- 检查
BF16 matvec / attention / swiglu的实际热点 - 优化并行粒度或数据布局
- 检查
-
如果热点在
lm_head projection- 优先优化
BF16classify head 路径 - 避免无价值的整块 materialize
- 但只有在 profiling 证明是主热点后才动
- 优先优化
-
如果热点在后处理
- 精简
fix_timestamps - 减少
Vec/String分配
- 精简
Phase 3 实施结果
本轮实际落地的是:
- 文件:
/opt/qwen3-asr/vendor/qwenasr/crates/qwen-asr/src/kernels/mod.rs
- 改动:
- 对 BLAS multi-token causal attention 增加长序列专用 batched 路径
- 仅在
seq_q >= 256时启用 - 从“每个 head、每一行 2 次小 GEMM”改为:
- 每个 head 1 次
Q @ K^T - 行级 causal softmax
- 每个 head 1 次
softmax @ V
- 每个 head 1 次
Phase 3 回归结果
来源:
- 优化后 benchmark:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_after_attention_opt.json
- 优化后 profiling:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_attention_opt_profile.json/opt/qwen3-asr/temp/test_logs/qwen_rust_runtime_concurrency_2min_attention_opt_profile.stderr
关键对比:
- 之前默认路径(2 分钟样本,基线文件):
total=128.25sasr=57.12salign=71.13s
- 优化后:
total=76.46sasr=45.32salign=31.14srtf=0.6372
attention 热点变化:
- 长段
seq_len=1801- 优化前:
decoder_prefill total_ms=79280.78attention_ms=73769.91
- 优化后:
decoder_prefill total_ms=16625.88attention_ms=9981.88
- 优化前:
结论:
- 当前这台 i5 上,
align的主矛盾已经从“attention 明显失控”收敛到了“attention 仍是第一热点,但已降到可接受量级” - 这一轮优化是有效的,应该保留
Phase 4: FFN 路径继续收敛
状态:已完成一轮,并保留有效部分
本轮动作:
- 文件:
/opt/qwen3-asr/vendor/qwenasr/crates/qwen-asr/src/decoder.rs
- 改动:
- 仅对 x86_64
BF16 prefill的 FFN 路径物化共享 F32 权重 - 范围只包括:
gate_up_fuseddown_weight
QKV和O-proj暂不纳入
- 仅对 x86_64
回归数据:
- 不带 profiling:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_after_ffn_opt.jsontotal=58.61sasr=29.50salign=29.11s
- 带 profiling:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_ffn_opt_profile.jsontotal=62.07sasr=32.15salign=29.92s
与上一轮 attention-only profiling 对比:
- attention-only:
total=64.58salign=30.86s
- FFN-opt:
total=62.07salign=29.92s
长段热点对比(seq_len=1801):
- FFN-opt 之前:
attention_ms=9981.88gate_up_ms=2080.10down_proj_ms=1183.64
- FFN-opt 之后:
attention_ms=9660.15gate_up_ms=2210.39down_proj_ms=1166.96
结论:
- FFN 这轮不是“大收益”,但
align_sec仍然有小幅下降 - 收益不像 attention 优化那样压倒性,更像是小幅收敛
- 当前可以保留,但不值得继续在同一方向上扩大复杂度
Phase 5: QKV / O-proj 试验与回退
状态:已完成并回退
尝试:
- 在 prefill 中进一步物化
QKV和O-proj的共享 F32 权重
回归数据:
- 试验版本:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_after_qkv_opt.jsontotal=63.12salign=30.07s
结论:
- 相比 FFN-opt 版本,没有形成净收益
- 因此这条路径已回退,不保留在主线代码里
Phase 6: attention 继续细化的两次试验
状态:已完成并回退
试验 A:query-block batched attention
尝试:
- 将长序列 batched causal attention 从“整段一次性
QK^T/SV”改成按 query block 分块执行
结果:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_after_attention_block_opt.jsontotal=79.18salign=33.40s
结论:
- 这条路径在当前 i5 + OpenBLAS 组合下没有收益
- 增加 GEMM 次数带来的额外调度开销,超过了小块缓存收益
- 已回退
试验 B:提高 batched 切换阈值到 512
尝试:
- 只改
BATCHED_CAUSAL_ATTENTION_THRESHOLD - 让中等长度序列继续走 row-wise 路径,只把更长的序列交给 batched 路径
结果:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_after_attention_threshold512.jsontotal=74.49salign=30.42s
结论:
- 相比当前主线最优版本也没有收益
- 说明当前阈值
256不是主要问题 - 已恢复回
256
Phase 7: K/V block + online softmax 试验
状态:已完成并回退
尝试:
- 仅替换长序列 batched attention 路径
- 改为按
K/Vblock 流式累积的 online softmax - 短序列和单 token 路径完全不动
目标:
- 不再一次性物化整块
scores - 降低大矩阵内存压力
- 观察是否能进一步压低长段
attention_ms
结果:
/opt/qwen3-asr/temp/test_assets/qwen_rust_runtime_concurrency_2min_after_kvblock_online_softmax.jsontotal=74.11salign=30.33s
结论:
- 当前实现下没有优于主线最优版本
- 在这台机器上,额外的 block 循环和 online softmax 合并开销,超过了减少大
scores矩阵带来的收益 - 已回退
当前判断
- 当前最值钱、且已验证有效的优化仍然是:
- 长序列 batched causal attention
- FFN selective F32 物化
- 继续扩大到
QKV/O-proj这一步暂时不划算 - query-block 化和 batched 阈值调优目前也不划算
K/V block + online softmax在当前实现形态下也不划算- 后续判断应优先看:
align_secdecoder_prefillprofiling- 尤其是长段
seq_len下的热点变化
补充:
asr_sec在多次回归中波动明显大于align_sec- 因此后续评估优化效果时,不应只盯总耗时,应优先以
alignprofiling 为准
暂不做的事
- 不再把 x86_64 默认路径改回
INT8 decode - 不继续做没有 profiling 支撑的
align结构性改写 - 不围绕
runtime concurrency数量盲调
下一步执行顺序
- 保留当前 batched causal attention 路径,继续观察不同长段下的稳定性
- 保留 FFN selective F32 物化,继续观察其稳定收益
- 如需继续优化,优先看:
- 更激进的
attention算法级改动,例如按K/Vblock 的 online softmax - 再其次才是
qkv_ms + gate_up_ms + down_proj_ms
- 更激进的
- 如果后续继续深挖,再考虑:
- batched attention 的 block 化,降低大 score matrix 的瞬时内存
decoder_prefill内的投影层进一步收敛
- 保持同一份 2 分钟样本持续回归,避免再次把回归误当成优化