Skip to content

fix(cuda): handle 64-lane warps in broadcast backward reduction - #216

Open
chen2021673 wants to merge 1 commit into
masterfrom
fix/cuda-64-lane-warp-reduction
Open

fix(cuda): handle 64-lane warps in broadcast backward reduction#216
chen2021673 wants to merge 1 commit into
masterfrom
fix/cuda-64-lane-warp-reduction

Conversation

@chen2021673

@chen2021673 chen2021673 commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

背景

BroadCast 反向传播 kernel 之前默认物理 WarpSize 为 32,在 32-lane 设备上运行正常,但在部分 WarpSize 为 64 的设备上会出现规约错误。

修改内容

  • 将物理 Warp 划分为独立的 32-lane 逻辑 Warp
  • 保留完整的 ballot mask,避免 64-lane mask 被截断
  • 统一 shuffle、CUB 规约和共享内存的逻辑 WarpSize
  • 新增跨多个逻辑 Warp 的广播乘法反向传播测试

测试

新增 {2, 64} * {2, 1} Broadcast Mul Backward 用例,验证:

  • grad_a 均为 2
  • grad_b{64, 64}

maca C550 WarpSize=64 环境中测试:
image

CUDA 测试:
image
image

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant