-
Notifications
You must be signed in to change notification settings - Fork 1.1k
Pull requests: deepseek-ai/DeepGEMM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Don't enable the assertion check for release build
#390
opened Jul 22, 2026 by
zhangqingshan373
Loading…
Make the JIT cache key independent of the install include path
#388
opened Jul 20, 2026 by
matteso1
Loading…
Optimize sm100 version with token-aware BM240 tiling, native UE8M0, and direct combine
#386
opened Jul 17, 2026 by
qinqinwo
Loading…
Add B200-Calibrated Sparse-Routing Adaptive Wave Sizing
#381
opened Jul 15, 2026 by
qinqinwo
Loading…
SM120: MoE grouped GEMM decode optimizations: skip padding I/O, BLOCK_M=32
#380
opened Jul 15, 2026 by
leavelet
Loading…
SM120: fix FP8 MQA logits swizzle mode for head_dim < 128
#379
opened Jul 15, 2026 by
leavelet
Loading…
Fix non-deterministic results in k_grouped_fp8_gemm_nt_contiguous
#375
opened Jul 9, 2026 by
Functionhx
Loading…
fix: index CUDA sources/headers when generating .pyi stubs
#361
opened Jun 14, 2026 by
Osamaali313
Loading…
Draft: Add green-context split-kernel MegaMoE features
#357
opened Jun 12, 2026 by
RayWang96
Collaborator
Loading…
fix: use importlib to load scripts/generate_pyi.py instead of package…
#355
opened Jun 5, 2026 by
Mikezhang001
Loading…
Add shape-aware recommended alignment for SM90 small-M grouped GEMM
#350
opened Jun 2, 2026 by
qescccczmr
Loading…
Fix TMEM lane address for debug mode across all SM100 kernels
#341
opened May 28, 2026 by
yejunjin
Loading…
Add SM90 FP8 paged MQA logits support for next_n=3Fix sm90 nextn3 paged mqa logits
#340
opened May 28, 2026 by
yangsiqt
Loading…
Fix ue8m0 packing: mask mantissa bits when extracting fp32 exponents
#337
opened May 19, 2026 by
yhyang201
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.