CUDA contribution — T4/SM75 access, background in CUDA kernel work #3505
amankarki151
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi, I'm looking to make a real contribution to CUTLASS and wanted to ask before diving in rather than guess.
A bit of background: I recently implemented and got merged a CUDA kernel (1D pooling, avg/max) into llama.cpp's CUDA backend — verified against 216 test cases, reviewed and merged by the project's lead maintainer (github.com/ggml-org/llama.cpp/pull/27573). I'm looking to go deeper into CUDA kernel work and CUTLASS specifically, since it's a step up in complexity from what I've done so far.
My hardware access is currently 2x NVIDIA T4 (Turing, SM75) via Kaggle Notebooks — I don't have access to Hopper or Blackwell GPUs right now. I know a lot of the recent CUTLASS work targets newer architectures, so I want to ask directly rather than waste anyone's time: is there anything real and useful I could contribute to that's testable on Turing/SM75, or should I be looking at older/simpler parts of the codebase (e.g. CUTLASS 2.x-era GEMM kernels, or non-arch-specific utility/test code) to start?
All reactions