Skip to content
Discussion options

You must be logged in to vote

For the pointer-only ABI, I think your current pattern is already the simplest one: keep cute.Pointer as the kernel argument and reconstruct the Tensor using a static layout inside the kernel.

Since the shape/stride are compile-time constants, they do not need to become PTX parameters. I wouldn't introduce a custom DynamicExpression unless you specifically need a reusable abstraction across many layouts.

For the shared-memory question, current CUTLASS has:

cutlass.utils.get_kernel_smem_size(kernel)

It returns the total static shared-memory allocation for the kernel. The important detail is that it must be called inside a @cute.jit body after the kernel's .launch() has been traced/registered.

Replies: 2 comments

Comment options

You must be logged in to vote
0 replies
Answer selected by peterbjorgensen
Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants