How to make Tensor static and only call kernel by pointer? #3169
|
I am trying to turn Python cute dsl kernels into TensorRT QDP Plugins. It restricts the kind of parameters the PTX can have to (input_pointers..., scalar_arguments..., output_pointers...). I have it working for a simple "add_one_kernel" (see the program listing below if interested). Instead of writing the kernel with pointer arguments, I would like to write it with some StaticTensorPointer class that automatically turns the layout etc. into compile-time variables and only the pointer ends up as a PTX parameter. I guess it's possible with a custom Another question: Is it possible to get the shared memory needed for a compiled kernel? I need it for the kernel launch parameters of the kernel in the plugin. |
Replies: 2 comments
|
For the pointer-only ABI, I think your current pattern is already the simplest one: keep Since the shape/stride are compile-time constants, they do not need to become PTX parameters. I wouldn't introduce a custom DynamicExpression unless you specifically need a reusable abstraction across many layouts. For the shared-memory question, current CUTLASS has: cutlass.utils.get_kernel_smem_size(kernel) It returns the total static shared-memory allocation for the kernel. The important detail is that it must be called inside a
|
|
The key distinction seems to be between making the tensor layout static at the CuTe DSL level and keeping the underlying data pointer dynamic at the kernel ABI level. If the goal is to generate a TensorRT QDP Plugin with only pointers and scalar arguments in the PTX parameter list, I would expect the tensor's shape/layout metadata to be represented as compile-time expressions while the base address remains a runtime parameter. For the shared-memory question, I think the important part is whether the compiled kernel's shared-memory requirement is exposed through the generated kernel metadata rather than being derived from the runtime tensor object. If the allocation size depends only on static layout/tile parameters, it should be possible to determine it during compilation. Would the recommended approach be to define a custom Also, for TensorRT integration, is there a supported way to query the generated kernel's required dynamic shared-memory size after compilation, rather than duplicating the shared-memory calculation in the plugin implementation? |
For the pointer-only ABI, I think your current pattern is already the simplest one: keep
cute.Pointeras the kernel argument and reconstruct the Tensor using a static layout inside the kernel.Since the shape/stride are compile-time constants, they do not need to become PTX parameters. I wouldn't introduce a custom DynamicExpression unless you specifically need a reusable abstraction across many layouts.
For the shared-memory question, current CUTLASS has:
cutlass.utils.get_kernel_smem_size(kernel)
It returns the total static shared-memory allocation for the kernel. The important detail is that it must be called inside a
@cute.jitbody after the kernel's.launch()has been traced/registered.