Replies: 2 comments
|
Great question! The reason can_implement uses StrideA{} instead of runtime args.dA boils down to compile-time static polymorphism and zero-overhead dispatch in the CUTLASS 3.x architecture. Here is why it is designed this way: Compile-Time Evaluation & TMA Descriptor Generation: Avoiding Runtime Overhead: Note that while runtime leading dimensions can vary, the layout structure (e.g., whether it's strictly row-major or column-major with standard strides) is captured by the StrideA template parameter, making the static check on StrideA{} both necessary and sufficient for structural capability validation. |
|
Thanks for the explanation. I looked a bit closer at the implementation, and I think there is an important distinction between the compile-time stride type and the runtime stride value here. In
while
So I understand why My remaining question is about cases where If so, is the assumption that the caller is responsible for guaranteeing that That seems to be the key distinction between checking the layout type and checking the actual runtime tensor layout. |
Uh oh!
There was an error while loading. Please reload this page.
cutlass/include/cutlass/gemm/collective/sm90_mma_tma_gmma_ss_warpspecialized.hpp
Line 246 in 8dbce01
why use StrideA{} check but not actual parameters at runtime args.dA to check?
All reactions