u/c-cul

▲ 5 r/CUDA

optimization of SASS stall counts, part 2

by relaxing delays for some small set of instructions we can get speed-up 0.2-0.3% for integer-heavy kernels

redplait.blogspot.com
u/c-cul — 13 days ago
▲ 8 r/CUDA+1 crossposts

optimization of SASS stall counts

- ptxas has enough good heuristic

- in average you can reduce ~3% of stall counts

- overall speed up is not equivalent to the number of optimized stall counts

redplait.blogspot.com
u/c-cul — 29 days ago

memory ssa for numa/gpu

sorry if this is wrong sub-reddit for such questions

Are there some papers/experiments about subj to automatically derive things like indices swizzling/caching in shared memory/pinned memory etc?

reddit.com
u/c-cul — 2 months ago