In this part of the Cache Lab, the mission is simple yet devious: optimize matrix transposition for three specific sizes: 32x32, 64x64, and 61x67. Our primary enemy? Cache misses. Matrix Transposition…
Louis Aeilot's Blog 的其他文章
- How Close Is FlashAttention to the Limit? Understanding Attention Through Data Movement
- Beyond FLOPs: How COSMA Builds Parallel Matrix Multiplication from Communication Bounds
- The Red-Blue Pebble Game: Why Faster Processors Still Have to Move Data
- Git Needs a Trash Can
- RoPE: Properties, Patterns, and Long-Context Behavior
讨论
登录后参与讨论
还没有评论,来说第一句吧。