Blog
Fragments of code, concepts, and curiosity.
The Leap of Sparse Attention in Transformers
A math-first history of sparse attention — the decade-long lineage that modern top-k and masked-softmax work keeps rediscovering. Part I builds the theory from sparsemax and Fenchel–Young losses to α-entmax, the dispersion theorems, and attention as kernel regression, with interactive D3 figures throughout. Part II covers the efficiency era.
SLURM in the Wild: A Practical Guide for Academic Labs
A complete guide from basic concepts to production deployment — covering multi-node setup, GPU scheduling, advanced monitoring, and hard-learned lessons from scaling a research lab from 2 to 30+ users across heterogeneous hardware.