Blog

Fragments of code, concepts, and curiosity.

· 2-part series · interactive

The Leap of Sparse Attention in Transformers

A math-first history of sparse attention — the decade-long lineage that modern top-k and masked-softmax work keeps rediscovering. Part I builds the theory from sparsemax and Fenchel–Young losses to α-entmax, the dispersion theorems, and attention as kernel regression, with interactive D3 figures throughout. Part II covers the efficiency era.

Sparse Attention Transformers α-entmax
· 50 min read

SLURM in the Wild: A Practical Guide for Academic Labs

A complete guide from basic concepts to production deployment — covering multi-node setup, GPU scheduling, advanced monitoring, and hard-learned lessons from scaling a research lab from 2 to 30+ users across heterogeneous hardware.

Infrastructure GPU Computing SLURM
More posts coming soon.