A sparse matrix is mostly zeros. In a neural network, we can create one by setting some learned weights to zero, a process called pruning. That can make a model smaller, but speeding it up depends on which calculations the GPU can actually skip.
GPUs work efficiently on regular blocks of numbers. NVIDIA's block-sparse approach skips blocks that are entirely zero, then does all the calculations inside the blocks it keeps. Any zeros left inside those blocks still get processed.
If we're already doing those calculations, could we let those weights learn again and improve the model without slowing it down?
Hugging Face's hybrid-filled pruning lets zeroed weights learn again within the attention heads it keeps. It reports recovering some accuracy at the same speed.
I'd test this while keeping the same blocks, and measure both accuracy and runtime. If the extra weights improve accuracy without slowing inference, keeping them seems worthwhile. The number of zeros alone doesn't tell us how efficiently a model runs.