
Are Megakernels Dead? Mixture of Kittens Says No
This episode explores why giant fused GPU kernels have become a maintenance and scaling headache for production inference, especially when tensor parallelism forces real communication boundaries. It then pivots to Mixture of Kittens, a new MoE training megakernel that delivered a reported 41% throughput boost and shows why low-level optimization still matters at frontier-lab scale.
Chapter 1
The Death of the 67000 Line Kernel
William Palmer
Sixty seven thousand lines of hand written C plus plus code. Just, just think about that for a second. That is what some research teams were putting into a single GPU kernel, a megakernel, trying to fuse an entire forward pass of a large language model together. All to save a few microseconds of launch overhead. But... well, last week on the Latent Space podcast, top inference engineers basically declared the whole approach dead on arrival for production. And, you know, as someone who spends weekends tearing down race car engines, I, I kind of get the obsession. When you are chasing pure speed, the urge to custom fabricate every single bracket and intake runner yourself is massive. But when your bespoke intake snaps at two hundred miles an hour because you bypassed the factory telemetry? Yeah, you realize why modular engineering exists.
William Palmer
So why did megakernels hit such a hard wall? Well, it turns out you run straight into a hard mathematical constraint called tensor parallelism. Imagine you are splitting a huge matrix across two GPUs. Half the matrix lives on GPU one, half lives on GPU two. Everything is fine until you hit a non linear operation, like attention softmax or exponentiation. To calculate the softmax for a single row, you cannot just look at your half of the row. You need the partial results from the other GPU to compute the full row denominator. So even if you write this ridiculously complex fused kernel, your thread blocks are forced to stop, wait, and communicate across the interconnect. A fused kernel simply cannot save you from the laws of distributed math.
William Palmer
And, and then there is the maintainability nightmare. Engineers who actually worked on these mega kernels at top labs admitted that in production, modular setups like TensorRT LLM almost always win out. Why? Because you can optimize each individual component independently, run them in parallel, and keep the code clean. Plus, NVIDIA is actively killing the main reason megakernels existed in the first place. Kyle Kranen over at NVIDIA recently pulled back the curtain on their upcoming Rubin GPU architecture. They are adding physical CTA dependency triggers right into silicon. So when kernel one finishes seven of its ten thread blocks, kernel two does not sit there idling waiting for the stragglers. The hardware itself launches those seven available blocks immediately. Silicon is eating software optimization for breakfast. So, case closed, right? Fused mega kernels are buried. Dead and buried.
Chapter 2
Mixture of Kittens and the 41 Percent Counter Attack
William Palmer
Well... not so fast. On the exact same day that engineers were writing eulogies for megakernels, a Stanford PhD student named Stuart Sul and his team released something that sent shockwaves through inference engineering. Building on Dan Fu's research group and Ben Spector's earlier work on ThunderKittens, Cursor and Hazy Research open sourced a brand new training megakernel called Mixture of Kittens. MoK for short. And the headline claim? A forty one percent overall boost in tokens per second on NVIDIA NVL72 hardware clusters. That is up to two point three seven times faster than the strongest public baselines!
William Palmer
Now, how on earth did they pull that off if megakernels are supposed to be dead? They did it by targeting the absolute most brutal bottleneck in Mixture of Experts models: communication. In a massive MoE architecture, token routing and expert compute constantly fight for GPU bandwidth. Mixture of Kittens strips out high level abstractions entirely. It uses pull based dispatch built on ThunderKittens tile primitives, fusing the MoE routing communication and compute directly into a single, deterministic execution loop. It completely eliminates CPU to GPU sync overhead.
William Palmer
So, wait, how do we reconcile these two realities? On one hand, megakernels are a maintenance trap that commercial serving engines avoid. On the other hand, Mixture of Kittens delivers a forty one percent speedup on hardware racks costing millions of dollars. The answer comes down to economics. When you operate at the scale of frontier labs or hyperscalers, a forty one percent efficiency jump translates directly into tens or even hundreds of millions of dollars saved in hardware and electricity. For that level of return, companies will happily hire world class CUDA wizards to maintain fifty thousand lines of terrifying low level assembly. So are megakernels dead? For standard modular deployments, absolutely. But for ultra high stakes training at cluster scale? Megakernels are so, so back.