MQA and GQA shrink the KV cache, the thing that bottlenecks LLM serving. The signal is knowing they share key/value heads across query heads to cut memory and bandwidth, with GQA as the quality-preserving middle ground. Here is the answer.
Unlock the other 750 answers · ₹2,000 / $25Your progress and mastery stay saved · 6 months · one payment · no auto-renew
