Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Paper • 2608.03796 • Published • 20
Democratizar el PLN en español e incentivar su aplicación para generar impacto social 💛
hf-mem release added a breakdown of Mixture-of-Experts (MoE) memory usage!hf-mem now splits MoE memory into base model weights, routed experts, and KV cache