Beyond Distillation: How to Build Smaller, Better Language Models
When a small open-weight model performs surprisingly well, one explanation tends to dominate the conversation: it must have distilled a larger U.S. frontier model. Knowledge distillation—training a smaller “student” model...
Continue reading
0 Comments