-
The Gemma Challenge and the Case for Agent Collabs
Running an open agent collaboration to speed up Gemma 4.
-
Distilling 100B+ Models 40x Faster with TRL
TRL distillation for 100B+ teachers, 40x faster.
-
Unlocking On-Policy Distillation for Any Model Family
Post hosted on Hugging Face Spaces.
-
Quantifying Prediction Difficulty
How can we tell if an example is difficult for a model?
-
Mean vs. Total Cross-Entropy Loss
Mainly a reminder to myself after wasting a full day debugging gradients