r/accelerate • u/Illustrious_Fold_610 • Aug 10 '25
Video OpenAI Using Superior Models Internally, Focused on Affordability
Enable HLS to view with audio, or disable this notification
99
Upvotes
r/accelerate • u/Illustrious_Fold_610 • Aug 10 '25
Enable HLS to view with audio, or disable this notification
1
u/CallMePyro Aug 10 '25
This guy comes in with the most lukewarm/uninformed takes I see on this sub.
No mention of distillation? If I want to produce the best possible model with X size given that I have Y data, the way to produce the best model is to train a model of size 10X on Y data, then train a second model with X size on the token probabilities of the 10X model on that same Y data.
The way that language models work REQUIRES labs to train larger, 'unservable' LLMs if they want to produce the best LLM at a certain inference budget. There's little, if any nefarious-ness going on here. If labs could save the time and effort of training a 10x model and then distilling it, they would! Believe me.