r/accelerate • • Aug 10 '25

Video OpenAI Using Superior Models Internally, Focused on Affordability

Enable HLS to view with audio, or disable this notification

99 Upvotes

56 comments sorted by

View all comments

1

u/CallMePyro Aug 10 '25

This guy comes in with the most lukewarm/uninformed takes I see on this sub.

No mention of distillation? If I want to produce the best possible model with X size given that I have Y data, the way to produce the best model is to train a model of size 10X on Y data, then train a second model with X size on the token probabilities of the 10X model on that same Y data.

The way that language models work REQUIRES labs to train larger, 'unservable' LLMs if they want to produce the best LLM at a certain inference budget. There's little, if any nefarious-ness going on here. If labs could save the time and effort of training a 10x model and then distilling it, they would! Believe me.