Hey, amazing, awesome people of the beautiful internet ππ₯°
Distillation has been (from my point of view) a main driving factor for the success of hashtag#LLMs - like distilling the knowledge of an amazing big model (say hashtag#DeepSeekv3, or hashtag#GeminiAI) into yours.
Probably, you have done it with minimising a KL divergence, and it somehow worked.
Well, not that well, right?
1οΈβ£ Your model tends to memorise! 2οΈβ£ Your model might get the right answer, but its reasoning might be flawed.
To fix those problems, we rethink distillation and process a new approach! A method that is based on constrained RL that comes with nice theoretical guarantees and excellent performance!
One of the hardest challenges in AI safety is finding the right balance: how do we protect people from harm without undermining their agency? This tension is especially visible in conversational systems, where safeguards can sometimes feel more paternalistic than supportive.
In my latest piece for Hugging Face, I argue that open source and community-driven approaches offer a promising (though not exclusive) way forward.
β¨ Transparency can make safety mechanisms into learning opportunities. β¨ Collaboration with diverse communities makes safeguards more relevant across contexts. β¨ Iteration in the open lets protections evolve rather than freeze into rigid, one-size-fits-all rules.
Of course, this isnβt a silver bullet. Top-down safety measures will still be necessary in some cases. But if we only rely on corporate control, we risk building systems that are safe at the expense of trust and autonomy.