ModernBERT is not a chat model.
@jeremyphoward
-
Developer tools and creator tooling: do devs count as creators?
By
–
Do devs count as creators? Because there’s a lot of dev tools companies that have done well. Actually I guess don’t really know what you mean by creator tooling, since I would have thought stuff like figma would count as tooling for creators?
-

Fine-tuning as continued pretraining improves medical AI
By
–
Your daily reminder that fine tuning is just continued pretraining. Super cool results from @antoine_chaffin who is putting this knowledge into practice to improve medical AI:
-
Claude’s summarization approach in code implementation
By
–
Of course – we use it a lot! Although afaict Claude code uses a basic summarization approach.
-
Fast.ai: A Powerful Foundation for Dedicated AI Learners
By
–
http://
fast.ai can be a pretty amazing foundation for folks willing to put in the time and effort to learn, like @Suhail
, who worked really really hard. -
Prefix Tuning Approaches: Layer-wise Activation Prepending Comparison
By
–
My understanding is that @percyliang
's prefix tuning approach prepended activations to *every* layer, but yours only prepends to the first layer — is that right? -
Prefix Tuning Initialization with Real Tokens and DoRA
By
–
Is the trick of initializing with real tokens the secret to making all prefix tuning approaches competitive? BTW @EyubogluSabri I don't see DoRA mentioned in your paper – I wonder if that would close the gap a bit?
-
LoRA vs DoRA vs Prefix Tuning: Fine-tuning Method Effectiveness
By
–
The big open question for me is why LoRA (and particularly DoRA) has been the most successful for fine-tuning, but it worked so badly for this use case. Or maybe the opposite – why did prefix tuning work so well?
-
Meta’s AI Talent Loss: Strategic Consequences of Leadership Decisions
By
–
If Zuck hadn't laid off Erik's team of exceptional AI talent a few years ago, they would have less of an AI talent problem today…
