Want to scale LLMs without skyrocketing compute costs? Samsung presents MeKi: Memory-based Expert Knowledge Injection. Instead of making models bigger to learn everything, MeKi gives LLMs a dynamic memory bank of expert knowledge. Think of it as a cheat sheet the model can
AI
-
Why Geoffrey Hinton called it Dark Knowledge in CIFAR meetings
By
–
I’m wondering why @geoffreyhinton called it Dark Knowledge in the earlier @CIFAR_News meetings??
-
Codex working on plugin, could ship as official
By
–
codex is on it! if this works, we can ship this as an official plugin 🙂
-

OrderGrad: optimization beyond the mean via order statistics
By
–
« OrderGrad: Optimization Beyond the Mean with Policy Gradient Estimation by Order Statistics » Most RL optimizes the average reward, but deployment often cares about the best sample, the worst tail, the median, CVaR, or
-

OPRD: On-Policy Distillation of Teacher Representations Before LM Head
By
–
"OPRD: On-Policy Representation Distillation" On-policy distillation usually matches teacher and student only at the token probability level, throwing away the teacher’s hidden states. This paper moves the loss before the LM head, aligning student and teacher representations on
-
Pedro’s earlier insight on Gato and universal AI imitation
By
–
Thanks for asking. Pedro pointed out issue much earlier when I was working on General AgenT One — Gato https://
arxiv.org/abs/2205.06175 and wrote about it https://
arxiv.org/abs/2110.10819 Then Pedro came up with this brilliant theoretical insight: https://
adaptiveagents.org/universal_ai_a
s_imitation
… And we -
Getting Smart With Virtual Assistants in Healthcare
By
–
Getting Smart With Virtual Assistants in Healthcare
#AI #AIio #AIInnovation #ML #DataScience #Futureofwork @Scobleizer @AndrewYNg @drfeifei @KirkDBorne @fchollet @rowancheung @antgrasso -
AI crosses the threshold of recursive self-improvement
By
–
AI has just crossed a threshold: recursive self-improvement. Models write the code that trains them, discover the algorithms that will make them better, design the chips that run them… The machine improves the machine, which will improve the machine.
-
Why code neural net from scratch for tabular ML when XGBoost exists?
By
–
Why are you, a grown man, still coding a neural net from scratch for tabular ML when XGBoost exists? @trainxgb
-
Request for precise explanation of causal training of world models
By
–
Could you please be more precise. Could you show us precisely how world models are trained causally. Thanks
