Imagine a Transformer model without normalization. This is exactly what's proposed in a new paper from Meta, NYU, MIT, and Princeton. The authors found that normalization layers can be replaced with something called Dynamic Tanh (DyT). It looks like this: DyT(x)=γ ∗ tanh(αx)+β.
LLMS
-
LLMs Enable New Spammer Tactics with Synonym Swaps
By
–
Spammers are now using LLMs to run multiple rounds of quasi-synonym swaps in order to avoid getting caught by spam filters due to repetitions. This is a new form of English — spammer-speak, where every sentence is nearly single-use
-

Automating Web Workflows with LLMs at FOSSASIA Bangkok
By
–
Still weakness but all ready for my talk today about automating web workflows using LLMs at @fossasia Bangkok. 🙂 pic.twitter.com/2gYkPHz5RL
— Saurav Jain (Open Source + Communities) (@Sauain) 14 mars 2025Still weakness but all ready for my talk today about automating web workflows using LLMs at @fossasia Bangkok. 🙂
-

ReAct: Synergizing Reasoning and Acting in Language Models
By
–
ReAct: Synergizing Reasoning and Acting in Language Models Yao et al.: https://
arxiv.org/abs/2210.03629 #ArtificialIntelligence #ChatGPT #ReinforcementLearning -
Speculation on Model Distillation and Gemma-Gemini Similarities
By
–
Good catch. Yes, maybe that's distilled down from a larger model they didn't open source. In that context, I am also curious how similar Gemma and Gemini models are.
-
Google Releases Gemma 3 and Gemini 2.0 Flash
By
–
AI News from the past week that I've found interesting. Let me know about (and link me up to) anything that I missed this week so I can be sure to cover it in my breakdown video… (March 7th – 13th)
– Google Releases open-weight Gemma 3 model
– Google releases Gemini 2.0 Flash -
Megaprompt Development Through Human Evaluation and Experimentation
By
–
This one has just been a megaprompt with tons of riffing/experimenting and working with a team of human evaluators on the test cases. It's not the kind of thing that can be accurately evaluated by LLMs or programmatically, so it's been fairly labor intensive to get right. Also,
-
OpenAI o1 and o3-mini Add Python Data Analysis to ChatGPT
By
–
OpenAI o1 and o3-mini now offer Python-powered data analysis in ChatGPT. You can now ask these models to perform tasks like running regressions on test data, visualizing complex business metrics, and conducting scenario-based simulations.
-
Deep Work: Multi-Agent AI Tasks for Pro Users
By
–
Convergence released Deep Work for Pro users. Deep Work operates with multiple AI agents to complete multi-step tasks.
— 🚨 AI News | TestingCatalog (@testingcatalog) 13 mars 2025
This is Operator and Deep Research in one 👀 https://t.co/JTqpnBSkhM pic.twitter.com/8172wbrCI5Convergence released Deep Work for Pro users. Deep Work operates with multiple AI agents to complete multi-step tasks. This is Operator and Deep Research in one https://
t.co/JTqpnBSkhM -

LLM Distillation: Training Efficient Models for Specialized Tasks
By
–
LLM distillation helps train SLMs that retain the reasoning power of large models, while slashing inference costs. Join us to learn: Fine-tuning vs. distillation Training SLMs for specialized tasks Reducing latency & compute costs Register: https://
snorkel.ai/webinar/improv
ing-the-accuracy-of-domain-specific-tasks-with-llm-distillation/
…
