For transformers text decoders it's not clear yet afaik – but at least having the encoder head at a higher lr seems to be reliable. I also suspect the first 2 and last 3 layers in the body should have higher lr but I haven't got rigorous tests.
CODE
-
GPT-4 Evolves Reward Function Code for Simulated Robot Control
By
–
GPT-4 for evolving the code of a reward function to control a simulated robot:
-
Weight Decay Regularization and Parameterised Norm Layers
By
–
This was the paper that pointed out wd doesn't regularise in the presence of parameterised norm layers, and effectively just adjusts the LR https://
arxiv.org/abs/1706.05350 -
Weight Decay vs Learning Rate Scheduling in Model Training
By
–
The norm/wd paper showed iirc that you can get exactly the same effect with an lr scheduler. I wonder if we'd get better results by carefully tuning our schedulers instead of using wd? It's something I've wondered about for years but never got around to…
-
Weight Decay Ineffectiveness in Llama and Mistral Models
By
–
Note that Llama and Mistral have parameterised norm layers, so weight decay has no regularizing effect
-
Real-time bug identification practice for continuous quality improvement
By
–
My favorite part of this practice is the last question is “What are we doing wrong?”, followed by a prod similar to “No, really. Something is broken. What is the most broken thing you’ve seen?” Bugs get filed in substantially real time where appropriate. https://
x.com/ZacharyDeWitt/
/ZacharyDeWitt/status/1715424514866311475
… -
Washington’s Software Project Challenges Echo Across Decades
By
–
Interestingly this software echoes a decade later, both in Washington’s learned helplessness about doing software projects and in a “fun” loop I am going through at the moment.
-
Backwards Compatibility: Software Development’s Double-Edged Sword
By
–
backwards compatibility is simultaneously the worst part of a software project (hard to make foundational changes) and the best (means you always have a known-good baseline to compare against)
-

Microsoft PromptFlow: Build Production-Ready LLM Applications
By
–
microsoft/promptflow: Build high-quality LLM apps – from prototyping, testing to production deployment and monitoring. https://
bit.ly/46iENqu #AI #MachineLearning #DeepLearning #LLMs #DataScience -

Building Chatbots with Chat RAG and Memory
By
–
We’re excited to host our next Maker Spotlight live demo session today! Join Cohere community champions Arjun Patel and @Coffee_and_NLP in conversation with AI engineer @an1z8 as he walks us through leveraging Chat + RAG to build chatbot applications with memory and context.