Nobody's been talking about it but it's rather *mind-blowing* imo that the open-source Flacon 40B model is topping LLaMa 65B on leaderboards and many evals while having required not even half the compute of LLaMa to train from scratch Quick back of the envelop calculations:
–
LLMS
-
Falcon 40B Outperforms LLaMa 65B With Half Training Compute
By
–
-

Brainformers: Trading Simplicity for Efficiency in Transformers
By
–
Brainformers: Trading Simplicity for Efficiency paper page: https://
huggingface.co/papers/2306.00
008
… develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms -

Understanding Transformer Internal Mechanisms Through Memory Analysis
By
–
Birth of a Transformer: A Memory Viewpoint paper page: https://
huggingface.co/papers/2306.00
802
… Large language models based on transformers have achieved great empirical successes. However, as they are deployed more widely, there is a growing need to better understand their internal mechanisms -

Analyzing Attention Glitches in Transformer Language Models
By
–
Exposing Attention Glitches with Flip-Flop Language Modeling abs: https://
arxiv.org/abs/2306.00946 identifies and analyzes the phenomenon of attention glitches, in which the Transformer architecture's inductive biases intermittently fail to capture robust reasoning. To isolate the -
AI Expectations Paradox: Short-term Hype, Long-term Underestimation
By
–
the paradox of ai expectations: expectations of what AI will do in the next 1-2 years are always too high but expectations of AI will do in 5+ years are usually too low 5 years ago we were in self-driving summer. we didn’t get L5 autonomy, but we got the FAR more powerful LLMs
-

ReviewerGPT: Using Large Language Models for Scientific Paper Review
By
–
ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing Given the rapid ascent of large language models (LLMs), we study the question: (How) can large language models help in reviewing of scientific papers or proposals? We first conduct some pilot
-
Karpathy’s GPT State Talk: Training and Prompting Strategies
By
–
Excellent talk by @karpathy on the State of GPT, breaking down everything from the training pipeline to the most effective prompting strategies for LLMs. https://
youtube.com/watch?v=bZQun8
Y4L2A
… -
The Evolution of Prompt Engineering Beyond Simple Model Nudging
By
–
People always ask if prompt engineering is going to go away over time. My short answer is "no". But, a more nuanced answer is that the goal of prompt engineering has evolved over time: from nudging a finnicky language model to do an "easy" task (2020/2021) to figuring out how to
-

LangChain CEO to Keynote Data AI Summit 2026
By
–
This year, we’re pumped to have @langchain CEO @hwchase17 as a #DataAISummit keynote speaker! LangChain recently integrated with our #LLM serving endpoint, making easy to connect to open source models like Dolly Register today to hear him speak https://
bit.ly/40bHaZr -

REVEAL: Retrieval-Augmented Visual-Language Model for Knowledge-Intensive Tasks
By
–
Learn how REVEAL, an end-to-end retrieval-augmented visual-language model that learns to use multi-source multi-modal data to answer knowledge-intensive queries, achieves state-of-the-art results on visual question answering and image caption tasks. https://
goo.gle/3qcZwwc