10/ Generative Pretraining in Multimodality – presents a new transformer-based multimodal foundation model to generate images and text in a multimodal context; enables performant multimodal assistants via instruction tuning.
LLMS
-

Training Small Transformers on Arithmetic with Chain-of-Thought
By
–
8/ Teaching Arithmetics to Small Transformers – trains small transformers on chain-of-thought data to improve accuracy and convergence speed and highlights the importance of high-quality instructive data for rapidly eliciting arithmetic capabilities.
-

LLMs as General Pattern Machines: In-Context Learning and Robotics
By
–
6/ LLMs as General Pattern Machines – shows that without any additional training, LLMs are general sequence modelers, driven by in-context learning; applies 0-shot capabilities to robotics and transfer the pattern among words to actions.
-

NaViT: Flexible Vision Transformer for Any Resolution
By
–
5/ Patch n’ Pack: NaViT – a vision transformer for any aspect ratio and resolution through sequence packing; enables flexible model usage, improved training efficiency, and transfers to tasks involving image and video classification among others.
-

Secrets of RLHF in LLMs: PPO Inner Workings Explained
By
–
3/ Secrets of RLHF in LLMs – takes a closer look at RLHF and explores the inner workings of PPO with code included.
-

LongLLaMA Extends Context Length via Contrastive Training
By
–
4/ LongLLaMA – employs a contrastive training process to enhance the structure of the (key, value) space to extend context length; presents a fine-tuned model that lengthens context and demonstrates improvements in long context tasks.
-
Writing Technical Books for LLMs Instead of Humans
By
–
I would still write books. I guess the difference is that I’d write them for LLMs rather than humans, who probably don’t need technical books anymore then
-

MotherDuck: Simplifying Data Scaling and Management
By
–
MotherDuck: The Simple Joys of Scaling Up https://
bit.ly/41TGOqs #AI #MachineLearning #DeepLearning #LLMs #DataScience -
Early Release Performance Benchmarks Expected Similar to Hugging Face
By
–
It's still an early release so we haven't done the benchmarks yet. We expect it'll be similar to the HF implementations.
-

Build Custom LLMs on Your Data with Abacus AI
By
–
You can now use @abacusai to build a custom LLM on your data and solve some of the following problems: • ChatLLM: Search your knowledge base
• DataLLM: Get insights from your data
• TaskLLM: Teach an LLM to solve your tasks
