10/ Imitating Reasoning Process of LLMs – develops a 13B model that learns to imitate the reasoning process of large foundational models like GPT-4; it leverages large-scale & diverse imitation data and surpasses Vicuna-13B.
AI
-

Hierarchical Vision Transformer Improves Efficiency Training
By
–
8/ Hierarchical Vision Transformer – pretrains vision transformers with a visual pretext task (MAE), while removing unnecessary components; enables a simple hierarchical vision transformer that’s more accurate and faster at inference and during training.
-
ChatGPT Humor Analysis Reveals Overfitting to 25 Jokes
By
–
9/ Humor in ChatGPT – explores ChatGPT’s capabilities to grasp and reproduce humor; finds that over 90% of 1008 generated jokes were the same 25 jokes and that ChatGPT is also overfitted to a particular joke structure.
-
LEACE Method Erases Gender Bias from BERT Embeddings
By
–
6/ Concept Scrubbing in LLM – presents a method called LEAst-squares Concept Erasure (LEACE) to erase target concept information from every layer in a neural network; it’s used for reducing gender bias in BERT embeddings.
-

Fine-Grained RLHF Improves LM Training and Safety
By
–
7/ Fine-Grained RLHF – trains LMs with fine-grained human feedback; instead of using overall preference, more explicit feedback is provided at the segment level which helps to improve efficacy on long-form question answering and reduces toxicity.
-

Augmenting LLMs with Symbolic Memory via SQL Databases
By
–
5/ Augmenting LLMs with Databases – combines an LLM with a set of SQL databases, enabling a symbolic memory framework; completes tasks via LLM generating SQL instructions that manipulate the DB autonomously.
-

AlphaDev Discovers Faster Sorting Algorithms via Reinforcement Learning
By
–
2/ AlphaDev – a deep reinforcement learning agent which discovers faster sorting algorithms from scratch; the algorithms outperform previously known human benchmarks and have been integrated into the LLVM C++ library.
-

Sparse-Quantized Representation enables 4.75-bit LLM inference
By
–
3/ Sparse-Quantized Representation – a new compressed format and quantization technique that enables near-lossless compression of LLMs across model scales; “allows LLM inference at 4.75 bits with a 15% speedup”.
-
Dense Motion Estimation Method Tracks Pixels Across Full Videos
By
–
1/ Tracking Everything Everywhere All at Once – propose a test-time optimization method for estimating dense and long-range motion; enables accurate, full-length motion estimation of every pixel in a video.https://t.co/7O3Z0Em7wE
— DAIR.AI (@dair_ai) 11 juin 20231/ Tracking Everything Everywhere All at Once – propose a test-time optimization method for estimating dense and long-range motion; enables accurate, full-length motion estimation of every pixel in a video.
-
Top ML Papers Week: AlphaDev, MusicGen, RLHF Advances
By
–
Top ML Papers of the Week (June 5-11): – AlphaDev
– MusicGen
– Fine-Grained RLHF
– Humor in ChatGPT
– Concept Scrubbing in LLM
– Augmenting LLMs with Databases
…
