Top ML Papers of the Week (Mar 13 – Mar 19): – GPT-4
– FlexGen
– NeRFMeshing – Resurrecting RNNs
– An Overview of Language Models
– Universal Prompt Retrieval for LLMs
…
@dair_ai
-
Top ML Papers of the Week: GPT-4 and Language Models
By
–
-
GPT-4 Large Multimodal Model Advances General Knowledge
By
–
1/ GPT-4 – a large multimodal model with broader general knowledge and problem-solving abilities. https://t.co/drO0DXfZK9
— DAIR.AI (@dair_ai) 19 mars 20231/ GPT-4 – a large multimodal model with broader general knowledge and problem-solving abilities.
-
LERF: Language Embeddings in 3D Radiance Fields
By
–
2/ LERF (Language Embedded Radiance Fields) – a method for grounding language embeddings from models like CLIP into NeRF; this enables open-ended language queries in 3D.https://t.co/0zbkWzzTUT
— DAIR.AI (@dair_ai) 19 mars 20232/ LERF (Language Embedded Radiance Fields) – a method for grounding language embeddings from models like CLIP into NeRF; this enables open-ended language queries in 3D.
-

GigaGAN: Scaling GANs for Fast High-Resolution Image Synthesis
By
–
10/ GigaGAN – enables scaling up GANs on large datasets for text-to-image synthesis; it’s found to be orders of magnitude faster at inference time, synthesizes high-resolution images, & supports various latent space editing applications.
-

OpenICL: Open-source toolkit for in-context learning and LLM evaluation
By
–
8/ OpenICL – a new open-source toolkit for in-context learning and LLM evaluation; supports various state-of-the-art retrieval and inference methods, tasks, and zero-/few-shot evaluation of LLMs.
-

MathPrompter Improves LLM Mathematical Reasoning Performance
By
–
9/ MathPrompter – a technique that improves LLM performance on mathematical reasoning problems; it uses zero-shot chain-of-thought prompting and verification to ensure generated answers are accurate.
-

Foundation Models for Decision Making: Tools and Methods
By
–
6/ Foundation Models for Decision Making – provides an overview of foundation models for decision making, including tools, methods, and new research directions.
-

Hyena Hierarchy: Subquadratic Attention Replacement for LLMs
By
–
7/ Hyena Hierarchy – a subquadratic drop-in replacement for attention; it interleaves implicit long convolutions and data-controlled gating and can learn on sequences 10x longer and up to 100x faster than optimized attention.
-

A History of Generative AI: From GAN to ChatGPT
By
–
4/ A History of Generative AI – an overview of generative AI – from GAN to ChatGPT.
-

LLMs Override Semantic Priors Through In-Context Learning at Scale
By
–
5/ LLMs do In-Context Learning Differently – shows that with scale, LLMs can override semantic priors when presented with enough flipped labels; these models can also perform well when replacing targets with semantically-unrelated targets.