5/ Teaching Small LMs To Reason – an approach to teach smaller language models to reason; specifically, the LM is thought to use reasoning techniques, such as step-by-step processing, recall-then-generate, recall-reason-generate, extract-generate,…
@dair_ai
-

System 2 Attention: LLM Reasoning for Selective Context Processing
By
–
1/ System 2 Attention – leverages the reasoning capabilities of LLMs to decide what to attend to; it regenerates input context to only include relevant portions before attending to the regenerated context to elicit the final response from the model.
-
Top ML Papers: Mirasol3B, System 2 Attention, Speculative Sampling
By
–
Top ML Papers of the Week (Nov 20 – Nov 26): – Mirasol3B
– System 2 Attention
– Parallel Speculative Sampling
– Advancing Long-Context LLMs
– Teaching Small LMs To Reason
– LLMs as Collaborators for Medical Reasoning
… -
LLM Trading Agent Hides Insider Trading Decisions
By
–
10/ LLMs can Deceive Users – explores the use of an autonomous stock trading agent powered by LLMs; finds that the agent acts upon insider tips and hides the reason behind the trading decision.
-
Learning to Filter Context for RAG Systems
By
–
8/ Learning to Filter Context for RAG – a method that improves quality of the context provided to the generator via two steps: 1) identifying useful context-based and 2) training context filtering models that can filter retrieved contexts at inference.
-

MART: Multi-Round Automatic Red-Teaming for LLM Safety
By
–
9/ MART – proposes an approach for improving LLM safety with multi-round automatic red-teaming; incorporates automatic adversarial prompt writing and safe response generation, which increases red-teaming scalability and the safety of LLMs.
-

Contrastive Chain-of-Thought Prompting Enhances Model Reasoning
By
–
5/ Contrastive CoT Prompting – approach to enhance reasoning by providing both valid and invalid reasoning demonstrations to guide the model to reason step-by-step while reducing reasoning mistakes.
-

Survey on Language Models for Code: 50+ Models Review
By
–
6/ A Survey on Language Models for Code – provides an overview of LLMs for code, including a review of 50+ models, 30+ evaluation tasks, and 500 related works.
-
JARVIS-1: Open-World Multimodal Agent for Minecraft
By
–
7/ JARVIS-1 – an open-world agent that can perceive multimodal input (visual observations and human instructions), generate sophisticated plans, and perform embodied control, within the open-world Minecraft universe.https://t.co/ukgTY6pNT6
— DAIR.AI (@dair_ai) 19 novembre 20237/ JARVIS-1 – an open-world agent that can perceive multimodal input (visual observations and human instructions), generate sophisticated plans, and perform embodied control, within the open-world Minecraft universe.
-

GPT-4 Revolutionizes Scientific Discovery Across Drug Biology Chemistry
By
–
3/ LLMs for Scientific Discovery – explores the impact of large language models, particularly GPT-4, across various scientific fields including drug discovery, biology, and computational chemistry.
