6/ GPQA – a graduate-level Google-proof QA benchmark consisting of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry; the strongest GPT-4 based baseline achieves 39% accuracy.
LLMS
-

Teaching Small Language Models Reasoning Techniques
By
–
5/ Teaching Small LMs To Reason – an approach to teach smaller language models to reason; specifically, the LM is thought to use reasoning techniques, such as step-by-step processing, recall-then-generate, recall-reason-generate, extract-generate,…
-

System 2 Attention: LLM Reasoning for Selective Context Processing
By
–
1/ System 2 Attention – leverages the reasoning capabilities of LLMs to decide what to attend to; it regenerates input context to only include relevant portions before attending to the regenerated context to elicit the final response from the model.
-
Top ML Papers: Mirasol3B, System 2 Attention, Speculative Sampling
By
–
Top ML Papers of the Week (Nov 20 – Nov 26): – Mirasol3B
– System 2 Attention
– Parallel Speculative Sampling
– Advancing Long-Context LLMs
– Teaching Small LMs To Reason
– LLMs as Collaborators for Medical Reasoning
… -
LLM limitations in rare cases and medical applications
By
–
We can hope but 3 cautions
rare circumstances have not been the strong point of LLMs, when studied under controlled conditions some of the work that @binajv points to is really about medical databases, and getting high quality data, rather than LLMs per se, and I hope that -

Build RAG with Mistral-7B and LangChain Locally
By
–
Build your own RAG with Mistral-7B and LangChain Using local models can enable chatbots over your data while ensuring your data never leaves your computer Great for enterprise or privacy centric use cases This blog post by Madhav Thaker is one of the best resources we've
-
Tool Use Methodology: Retrieval Approach in AI Systems
By
–
There’s a paper describing this approach, for one particular tool (retrieval); idea is the same however for other tools https://
arxiv.org/abs/2310.11511 -

GAIA: Benchmark for Evaluating General AI Assistants
By
–
GAIA: a benchmark for General AI Assistants Mialon et al.: https://
arxiv.org/abs/2311.12983 #ArtificialIntelligence #DeepLearning #MachineLearning -

Master ChatGPT Prompt Crafting: AI Dialogue Art Guide
By
–
Dive into the art of #ChatGPT Prompt Crafting with this insightful infographic by GPT AI! Master the dialogue with AI by understanding roles, tasks, formats, and more. Elevate your tech game! Follow @ingliguori for a journey through #DigitalTransformation and #AI insights.
-

GAIA: Benchmark for General AI Assistants Performance
By
–
GAIA: a benchmark for General AI Assistants Mialon et al.: https://
arxiv.org/abs/2311.12983 #ArtificialIntelligence #DeepLearning #MachineLearning
