You thought that you can go to sleep now?? v/ @Yampeleg ***********
Orca 2 Just dropped. Paper: https://
arxiv.org/pdf/2311.11045
.pdf
… Results:
Orca 2 13B beats LLaMA-Chat-70B TL;DR:
Training smaller model to reason by using multiple techniques: step-by-step, recall then
LLMS
-

Orca 2 LLM Outperforms Larger Models with Advanced Reasoning
By
–
-
New Fine-Tuning Suite Enables Custom AI Models in 30 Minutes
By
–
We are excited to introduce our comprehensive fine-tuning suite. Businesses can now train and deploy a unique chat, reranking, or multi-label classification model in less than 30 minutes.
-

LoRAX: Open-Source Framework for Efficient Multi-Model LLM Deployment
By
–
We recently open-sourced LoRAX—a novel LLM deployment framework that makes it possible to serve 100s of #finetuned models at the cost of one #GPU. So how does #LoRAX achieve these great results? Check out the blog to find out: https://
pbase.ai/49Q2Fof -

Groq Demonstrates 10X Performance Advantage Over Nvidia GPUs
By
–
"The company’s demo was nothing short of amazing, demonstrating what looked to be at least a 10X performance advantage over (Nvidia) GPUs in performing GPT-3 inference queries." Read more from @karlfreund in @Forbes here: https://
forbes.com/sites/karlfreu
nd/2023/11/20/the-good-bad-and-ugly-from-supercomputing-23-or-nearby/?sh=67a9ec7d575b
… -
Branch-Train-Merge: Modular LLM Training Innovation
By
–
We would like to thank @ssgrn and @margs_li from @UW for presenting their recent research at @SambaNovaAI
. Great to hear about their research ideas – Branch-Train-Merge (BTM) and cBTM. Work around modular LLM training and merging during runtime can unlock massive efficiency! -
Composition of Experts Vision Accelerated by SN40L Research
By
–
These ideas are very aligned with the vision of composition-of-experts that SN40L can help accelerate! Papers covered in the talk – 1. https://
arxiv.org/abs/2303.14177
2. https://
arxiv.org/abs/2208.03306 Thank you for the wonderful talk and looking forward to more of your research! -

RAG Stack Documentation: Organizing Key Strategies
By
–
Deconstructing the RAG stack We repeatedly hear that navigating the idea maze around RAG is a challenge. We've overhauled our RAG docs and made a series of guides to organize / explain key RAG strategies. As an entry point, see our new RAG docs: https://
python.langchain.com/docs/use_cases
/question_answering/
… -
Tuna: No-Code LLM Fine-Tuning Dataset Generation Tool
By
–
Tuna: a no-code tool for quickly generating LLM fine-tuning datasets from scratch built by Student Hacker, @itsandrewgao
-
Claude 2.1 API now available with 200K context window
By
–
Claude 2.1 is available now in our API, and is powering http://
claude.ai for both the free and Pro tiers. Usage of the 200K context window is reserved for Claude Pro users. Read more in our blog post: -
Claude 2.1 Introduces Tool Use for API Integration
By
–
Claude 2.1 includes a new tool use feature that allows the model to integrate with users' existing processes, products, and APIs. This means that Claude can now orchestrate across developer-defined functions or APIs, web search, and private knowledge bases.