A great overview of what's going on under the hood in a RAG pipeline Thanks for writing this @czue
!
LLMS
-
RAG Pipeline Technical Overview and Implementation Guide
By
–
-

Groq Day: Ultra-Fast LLM Performance and Hardware Ecosystem Demo
By
–
Join us at #GroqDay to see ultra-fast #LLM performance for #AI applications at scale on our live demo; learn about the synchronous Groq hardware and software ecosystem and roadmap; & hear from our leaders directly. Register at http://
groq.link/groqday5 -
Clarification between inference and training processes
By
–
Thanks for tagging me, but that’s only inference not training right?
-
Code Llama Now Available in Hugging Face Playground
By
–
You can try Code Llama now in the Code Llama playground @huggingface space — it's also available in the Hugging Face ecosystem, starting with transformers version 4.33.
-

Groq Achieves 100 Tokens/Second Milestone with Llama2 70B
By
–
@GroqInc recently became the world's first to achieve the 100 t/s/u responsiveness milestone for text generation on @MetaAI
's #Llama2 70B model. Before the celebratory swag made it to Toronto, the team already blew past their record. Stay tuned for news about our latest milestone -

The 5 Best ChatGPT Prompts for Entrepreneurs
By
–
Starting with no plan on ChatGPT can feel overwhelming. Entrepreneurs harness ChatGPT for ideation, drafting emails, business coaching, and SEO insights. Here's top 5 Prompts Used by Entrepreneurs Follow the thread below:
-
Alignment Engineering: Bridging LLM and Human World Models
By
–
"Alignment engineering" might be a better term than "prompt engineering". It's not just about instructing LLM but about aligning your world model with LLM's world model. If LLM's response strays from your expectations, it means your world models differ, and the prompt needs
-

Perplexity Adopts Fine-Tuned GPT-3.5 for Cost Efficiency
By
–
The transition to the fine-tuned GPT-3.5 model also reduces inference costs. This efficiency enables us to continue investing in enhancements, making your Perplexity experience even better. Try Perplexity Copilot now and see the difference for yourself! https://
pplx.ai/yVdg4Av -
FT-GPT-3.5 Latency Reduced 4-5x Faster Inference Speed
By
–
We’ve reduced the model latency by 4-5x, serving results on average in 0.65 seconds instead of 3.15 seconds (FT-GPT-3.5 compared to GPT-4). You may notice the speedup when Copilot prompts you for user input. Every second counts, and we’re here to make them all productive. pic.twitter.com/dGTu8aYtXw
— Perplexity (@perplexity_ai) 25 août 2023We’ve reduced the model latency by 4-5x, serving results on average in 0.65 seconds instead of 3.15 seconds (FT-GPT-3.5 compared to GPT-4). You may notice the speedup when Copilot prompts you for user input. Every second counts, and we’re here to make them all productive.
-

Fine-tuned GPT-3.5 matches GPT-4 performance accuracy
By
–
Our fine-tuned GPT-3.5 model ties with the GPT-4-based model in human ranking on our task. This isn’t just about speed; it’s about delivering precise and accurate responses to your complex queries. Now you get the best of Perplexity at your finger tips.