DeepSeek API introduces Context Caching on Disk, cutting prices by an order of magnitude | DeepSeek API Docs https://
bit.ly/4dL0sv7
#AI #MachineLearning #DeepLearning #LLMs #DataScience
LLMS
-

DeepSeek API Context Caching Reduces Costs Significantly
By
–
-
Groq CEO: Open Models Win in AI Infrastructure Market
By
–
“Open always wins,” declares Jonathan Ross, CEO of Groq, a provider of specialized AI processing infrastructure that has seen massive uptake of customers using open models. — @mmarshall / @VentureBeat
-

Retrieving the Right Data from Multiple Sources in LLM Applications
By
–
A major challenge in LLM applications is retrieving the right data from multiple sources based on the prompt. Here’s an informative guide on this.
-

SambaNova RDUs deliver 10x GPU speed with tenth power
By
–
🎥 @RodrigoLiang, shares how SambaNova's propriety RDUs can deliver 10x the speed of traditional #GPUs using 1/10th the power.
— SambaNova (@SambaNovaAI) 25 octobre 2024
Experience fast #AI inference on @AIatMeta's Llama 3.2 ⤵️https://t.co/zm6RCXXsaP@RodrigoLiang
, shares how SambaNova's propriety RDUs can deliver 10x the speed of traditional #GPUs using 1/10th the power. Experience fast #AI inference on @AIatMeta
's Llama 3.2 http://
cloud.sambanova.ai -

Aya-Expanse: Impressive Multilingual LLM from Cohere
By
–
Aya-Expanse is a really impressive model if you need a highly performant pre-trained LLM (or speech) in one of the below 32 languages. Congratulation @CohereForAI and @cohere Easy to combine with speech and images as well (just play with the HF Space below) Languages: Arabic,
-
Alibaba Presents Ovis: Novel Multimodal Language Model Architecture
By
–
Alibaba presents Ovis a novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings https://
x.com/i/broadcasts/1
zqKVYbPvRPxB
… -

Combining Bi-Encoders and Cross-Encoders in RAG Systems
By
–
Combining Both in RAG Systems To get the best of both worlds: Retrieve: Use Bi-Encoders as embedding-model to efficiently retrieve top candidate documents. Rerank: Use Cross-Encoders to rerank these candidates for better accuracy. Check this out 7/n
-

Bi-Encoders in RAG: Efficiency and Scalability Benefits
By
–
Bi-Encoders in RAG Bi-encoders are a great choice for embedding models! Why • Efficiency: Pre-compute embeddings for all documents and store them in a vector store.
• Scalability: Then use fast ANN to search over millions of vectors 4/n -
ORPO vs SFT/DPO: Training Technique Comparison Analysis
By
–
I wouldn't necessarily recommend ORPO in practice vs. SFT/DPO, but 3k samples is very small.
-

OpenAI’s Orion Model to Launch November 30th, Training with Synthetic Data
By
–
OpenAI's next flagship model, Orion, slated to be released on ChatGPT's 2 year anniversary on November 30th, sources tell The Verge. It was reported that OpenAI was using o1 to provide synthetic data to train Orion.