with Griffin and Infini-attention, it increasingly feels like Google leapfrogged Together and RWKV in the race for scaling up linear attention, and they shared a watercooler conversation with Anthropic or something papers are trickling out now but it's hard to know what's
LLMS
-
Emerging Consensus: 8x22B Model is Mixtral Medium
By
–
think emerging consensus is 8x22B -is- mixtral medium
-

Claude Opus Maestro: Search Integration for Enhanced Subagent Performance
By
–
New update in Maestro!
— Pietro Schirano (@skirano) 11 avril 2024
Introducing Search 🔍
Now, when creating a task for its subagent, Claude Opus will perform a search and get the best answer to help the subagent solve that task even better.
See it in action.
It's honestly so cool to see its thinking. 💭 pic.twitter.com/rJ3Y493uFBNew update in Maestro! Introducing Search Now, when creating a task for its subagent, Claude Opus will perform a search and get the best answer to help the subagent solve that task even better. See it in action. It's honestly so cool to see its thinking.
-
Snowflake Partnership Launches Specialist AI Models
By
–
Super excited by this partnership with @SnowflakeDB
, with specialist models! -
Bixtral deployment errors on A100 GPUs HuggingFace endpoints
By
–
Hey @michelleyhbn @LucSGeorges just wanted to share that there seem to be errors when using 4x A100 80gbs to host Bixtral on HF endpoints… seems like more than enough RAM, but still crashed — tried gptq etc. No dice, all failed. Not a huge deal for me, but wanted to make sure
-
DeepSpeed optimization successfully implemented for AI training
By
–
@winglian wrangled deepspeed till it worked
-

GPT Capability Leaps: Research Analysis from GPT-3 to GPT-4
By
–
I am not making a random statement, I am citing this paper: https://
arxiv.org/pdf/2403.05812
.pdf
… So it might be better to discuss the research here rather than having subjective conversation. But leaps on many measures from GPT-3 to 3.5 to GPT-4 were more than a doubling of capabilities. -
Fine-tuned Mixtral 8x22B Inference Optimization Guide
By
–
Fine-tuned Mixtral 8x22B is ready. Curious what the easiest way to do inference is? Spinning up some GPUs for vLLM right now, is there something better?
-

ChatGPT vs. Copilot: Which AI Chatbot is Better?
By
–
#ChatGPT vs. Copilot: Which #AI chatbot is better for you?
by @sabrinaa_ortiz @ZDNET Read more: https://
buff.ly/3SFZnfe #Chatbots #ArtificialIntelligence #MI #DL #Tech #Technology #Data cc: @terenceleungsf @yvesmulkers @pascal_bornet -

Prompting Mixtral-8x22B: Effective Techniques Thread
By
–
Man I forgot how much more intensive prompting base models is All of this, just to shorten some text. Would anyone find value in a quick thread on how to prompt Mixtral-8x22B effectively?
