A simple post while you are easing back in after the holiday break The fastest #LLM #Inference speed is available to try right now at http://
Groq.com running #Llama 2, 70B with 4k sequence length. No "tricks" for our speed, we're an LPU™ based system. Try it out.
AI HARDWARE
-
Groq Offers Fastest LLM Inference Speed with Llama 2 70B
By
–
-
Friend launches GAN project for home GPU testing
By
–
Exciting news from my friend and GAN co-author, I’m looking forward to trying this out on my own GPU at home
-
Oracle Cloud reduces inference latency 51% with NVIDIA Triton
By
–
Explore how @OracleCloud uses NVIDIA Triton Inference Server to deliver its #computervision and #datascience services to enterprises across more than 45 regional data centers reducing #inference latency up to 51% and TCO by 10%.
Read the new blog -
This Curious Robot Should Be Impossible
By
–
This Curious Robot Should Be Impossible! https://
youtu.be/Nnpm-rJfFjQ?si
=TuFGPpS9XyVYTT8_
… via @YouTube -
Llama 2 70B on Groq LPU Delivers Fast Cocktail Recipes NYE
By
–
Side-by-side, #Llama 2, 70B on the Groq LPU™ Inference Engine and the @lmsysorg. If you have to become an instant bartender for your #NYE #party tonight, #prompt for some cocktail & mocktail recipes at https://t.co/4FkcJN9SgW learn on the go. #GenAI #LLMs #groqspeed #inference pic.twitter.com/OUnHeFIYGS
— Groq Inc (@GroqInc) 31 décembre 2023Side-by-side, #Llama 2, 70B on the Groq LPU™ Inference Engine and the @lmsysorg
. If you have to become an instant bartender for your #NYE #party tonight, #prompt for some cocktail & mocktail recipes at http://
Groq.com learn on the go. #GenAI #LLMs #groqspeed #inference -

Survey of 300+ Papers: Generative AI from Gemini to Q-Star
By
–
2/ From Gemini to Q-Star – surveys 300+ papers and summarizes research developments to look at in the space of Generative AI; it covers computational challenges, scalability, real-world implications, and more.
-
NVIDIA’s Advancements in Image Generation in 2023
By
–
8. Advancements in Image Generation • NVIDIA's Perfusion, Drag Your GAN, and Neuralangelo elevate image generation with superior control and realism. #whatsai #ai2023 #GPT4 #futuretech #aiinnovation #ai
-

Fast Inference of Mixture-of-Experts Language Models with Offloading
By
–
Fast Inference of Mixture-of-Experts Language Models with Offloading Artyom Eliseev, Denis Mazur : https://
arxiv.org/abs/2312.17238 #ArtificialIntelligence #DeepLearning #MachineLearning -
tDCS and TMS: Emerging Neurotechnology Applications
By
–
very early but tDCS is pretty interesting overall. another interesting one is TMS (which uses targeted magnets)
-
EEG Sensor Conductance Challenges in Brain Computer Interfaces
By
–
yeah, and hold your breath / don't blink! I saw a mention of muse which was one I've tried a few years ago. form factor is nice but has similar issues. most eeg research uses tight sensor caps with gel / saline because conductance is another issue with dry sensors