FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning paper: https://
tridao.me/publications/f
lash2/flash2.pdf
…
github: https://
github.com/Dao-AILab/flas
h-attention
… Scaling Transformers to longer sequence lengths has been a major problem in the last several years, promising to improve performance
AI HARDWARE
-

FlashAttention-2: Faster Attention with Better Parallelism
By
–
-
Token Output Rate Matters for LLM Inference Processor Evaluation
By
–
When evaluating inference processors for deploying autoregressive #LLMs, it is crucial to consider the rate of tokens output per second, not just the rate of tokens input and processed per second. For more info: https://
groq.com/inference/ -

Falcon-40B Achieves 5X Faster Inference Than GPT-3
By
–
#Falcon-40B can perform inference 5X faster than #GPT-3 & requires less compute for pre-training compared to comparable models like #PaLM-62B or #Chinchilla.
V/ @cwolferesearch @FrRonconi @labordeolivier @Nicochan33 @Shi4Tech @KanezaDiane @enilev @kalydeoo @Khulood_Almani -

Groq and BittWare Design Efficient AI Inference Chip for Low Latency
By
–
For #MachineLearning #inference, GPU inefficiencies lead to latency, low silicon resource usage & unpredictable performance. @GroqInc & @BittWareInc designed an #AI deep learning chip to provide predictable, efficient, low-latency inference. For more info: http://
bittware.com/products/groq -
GPU Support Request for PyTorch Integration
By
–
I hope it will support GPUs some time so we can use it with PyTorch etc. (Thanks for the link, but sry, as a personal thing, I am not watching content from that interviewer due to his blocking spree of AI and deep learning researchers and content creators)
-
Autonomous Robots Recharge Without Human Assistance Using ML
By
–
It can open its doors and find the nearest electric outlet to recharge without human assistance. Read more about this interesting development here: https://
bit.ly/46GSjFH @uofcincy @_DigitalIndia @GoI_MeitY @NeGD_GoI @nasscom #Robots #Robotics #machinelearning #INDIAai -
Which AI startups face GPU access constraints today?
By
–
Outside of “foundation model” AI startups, which AI companies are actually rate-limited by lack of sufficient GPU access today?
-
Unitree Robots Lead in Price and Robustness
By
–
I +1 this prediction. Unitree is in a league of its own with its robots in terms of price and robustness
-

Groq Hiring Senior Product Manager for Generative AI
By
–
Groq is #hiring a #SeniorProductManager! Join us and shape the future of #GenerativeAI. https://
groq.com/careers/?gh_ji
d=5662281003
… -

Groq Enhances LLM Inference Strategy for Federal AI ROI
By
–
Following a sound #inference strategy will be the difference between success and failure when it comes to deploying #LLM workloads. We're thrilled to have Marc Wilson on our team to help Federal customers achieve a generational leap in the ROI of #AI solutions. pic.twitter.com/ZjRShiNMwo
— Groq Inc (@GroqInc) 12 juillet 2023Following a sound #inference strategy will be the difference between success and failure when it comes to deploying #LLM workloads. We're thrilled to have Marc Wilson on our team to help Federal customers achieve a generational leap in the ROI of #AI solutions.