We trained unstructured sparse 1.3B GPT-3 models on CS-2 systems and demonstrated how we achieve competitive results at a fraction of the inference FLOPs – our 83.8% sparse model achieved a 3x reduction in FLOPs at matching performance Learn more here:
LLMS
-
LangChain 0.0.25: Bug Fixes and Feature Improvements
By
–
LangChain version 0.0.25 Bug fix for the semantic similarity example selector (@AkashSamant4)
Support SQL statements that return no results (@andrewgleave – 1st PR! Welcome!!)
Tidying up SerpAPI and PALChain
Adding examples of demos that use LangChain -

BLOOM: 176B Parameter Open-Access Multilingual Language Model
By
–
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Le Scao et al.: https://
arxiv.org/abs/2211.05100 #ArtificialIntelligence #DeepLearning #MachineLearning -
Text-davinci-002 derived from code-davinci-002 with code training
By
–
News to me that text-davinci-002 derives from code-davinci-002. I imagined the code model had code training the text one lacked. It looks like they both have code training, but instruction tuning interferes with code completion use cases maybe.
-

OpenAI retroactively updates training methods for GPT-3.5 models
By
–

OpenAI update on training methods for “GPT‑3.5” models (retroactively including text‑davinci‑002): https://
beta.openai.com/docs/model-ind
ex-for-researchers
… -
Whisper Fine Tuning Events with OpenAI and Meta talks
By
–
Yess! Join @sanchitgandhi99 as he unpacks the Whisper Fine Tuning Events this Friday! Followed by talks from @OpenAI and @MetaAI on Monday!!
-
GPT-4 Adapts Users Instead of Being Fine-Tuned
By
–
GPT-4 doesn’t need to be fine-tuned for downstream tasks. It fine-tunes the humans that intend on leveraging it.
-
Sparse Pre-training Dense Fine-tuning Reduces GPT Training FLOPs
By
–
At @NeurIPSConf
, we introduced Sparse Pre-training and Dense Fine-tuning to reduce the computational FLOPs of training GPT models using weight sparsity without significant loss in downstream task metrics Read more here: https://
hubs.li/Q01twKNt0 #NeurIPS2022 #sparsity #ai #ml -
GPT-4 and the Turing Test Comparison
By
–
GPT-4 could treat the Turing test the same way we treat CAPTCHA
-
Cerebras and Cirrascale Launch Affordable LLM Cloud Service
By
–
If you’re working on large language models on a budget, Cerebras and Cirrascale offer a cloud service starting at $2,500. Large language models are about to be a giant commercial phenomenon, says Cerebras CEO Feldman. // @cerebras // #AI #deeplearning #machinelearning