just software optimizations. still FP16 weights.
@cerebras
-
Compiler Optimizations and Memory Management for FP16 Models
By
–
compiler optimizations, better memory management etc. same FP16 weights.
-
Cerebras Accelerates AI Inference Speed by 10x for Deeper Reasoning
By
–
o1 takes takes multiple minutes for inference. cerebras reduces inference time by 10x
faster inference = deeper CoT / rollouts -
Scaling LLM Training: Double Digit Batch Size for 70B Models
By
–
double digit batch size, eventual goal is one rack for 70B.
-

Cerebras Inference Speed Update: Llama 3.1 Performance Benchmarks
By
–
Cerebras Inference perf update:
Llama3.1-8B: 1,8001,927 tokens/s
Llama3.1-70B: 450481 tokens/s
Stillfor the most popular open model in the world. https://
inference.cerebras.ai -
PyTorch Training and CUDA Low Level Optimizations
By
–
Training is already done in PyTorch. CUDA is for low level optimizations.
-

NANDA: Advanced Hindi Large Language Model Unveiled
By
–
Excited to share the development of NANDA, a cutting-edge Hindi Large Language Model, created with our partners G42, Inception, and MBZUAI (Mohamed bin Zayed University of Artificial Intelligence). NANDA was trained on Condor Galaxy, one of the world’s most powerful AI
-

Cerebras Releases DocChat: GPT-4 Level Document AI Model
By
–
Cerebras Announces DocChat: GPT-4 Level Conversational QA Trained in a Few Hours DocChat is our latest document-based conversational AI model series, that includes two models: Cerebras Llama3-DocChat – An LLM that was trained to GPT-4 level document-based Q&A
-

Cerebras Accelerates GenAI Progress in Life Sciences Healthcare
By
–
Cerebras Accelerates GenAI progress across Life Sciences and Healthcare Generative AI is transforming the life sciences and healthcare industries. Cerebras is helping leaders like GSK and Mayo develop AI breakthroughs by training their own models for molecular design, drug