(1/2) We are excited to share that our paper "Sparse Iso-FLOP Transformations for Maximizing Training Efficiency" has been accepted by the Workshop on Advancing Neural Network Training (WANT) at #neurips2023! Read our blog for more information: https://
cerebras.net/blog/can-spars
ity-make-ai-models-more-accurate
…
@cerebras
-

Sparse Transformations Maximize Neural Network Training Efficiency
By
–
-

State-of-the-Art LLM Training Techniques and Learnings
By
–
(2/2) Our paper details learnings from training state-of-the-art LLMs, including: -Rotary and ALiBi position embeddings
-Swish-gated linear unit (SwiGLU)
-Overtraining on many Tokens-Per-Parameter
-LR Decay Ratio
-Maximal update parameterization (muP) https://
arxiv.org/abs/2309.11568 -
Cerebras Accelerates 175B LLM Training with PyTorch Single Device
By
–
Cerebras presented at the PyTorch Developer Conference and covered how we accelerate Large Language Model (175B+) training using PyTorch, Torch-MLIR, and the simplicity of single device training without a single torch.distributed instruction Watch here:
-

Cerebras Co-Founder Presents AI Hardware Innovations at OktoberTech
By
–
Cerebras Co-Founder JP Fricker had a great time presenting at OktoberTech 2023! This annual event by @Infineon Technologies brought together some of the brightest minds in the tech industry to share their perspectives on the latest trends and innovations. #Digitalization
-

ALiBi Models Performance Extension Without Fine-Tuning
By
–
Unlike existing techniques – this method requires no additional fine tuning. By setting a few parameters in your HuggingFace config.json, you can instantly extend the performance of ALiBi models such as BTLM-3B-8K by ~2x: https://
huggingface.co/cerebras/btlm-
3b-8k-base#during-inference-without-fine-tuning
… -

Position Interpolation Extends ALiBi Model Context From 8K to 16K
By
–
We show that position interpolation works just as well in the context of ALiBi – models trained with 8K of context can now extrapolate up to 16K.
-
Position Interpolation Improves ALiBi Model Extrapolation
By
–
To make ALiBi extrapolate better, we applied position interpolation – an idea popularized in models that use RoPE. Position interpolation scales the input to fit in the context length used during training, thus generating more stable results.
-
ALiBi Context Extrapolation Limits in Production Language Models
By
–
ALiBi promises fast, long extrapolation but in production models only extrapolate 10-20% longer than its training context. This is why ALiBi models such as BTLM-3B-8K and MPT-7B-8K are trained on a mix of 2K and 8K contexts – ALiBi on its own cannot extrapolate from 2K to 8K.
-
Position Interpolation Doubles Context Length for ALiBi Models
By
–
Paper drop: Position Interpolation Improves ALiBi Extrapolation We found a simple method to 2x the context length of models that use ALiBi. This lets models like BTLM-3B-8K and MPT-7B-8K run high quality inference at up to 16K with no additional fine tuning.
-
India Team Doing Fundamental Engineering Work Says CEO
By
–
"We don't believe in an offshore development center model; we believe in hiring good engineers and giving them big projects to run. Our India team ..is doing fundamental work for us." @andrewdfeldman said this to @timesofindia in the article found here: