Introducing long-context transformer using mini sequences. It is a simple and effective method for highly efficient and accurate LLM training with extremely long sequences. Our research demonstrates that the Llama3-8B model can be trained with context lengths up to 60k tokens on
LLMS
-
PaliGemma Fine-tuned Model for Visual Question Answering
By
–
https://
huggingface.co/abhishek/autot
rain-paligemma-finetuned-vqa
… -
LLMs Have Become a Commodity in Their Current Form
By
–
LLMs, in its current form, has truly become a commodity.
-
Llama 3.1 405B Launch: Open Weights, GPT-4o Comparable Performance
By
–
Llama 3.1 405B launches • Comparable to GPT-4o and Claude 3.5 Sonnet, according to the benchmarks
• The weights are publicly available
• 128k context even on the smaller models Discussion on r/ML -
GPT-4 Turbo Outperforms GPT-4o on Key Evaluation Dimensions
By
–
indeed. we haven’t finished our 4o mini eval, but our evals indicate that 4 turbo is still better than 4o on many important dimensions
-

Infrastructure Setup and Scripts for 70B Model Deployment
By
–
From bare metal to a 70B model: infrastructure set-up and scripts https://
bit.ly/45XsNfi
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
GPT-4o mini fine-tuning now live for users
By
–
and also, GPT-4o mini fine-tuning now live! https://t.co/WHWTvNwRNR
— Sam Altman (@sama) 23 juillet 2024and also, GPT-4o mini fine-tuning now live!
-

GPT-4o mini achieves near GPT-4o performance at fraction of cost
By
–
we try not to get too excited about any one eval, but excited to see GPT-4o mini so close to GPT-4o performance on lmsys at 1/20th the price!
-
Meta Releases Llama 3.1 Models to Community
By
–
We're thrilled to make these models available to the community and can't wait to see what they do with Llama 3.1!
-
Offline vs Online Distillation: Implementation Discrepancy Analysis
By
–
It feels like they're talking about "offline" distillation here, but actually implemented online distillation as described in the paper
