To better enable the community to build on our work — and contribute to the responsible development of LLMs — we've published further details about the architecture, training compute, approach to fine-tuning & more for Llama 2 in a new paper. Full paper https://
bit.ly/44JAELQ
LLMS
-

Meta Publishes Llama 2 Technical Paper for Community Development
By
–
-

Cerebras Opentensor Host AMA on BTLM-3B-8K Model
By
–
Reminder that at 10:00 am PT today, Cerebras and Opentensor will host an AMA on our Discord server to talk about BTLM-3B-8K. Come to ask questions, engage in a discussion, or simply enjoy the conversations! Join our Discord here: https://
hubs.li/Q01ZhwD80 -
Techniques for Constructing Prompts Command Model
By
–
Prompts can be as simple as a one-liner, or they can be as complex as multiple layers of specific information. In this article, we look at some techniques for constructing prompts for the Command model. https://
short.cohere.ai/construct?utm_
source=twitter&utm_medium=social
… -
Llama 2 Coding Limitations Impact on Agentic AI Tasks
By
–
Had an awesome webinar with @RLanceMartin and @jamescalam on Weds around using Llama 2 Llama2 is significantly worse on coding tasks – is that why it's not nearly as good at agentic tasks as GPT models? Up on YouTube now:
-
Leandojo Democratizes Theorem Proving with Large Language Models
By
–
Leandojo democratizes theorem proving with LLMs. Visit @KaiyuYang4 poster today at @icmlconf
-
Lit-GPT: Unified Codebase for Decoder Analysis and Comparison
By
–
It’s kind of hard to analyze across repos due to implementation differences and details along the data loading and finetuning pipelines. That’s where I’d say Lit-GPT is useful because it makes the set of relevant decoders available in the same unified code base
-
Model Documentation Gaps and Empirical Analysis Importance
By
–
Not that I am aware of. On top of that several models don’t even have thorough papers themselves because people are currently rushing them out. I guess the best way is really some empirical analysis coupled with some knowledge bits like falcon and llama2 use multiquery attention
-
Mastering LLM Basics as Foundation for Innovation
By
–
I don't disagree, but I think that in order to innovate it can also be helpful to have the basics down. I.e., implementing and running state-of-the-art LLMs is kind of part of the baseline exercise before going further from there.
-

NeurIPS LLM Efficiency Challenge: Train Model in One Day
By
–
Trying to develop the next huge LLM is fun, but how about a side project with a more reasonable scope like the 1 LLM + 1 GPU + 1 Day NeurIPS efficiency challenge: https://
llm-efficiency-challenge.github.io Below a list of the approved models (PS: Lit-GPT was chosen as the official starter kit ) -

LLaMA-like Model Pretraining Accelerated 38% Open-Source
By
–
65-Billion-Parameter Large Model Pretraining Accelerated by 38%, Best Practices for Building LLaMA-like Base Models Open-Source | https://
syncedreview.com/2023/07/18/65-
billion-parameter-large-model-pretraining-accelerated-by-38-best-practices-for-building-llama-like-base-models-open-source/
…
#AI #ML #ArtificialIntelligence #MachineLearning #DeepNeuralNetwork