Oh, reading a bit more about llama.cpp (
https://
github.com/ggerganov/llam
a.cpp
…), that's only inference, not training? I haven't tried since I don't have the model checkpoints on my laptop, but you may be able to use gptq.int4 quantization then: https://
github.com/Lightning-AI/l
it-gpt/blob/main/tutorials/quantize.md
…
LLMS
-
llama.cpp inference and gptq quantization techniques exploration
By
–
-
QLoRA 4-bit NormalFloat format supported only on Nvidia GPUs
By
–
Ah sorry, I meant M1/M2 chips (not specifically M1/2 CPUs). As far as I know, the 4-bit NormalFloat format that is used in QLoRA is currently only supported on Nvidia GPUs (
https://
github.com/TimDettmers/bi
tsandbytes/issues/485
…). Maybe the repo you mentioned uses a different type of quantized training. -

LoRA Compared to Llama-Adapter and Llama-Adapter v2
By
–
LoRA is a parameter-efficient finetuning technique, yes. I recently compared to Llama-Adapter and Llama-Adapter v2:
-

Enterprise AI ROI: Master LLM Inference Strategy for Maximum Impact
By
–
Get your inference strategy right & your enterprise will achieve a generational leap in the ROI of AI solutions for LLMs & other revolutionary workloads. For guidance, get our white paper, Key Enterprise Considerations for Inference Deployment of LLMs. http://
groq.com/inference -

UT Austin NLP Lectures: From Basics to LLMs and Beyond
By
–
Natural Language Processing – UT Austin A concise series of NLP lectures from UT Austin. Covers a vector of topics from basics of machine learning, NLP fundamentals, models(BERT, BART, T5, GPT-3…), and hot topics/trends in LLMs including instruction tuning, chain-of-thoughts,
-

SeamlessM4T Breakthrough in Multilingual Speech Translation
By
–
SeamlessM4T represents a significant breakthrough in the field of speech-to-speech & speech-to-text by addressing the challenges of limited language coverage & a reliance on separate systems. More details https://
bit.ly/45g2pMq -

Fine-tuning LLaMA2 Workshop: One-Day In-Person Event
By
–
If you want to explore finetuning LLaMA2, we'll be talking at and helping out with a 1 day in-person event focused explicitly on finetuning OSS models Hopefully this guide will come in handy! RSVP here (s/o @swyx and @NaderLikeLadder for organizing): https://
partiful.com/e/T4ngRPaU2uUT
XM8pN17d
… -
Apple App Store Model: GPT Derivatives Ecosystem Development
By
–
Also what became really obvious after the app store was that Apple was not going to be designing/building/discovering the killer use cases for the iPhone (except maybe passbook). I think the same thing is going to probably be true about GPT/Llama derivatives.
-
ChatGPT Works But Far From Fully Baked Says Expert
By
–
I mean you could say the same thing about ChatGPT. It actually *works* and is (sort of) immediately useful after a million false starts in AI, and you can look at it and see the potential for an industry-altering tech. But it's nowhere close to fully baked.
-
New LLM Survey Report on Production Adoption and Customization Techniques
By
–
Check out our new #LLM survey report to get a lowdown on the current state of LLMs in #production. Topics covered: LLM adoption Top use cases Key challenges when productionizing LLMs Techniques for customization incl. #finetuning and more! https://
pbase.ai/44jLmba.