The main options are fp32 (or tf32), bf16, or fp16+mixed prec afaik.
CODE
-
FP16 Training Optimization: Dynamic Loss Scaling Without Mixed Precision
By
–
I suspect the problem you're testing here might be too easy to see differences – you're getting nearly 100% either way. Would be interesting to see fp16 with dynamic loss scaling but without mixed precision to compare to there.
-
BF16 vs FP16 Training Performance Comparison Analysis
By
–
There was some comparison of training performance by @StasBekman recently that showed bf16 non-mixed training was quite a bit worse. Which I'd expect, because bf16 has much less precision than even fp16.
-

Open-Source RAG ChatGPT Web App AI Tutor Launch
By
–
How we Built an Open-Source RAG-based ChatGPT Web App: Meet Our new AI Tutor! This is Towards AI's new AI tutor, a question-answering chatbot built to answer anything about LLMs with up-to-date information! (work in progress, please try it and give us feedback, its free!) Learn
-
Multi-Agent Simulation Framework for Strategy Testing
By
–
Let's recap. Agent: You describe your agents, their objectives and goals Environment: Define the environment using just text Simulation: Start the simulation and observe agents' strategies, gain valuable insights in a simulated environment
-
BF16 Benefits Beyond Loss Scaling in Mixed Precision
By
–
Other than avoiding loss scaling, are there other benefits to using bf16? It still seems to require mixed precision afaict, right?
-
Distributed GPU Processing: File Synchronization and Aggregation
By
–
because in distributed mode, each GPU runs its own process. so this code will run simultaneously multiple times, once per GPU. so in this setting, each GPU writes a file, waits for everyone, then aggregates all the files, waits again, then deletes its own file 🙂
-
Building RAG-Powered Chatbots with LLM Chat Endpoints
By
–
What can you build with LLM chatbots powered by retrieval-augmented generation (RAG)? The Chat endpoint makes it easy to build RAG-powered chatbots. In this LLM University chapter, learn the foundations of LLM chatbots and the Chat endpoint. https://
txt.cohere.com/exploring-chat
-rag/
… -

Parallel Dataset Embedding Computing with HuggingFace and DDP
By
–
most useful bit of code I've written all year: call map() on a HuggingFace dataset in torch distributed mode (like DDP) as one example, this will let you compute embeddings for a dataset in parallel, using all the GPUs you have http://
gist.github.com/jxmorris12/69a
730fee174f5309968e984c298f8f2
…