you can skip the DL chapter about boltzmann machines. replace it with material on Transformers
@jxmnop
-

Master AI: Read Two Essential Textbooks Cover to Cover
By
–
take a year out of your life and read these two textbooks cover-to-cover. you will already know more than 90% of people in AI
-
Meta Trains Sonar AI Model for Advanced Applications
By
–
apparently they do train a model called Sonar?
-
Different Architecture and Pre-trained Model Approach
By
–
it's a very different architecture and starting from a different pre-trained model!
-

Can Orange Line Catch Blue Line Without Waiting?
By
–
POLL: is the orange line going to catch up to the blue line? is there any way to know without just waiting? (it will take 5 days) this feels like the Halting Problem
-
Formalizing LLM Vibes: Toward Empirical Definition and Evaluation
By
–
i'm sure it won't be too long until someone writes a paper claiming to propose a formal definition & evaluation of LLM "vibes"
-
Fine-tuning Embedding Models: Why BERT Outperforms Larger T5 Encoders
By
–
question: i’m finetuning embedding models for a retrieval. i found large T5 encoders from GTR and SentenceT5 families significantly underperform BERT base how is this possible? feels it must be hyperparameters? unless i found the one NLP task ever where scaling doesn’t help
-
Combining DDP, bf16, gradcache and accelerate techniques
By
–
lol kind of. it was nontrivial for me to combine DDP, bf16, gradcache, and accelerate. maybe we could work together
-
GradCache Optimization for Distributed GPU Training
By
–
oh yeah. I think his trick was just GradCache. and sharing negatives between GPUs isn’t trivial
-
Framework Features Beyond PyTorch Lightning Trainer
By
–
what would u expect from this kind of framework that u don’t get from eg pytorch lightning trainer?