Sentiment analysis is not just a binary classification task. It's kind of like saying [insert any disease here] research is just a binary classification problem.
@rasbt
-
New scikit-learn Version Released with PyTorch Support and Features
By
–
It doesn't get much attention these days (in both senses) but a new version of @scikit_learn
, my favorite machine learning library, is out! – PyTorch support for LinearDiscriminant Analysis
– Validation Curves
– Decision tree with N/A features
– and more -
Appreciation for MPT Article and FSDP Implementation Question
By
–
Thanks for writing & sharing your latest article on MPT. Really well written and has just the right level of detail .
One question though. You wrote > "The entire training framework is based upon PyTorch’s Fully Sharded Data Parallel (FSDP) package and uses no pipeline or -

Stable Diffusion Based on Latent Diffusion Models Knowledge
By
–
I haven't read the full article, yet, and I also don't know much about the person in question (not familiar with any statements besides the occasional tweets that pop up on my timeline). That being said, I thought it was known that Stable Diffusion is based on the Latent
-
Stage 3 Offloading and CPUAdam Performance Comparison with FSDP
By
–
Relatively similar. I think stage 3 with offloading and CPUAdam was even a tad better but I’d have to double check again on Wed when I am back at my computer. I usually use DeepSpeed but opted for FSDP here to reduce external dependencies.
-
Why Transformers and Self-Attention Over Convolutional or RNN Layers
By
–
Sure. But why specifically for transformer layers and self-attention, not say convolutional or RNN layers?
-
Multiquery Attention Challenges with LoRA Porting Between Models
By
–
That’s absolutely true and a good point. I remember multiquery attention struggles when trying to port LoRA from LLaMA to Falcon.
-

TransformerEncoderLayer Abstraction Compared to Conv2D and RNN
By
–
Ok, fair. But if you look at the TransformerEncoderLayer class, for example, I don't think it's more abstract than nn.Conv2D or nn.RNN, for example. For some reason people just like implementing Transformer layers from scratch way more than convolutional or recurrent layers .
-
Everyone’s Using nn.Linear and Similar Neural Network Layers
By
–
But everyone's using nn.Linear etc.
-
PyTorch Transformer Classes Predate Most Online Tutorials
By
–
Interesting! I'd say the nn.TransformerEncoderLayer and nn.TransformerDecoderLayer classes etc. actually precede most tutorials out there. Chicken-egg problem maybe