thank you so much for the quick turnaround! Feature Request: 2factor, and the 2fac enforced for Admins of orgs — just like Github does as of last year.
@soumithchintala
-
Rust prevents memory safety issues in software development
By
–
if you used rust, there wouldn't be any sliding.
-
AMD Single Node Training: Scaling Challenges for Large Models
By
–
single node training. and that makes a difference. i’m not fully sold about amd and multinode large scale training, but one step at a time
-

MosaicML Demonstrates LLM Training on AMD MI250 MI300X
By
–
Here's @MosaicML showcasing their results with @PyTorch 2.0 + AMD for LLM Training.
They made "Zero Code Changes" to run on AMD.
MI250 is already trending well, and IMO MI300X will be very competitive. -
CoreWeave Infrastructure Deals: Opex, Equity, and Risk Trade-offs
By
–
usually such deals are frontloaded-opex for discount or in exchange of equity. Whether CoreWeave builds it for Inflection (In this case) or builds it and sells it later to someone else is a mere detail that does some trade-off on debt, equity, leverage and risk
-
Software optimization and co-design for H100 GPU maximum potential
By
–
yea, I expect a healthy amount of software work and co-design to use the H100 to its maximum potential!
-
H100 Training Speed Triples A100 with FP8 Optimization
By
–
As @MosaicML showcased in April, on GPT training H100 is ~3x the speed of A100, if you use FP8 training, which is both seriously impressive and grounds you in relative improvement versus the previous generation! https://
mosaicml.com/blog/coreweave
-nvidia-h100-part-1
… -
MLPerf GPT-3 Training Results Clarification Explained
By
–
Context: People are misunderstanding the GPT-3 – 11 minutes training result from the latest MLPerf.
reference: -
GPT-3 Training Speed: Clarifying the 11-Minute Claim
By
–
No, GPT-3 wasn't trained in 11 minutes. The GPT-3 architecture was trained on the C4 dataset to 2.69 log-probability in 11 minutes on 3584 H100 GPUs. Don't focus on the "11 minutes" — because it's like saying "ResNet-50 was trained in 5 seconds on MNIST to 80% accuracy"
-
AI Research Integrity Crisis: Academic Fraud Parallels Emerge
By
–
What a hustle!
I feel sorry for the students on the paper for starting their journey in such a careless, low-integrity environment. The last time AI hit a bubble was ~2014, and this paper reminds me of Baidu cheating on the Imagenet competition. Similar playbook — threw a https://
x.com/ML_PhDer/statu
/ML_PhDer/status/1672750801234857984
…