More References:
6. torch-autograd: https://
github.com/twitter-archiv
e/torch-autograd
… 7. HIPS/Autograd: https://
github.com/HIPS/autograd
8. THNN refactors: * https://
github.com/torch/nn/pulls
?q=ispr+isclosed+THNN
… * https://
github.com/torch/cunn/pul
ls?q=ispr+isclosed+THCUNN
…
9. Online chat where the THNN organization happened: https://
app.gitter.im/#/room/#apaszk
e_THNN:gitter.im
…
@soumithchintala
-
PyTorch Autograd Framework References and THNN Refactoring
By
–
-
PyTorch Origins: Clarifying Ryan’s Lab Software Attribution
By
–
I've given a full and comprehensive answer here, but the tl;dr for @roydanroy is that PyTorch did not start with a git clone from Ryan's lab's software.
-
Clarifying PyTorch Origins: A Historical Deep Dive
By
–
this post was inspired by this thread that I was tagged into asking to (yet again) to clarify some details about the origins of PyTorch 🙂
-
PyTorch Origins: From Torch7 to Modern Deep Learning
By
–
PyTorch's design origins, its connection to Lua, its intertwined deep connection to JAX, its symbiotic connection to Chainer The groundwork for PyTorch originally started in early 2016, online, among a band of Torch7's contributors. Torch7 (~2010-2017)
These days, we also -
Reading Three Key Papers on Open-Source LLMs and Attention
By
–
oh man, people are focusing on "which paper" I read. I actually read three papers yesterday:
1. Are Open-Source LLM's catching up by @HailinChen3 et. al.: https://
arxiv.org/abs/2311.16989
2. System 2 Attention by @jaseweston and @tesatory : https://
arxiv.org/abs/2311.11829
3. (Part-way through -
Reading Three AI Papers: Open-Source LLMs and System 2 Attention
By
–
oh man, people are focusing on "which paper" I read. I actually read three papers yesterday:
1. Are Open-Source LLM's catching up by @HailinChen3 et. al.: https://
arxiv.org/abs/2311.16989
2. System 2 Attention by @jaseweston and @tesatory : https://
arxiv.org/abs/2311.11829
3. (Part-way through -
Academic Writing Brevity Debate in AI Research Community
By
–
Yesterday I read an 8-page paper. Breezed through it like a Netflix episode.
Clear, concise, and considerate of my time.
Somehow we've regressed to writing 30+ page epics (i'm guilty too). -
Batch Size Constraints: GPU SM Saturation and Memory Limits
By
–
For a large enough batch size on a given expert, you'll either saturate the SMs or run out of memory
-
MoE Inference Benefits: Why Mixture of Experts Improves Performance
By
–
It can be unintuitive why the Transformer-style MoE (in Mixtral/GPT4) has inference benefits.
Dima simplifies it with a clear explanation showcasing that MoE help inference once there's sufficient volume of requests (which hopefully are diverse enough that they don't hit the same -

Open-source models will outpace closed-source in the long term
By
–
the most convincing evidence I've seen so far that open-source will in-due-time accelerate away far beyond closed-source models.
— Soumith Chintala (@soumithchintala) 9 décembre 2023
People just want to use these models to best fit their use-case; a single company's closed-source effort rate-limits against people's imaginations and… https://t.co/Vw9Yp131lfthe most convincing evidence I've seen so far that open-source will in-due-time accelerate away far beyond closed-source models. People just want to use these models to best fit their use-case; a single company's closed-source effort rate-limits against people's imaginations and