At first, in the early 2020s, I worried that LLMs were often confidently wrong, calling them “fluent spouters of bullshit”. (And I was right; they have been and continue to be.). But now we have a new problem which is that *people* who *learn* from LLMs are also often
RESEARCH
-
Links to Reasoning from Scratch project and book
By
–
Ah yes, here are the links: GitHub: github.com/rasbt/reasoning-f… Manning: mng.bz/Nwr7 Amazon: amzn.to/4aAKiFY
-
Fine-tuning strategies and future plans for expanding capabilities
By
–
Thanks! The fine-tuning section is more general in the first one, and the second one is focused on math because it makes it easier from a security perspective. But I plan to add bonus materials over time, and function calling would be an interesting one.
-
Cognitive Surrender Research Covered by New York Times
By
–
"Cognitive surrender is clearly real, and with it will come the atrophy of certain skills and capacities, or the absence of their development in the first place." Fantastic coverage of my recent research with @steveshaw2020 by @ezraklein at @nytimes bit.ly/4lYzpBT
→ View original post on X — @sallyeaves, 2026-03-29 18:31 UTC
-

Self-Distillation Hidden Layers Self-Supervised Vision Models
By
–
"Self-Distillation of Hidden Layers for Self-Supervised Representation Learning" This paper shows that self-supervised vision models would work much better when they predict a teacher's hidden layers across the entire visual hierarchy. As they showed that just learning from the
-

HyperOffload: Compiler Framework Optimizes LLM Memory Management
By
–
Tired of LLMs maxing out memory even on powerful supernode architectures? Shanghai Jiao Tong University and Huawei Technologies Co., Ltd. introduce HyperOffload! This new compiler-assisted framework intelligently plans data movement for large language models. By treating
-
Human Science’s Ability to Generalize Models from Available Information
By
–
Basically, consider a high-profile scientific problem (one that is getting enough attention from smart humans and their externalized cognitive infrastructure). How good is human Science at converting the information available about the problem into generalizable models of the
-
AI Educational Outreach: Lectures, Essays, Blogs, and Social Media
By
–
KVCache quantization is a no-no as well I’d rather quantize the model to 2-bit rather than quantize the KVCache to 4-bit or even 8-bit
-

Gradient Descent Can Lead to Local Optima
By
–
I have already said my part on the newcomer community members Not that I think of myself as the OG, but I believe I have the track record to override anybody else in this space All that I want is to bring opinions closer, let's stop unnecessary toxicity
-

Getting Started with 3D Gaussian Splatting Tutorial
By
–
Wondering how to get started with 3d gaussian splatting? nitter.net/bilawalsidhu/status/20… Bilawal Sidhu (@bilawalsidhu) 3d gaussian splatting is really fun! If you've been curious about how to create radiance fields using your phone, DSLR, 360 cameras, or high-end LiDAR-based rigs, then I think you'll enjoy my new deep dive. Watch it here: piped.video/ctraRclNiZA — https://nitter.net/bilawalsidhu/status/2014775576981098748#m
→ View original post on X — @bilawalsidhu, 2026-03-29 17:38 UTC
