I think at least one person is working on an STL model but suspect it will be a while before anyone gets to commercializable results. (Thought of doing it myself; seems like not the thing to spend years of life on.) In principle it’s ~inevitable it happens.
HARDWARE
-
Running Large Models on Consumer GPUs: Parallelism and Quantization
By
–
I don't really use GPUs, as most of our work uses TPUs. I was just confused by @jeremyphoward 's statement that it wouldn't run on consumer GPUs, when the normal techniques of model parallelism, quantization, right compression, etc. seem like they should work just fine to run it
-
Model Efficiency on Consumer-Grade GPU Hardware
By
–
Why? It seems as though the model can run on 4 consumer-grade GPUs, no?
-
Running AI Models on Consumer GPUs: Feasibility Analysis
By
–
You can probably run the model on 4 (or maybe 8?) consumer GPUs, no?
-
Google’s Integrated AI Stack: Hardware to Cloud Services
By
–
Google has the entire AI stack: hardware (TPUs), data centers, large web datasets, training infrastructure, research, and cloud services to serve models. The ability to optimize across all parts of the stack end-to-end has incredible advantages.
-

Llama 4 Now Available on Cerebras Inference Platform
By
–
Llama 4 is here and it’s coming to Cerebras! Starting next week, we will be serving Llama 4 on Cerebras Inference at instant speed. Thank you to the @AIatMeta team for your partnership. Be the first to get access here: https://
cerebras.ai/build-with-us -
New AI Models Released for Wafer Scale Hardware Performance
By
–
Amazing release! Can't wait to show the world how fast these models can go on wafer scale hardware!