Qué barbaridad. Un chip especializado sólo para ejecutar Transformers! Si nos parecía rápido Groq con los 800 tokens/segundo, esto llega a los 500.000 tokens/segundo ejecutando Llama 70!!!
COMPUTING
-
OpenAI leads AI models; Microsoft expands cloud GPU infrastructure dominance
By
–
'OpenAI’s models continue to lead the pack [for AI models], with open source models running second. Among the cloud infrastructure providers, Microsoft has maintained and actually even materially expanded its lead for hosting AI workloads/GPUs.' -UBS, citing IT executive survey
-
Exploring heyglif AI inference infrastructure and capabilities
By
–
wow @heyglif is very very cool Where do you run inference, @fabianstelzer
? -
Energy Panel at Reindustrialize Detroit Conference
By
–
Wheels down in Detroit for Reindustrialize.
— Packy McCormick (@packyM) 25 juin 2024
Moderating a panel on ENERGY tomorrow.
Let’s go @austinbishop @MikeSlagh @aphysicist 🇺🇸
pic.twitter.com/nRAE9RtAY1Wheels down in Detroit for Reindustrialize. Moderating a panel on ENERGY tomorrow. Let’s go @austinbishop @MikeSlagh @aphysicist
-
Compilation Performance Trade-offs in Modern Code Development
By
–
While making compilation slower and code more complex for 99% of users?
-

MathWorks Accelerates Engineering and Science Innovation Worldwide
By
–
MathWorks is dedicated to accelerating the pace of engineering and science worldwide. Learn more about what we do https://
spr.ly/6014ggsBa -
Undersized overtrained models optimize for inference compute
By
–
This is expected, though — the best models are undersized and overtrained because you expect to use them; you optimize for inference compute instead of just train
-
Luca Trevisan’s Final Talk on Spectral Graph Theory
By
–
In case you missed it, Luca Trevisan's final talk was delivered today by @praveshkkothari and Salil Vadhan. Part 1 is on spectral graph theory, and part 2 is about his experiences as an outsider. Things start around 22:30. https://
youtube.com/live/FdP_n_x0a
1k?si=5BNaFKS9cFvjwzZo
… -
llm.c Memory Efficiency Advantages Over PyTorch
By
–
If you match the parameters you're actually under-estimating the improvement, because llm.c takes up a lot less space so you can crank up the batch size. I haven't fully dug into why PyTorch takes up that much space, slightly worried I'm doing something wrong but not sure what
-
MLX Framework as Optimized Alternative to PyTorch MPS
By
–
I'd have a look at MLX, cc @awnihannun , from some recent chatter I understand it is a lot more optimized that PyTorch mps atm. This would need a re-write of the build-nanogpt code into mlx, but possibly it's not that involved. I'd be happy to accept a PR for mlx clone.
