As an AI researcher, ordering a trash can and it coming from “Qualia zero” seems like something is telling me I am still in the Matrix.
@id_aa_carmack
-
Anisotropic processor dies could improve memory bandwidth efficiency
By
–
I wonder if there would be advantages to making some highly anisotropic processor/GPU/TPU dies — long and skinny instead of square-ish. The “shoreline” around the chip is often a limiting factor for memory/IO bandwidth, and the defect rate should be the same by area.
-
Core Algorithm Tiny, Substantial Scaffolding Required Around It
By
–
Yeah, the core algorithm will be tiny, but there will be a decent chunk of scaffolding around it. For the people that bring up the bulk of PyTorch/cuda/drivers around it, which are millions of lines of code — the path actually used will just be… a few tens of thousands of lines
-
Higher Dimensional Spaces in Modern Machine Learning Insights
By
–
Completely off topic with no subtext whatsoever, but since you are clearly familiar with the principles and the classic thinking about them, you may be interested in one of the critical insights from modern machine learning— in higher dimensional spaces, there basically aren’t
-
Finding a Better Term for Speed of Light Analysis in System Engineering
By
–
I often use “speed of light analysis” when talking about system engineering for a task, but I should probably find another term. I use it around questions like “what is the minimum latency for pass through video and synthetic frames in this architecture”, but “speed of light”
-
Semiconductor fab complexity and potential simplification possibilities
By
–
A state of the art fab today may be the most complex and sophisticated thing built by humans; you can’t “just go build one” and compete with TSMC given any amount of resources. However, they are general purpose systems with a lot of flexibility. How much simpler could they get
-
Modern computing reliability challenges compared to Cray 1 era
By
–
It just feels like we are back in the days of the Cray 1 with miles of not-very-reliable wires all over the place, waiting for VLSI to improve both performance and reliability.
-
GPU Cloud Instances Should Be Auctioned Instead of Flat-Rate Pricing
By
–
GPU cloud instances are a scarce resource, and you generally can’t get what you want when you want it. Why isn’t GPU time being auctioned, instead of sold at flat rates? Yes, consumers hate variable pricing, but businesses understand it, and auctions underpin a good chunk of the
-
ML Advances 3D Capture Quality Despite Performance Loading Challenges
By
–
ML can finally clean up 3D captures to be actually good now instead of just novel — 3D photos will be a significant value to almost everyone that likes 2D photos. Taking 100x as long to load as a normal photo in social media feeds will be a significant headwind, though.
-
Criticism of Grace Hopper’s 1:1 CPU-GPU Ratio Architecture
By
–
I actually don’t like the path they are taking — I would like to see them enable ever larger GPU to CPU ratios, with one CPU commanding 256 NVLINK GPUs, but instead they are going 1:1 with Grace Hopper, which is a step back from todays baseline.
