the way i see it, the last twelve months of AI research can be summed up in just two big breakthroughs: [i] reasoning ('test-time compute') – new ways to train models that can use more tokens to generate better answers. they mostly rely on RL with verifiable rewards [ii]
@jxmnop
-
Future Computers: Do We Still Need CPUs?
By
–
really dumb question. will future computers even need CPUs? seems they mostly exist to load data on and off the GPU. what’s the point of that
-
LLM Reviewers Raise Questions About Evaluation Authenticity
By
–
unfortunately many of my fellow reviewers are LLMs as well
-
Complicated ML Techniques Often Prove Less Useful Than Expected
By
–
a lot of the complicated stuff turned out not to be very useful or important. a few random examples: graphical models, variational autoencoders, RNN/LSTM, anything bayesian
-
NeurIPS Review Experience: LLM-Generated Papers Quality Issues
By
–
i just reviewed five papers for NeurIPS and it was an awful experience: – first paper was clearly LLM-generated. it was too short, the references didn't work, had no experiments or theory at all, and a ton of obvious mistakes. the more i read the less it made sense
– two were -
Interpretability as AI subfield: leading researchers and career opportunities
By
–
interpretability is definitely its own subfield, and imo very interesting & important, there just aren't as many open roles for interpretability work here are some people that come to mind @NeelNanda5 @hendrycks @ch402 @soniajoseph_ @csinva
-
Follow these AI robotics and research experts
By
–
not my area, but these accounts seem good to follow @chris_j_paxton @DrJimFan @drfeifei @BostonDynamics @physical_int @pabbeel
-
How Coders Can Learn AI: A Practical Learning Path
By
–
if you know how to code, and want to learn AI, this is what you should do: lucky for you the field has gotten *less* deep and much easier to learn about over the last few years. most lab roles require specific knowledge (model training, CUDA kernels, etc.). each takes ~18 months
-
Performance gains from optimized implementation in AI systems
By
–
i'll have to page @lo_LB_La for engineering details but iirc most of the performance gain comes from the optimized implementation rather than architectural details

