For small training sets, models use superposition to memorize more data points than the two available neurons. For large training sets, models learn features in superposition, as observed in our previous work, allowing the model to generalize. https://
transformer-circuits.pub/2022/toy_model
/index.html
…
RESEARCH
-

Superposition in Neural Networks: Memorization vs Feature Learning
By
–
-

Superposition Strategy: How Neural Networks Embed Features in Hidden Space
By
–
Our prior work showed that these toy models use a strategy called “superposition” to learn more features than available neurons. Here we observe how training data points, as well as features, are embedded in the hidden space.
-

Understanding Deep Learning Overfitting Through Mechanistic Analysis
By
–
We have little mechanistic understanding of how deep learning models overfit to their training data, despite it being a central problem. Here we extend our previous work on toy models to shed light on how models generalize beyond their training data. https://
transformer-circuits.pub/2023/toy-doubl
e-descent/index.html
… -
LLMs Will Get Better This Year, Solving Trivial Problems
By
–
Prototypes are easy, production is hard. However, LLMs will get a **lot** better this year. Expect many unsolved problems to become trivial. But not all.
-
The fundamental difference between software and non-software AI development trajectories
By
–
Yes, that’s the real explanation for it — sci-fi AI was imagined as software for consumer hardware, so once it eats the tail of its development it goes FOOM. The fact it isn’t software changes everything.
-

Digital Transformation Origins: Data Analysis 20000 Years Ago
By
–
Fascinating! When did #DigitalTransformation really begin? How humans developed #insights with #data & #analysis 20,000 years ago @BBCNews See http://
bit.ly/IceDataAge #tech #innovation #research #art @JolaBurnett #STEM @Hana_ElSayyed @EstelaMandela @AudreyDesisto @YvesMulkers -

Top 10 Tech Trends 2023: Large Models to Quantum Computing
By
–
Get a sneak peek at the future of tech with our annual predictions for the top 10 #TechTrends of 2023! From big models to virtual-real symbiosis, and quantum computing, there's so much to look forward to. For more: http://
research.baidu.com/Blog/index-vie
w?id=178
… -
AI’s Next Leap: Scaling and Balancing Resource Allocation
By
–
Yeah, figuring out better ways to scale and balance how much of something (or nothing of something) you get is maybe the next big AI leap to look forward to
-
Right Data Essential Updated Monthly For AI Systems
By
–
The right data is essential, updated monthly
-

Narrow AI vs General AI: ChatGPT on AI Development Challenges
By
–
Here are #chatgpt3 answers on @Twitter ownership. @elonmusk is it fair to say narrow AI is primetime and general AI will be as difficult to harness as human will and discretion?