Gradient descent has it backward: the steeper the slope, the riskier and less necessary large steps are, but the larger the steps it takes.
@pmddomingos
-
LLMs Face Unclear Use Cases Like Previous NLP Technologies
By
–
The problem with LLMs is what it's always been with NLP: if people can't tell what they're good for and not, they'll wind up not using them for much.
-
Research Boldness Inversely Proportional to Review Panel Size
By
–
The boldness of the research funded is inversely proportional to the size of the review panel.
-
Hierarchical Living Systems: P≠NP and Evolutionary Constraints
By
–
Living systems have hierarchical structure because P ≠ NP, and evolution can only try exponential combinations of a small number of things at a time.
-
Crypto Collapse Frees Hype Capacity for Language Models
By
–
The collapse of crypto is great news for large language models, because it frees up a lot of hype capacity.
-
Transformers Possess Greater Compositional Power Than MLPs
By
–
Those tasks are irrelevant to my point. You seem to be denying that transformers have more compositional power than MLPs.
-
Transformers Outperform MLPs for Parsing Tasks
By
–
Transformers are better than MLPs far from the limit. And these papers are about something else. All it takes for my statement to be correct is that a transformer can learn to parse correctly in some cases.
-
Transformers Represent Grammar Better Than MLPs
By
–
No, transformers can represent grammar better than (say) MLPs, which is an advance. And they can learn it to a very limited degree, so what I said is accurate.
-
Encoding Grammar in Transformers: Gradient Descent Learning Challenges
By
–
You can encode a grammar into a transformer. The problem is gradient descent is not very good at learning it.