No, I asserted here that SGD was *a* thing. I mention elsewhere that the fact that a classical 100-layer transformer network obeys its own "100-step rule" is another thing that matters (and also is not inaccessibly deep mathematics).
AI
-
XR and spatial tech terminology co-opted during hype cycles
By
–
lol ya the word totally got co-opted. As someone who's been working in XR and then "spatial" tech since 2014 — it is a shame to see. But hey — it happens with every hype cycle.
-
Regulators and courts set AI precedent for closed companies
By
–
It's not about what people want, this will be decided by regulators or the courts. It's a nice way to set a precedent which the closed companies will then have to follow.
-
Human creativity transcends commodity content in AI era
By
–
you're assuming consumption experiences stay static. IMO, humans will always conjure up more complex canvases for our creativity — and thus that'll always mint a small number of creators who can enrapture the creators beyond the commodity content
-
Dataset Filtering for Large Language Model Training Runs
By
–
Possibly yes. More broadly, if we're going to be doing more large training runs at all, I'd guess we should start filtering the datasets soon (presumably using a previous-generation LLM finetuned for that?). It's hard to finetune out a cognition once it's learned by the base.
-
Open Source Translation Models: Transformers vs English Pipeline
By
–
yes open source, vrai travail, par contre sur la fin il parle de traduction et qu'il dit que ca passe par de l'anglais et là c'est pas du tout ça, ils passent par les transformers et pour l'entraînement et Meta a les même résultats en terme de traduction
-
LLMs Confabulate Inherently Despite Infinite Data
By
–
Conjecture: even with infinite data, LLMs would still confabulate, because they blur the inputs and don’t reliably create precise representations of individuals and their properties. (see The Algebraic Mind, 2001 for related discussion which has thus far has held true)
-

Baby GPT as Finite State Markov Chain Visualization
By
–
This is a baby GPT with two tokens 0/1 and context length of 3, viewing it as a finite state markov chain. It was trained on the sequence "111101111011110" for 50 iterations. The parameters and the architecture of the Transformer modifies the probabilities on the arrows. E.g. we
-
AGI Alignment Knowledge and LLM Developer Expertise Skepticism
By
–
Maybe I shouldn't, but one last shot at explaining my position on AGI gatekeeping. I'm not saying that I know everything known to the high-status inventor or engineer of the largest LLM. I'm saying that I'm dubious that they have secret knowledge deeply relevant to *alignment*,