The D in a GAN will, in principle and if ground hard enough, try to learn exactly which errors humans make, and knowing this would make D smarter than a human. That said, I still think there's an important in-practice sense where GANs have more imitation-nature than predictors.
MACHINE LEARNING
-
PCA’s Relevance to AI Alignment Work Questioned
By
–
*Sigh.* I knew that meant "principal components analysis" without looking it up. I also knew about the controversy with ICA being patented. Now, what specific implication do you think PCA has for alignment work, exactly?
-
SGD and Transformer Network Depth Rules Matter
By
–
No, I asserted here that SGD was *a* thing. I mention elsewhere that the fact that a classical 100-layer transformer network obeys its own "100-step rule" is another thing that matters (and also is not inaccessibly deep mathematics).
-
Dataset Filtering for Large Language Model Training Runs
By
–
Possibly yes. More broadly, if we're going to be doing more large training runs at all, I'd guess we should start filtering the datasets soon (presumably using a previous-generation LLM finetuned for that?). It's hard to finetune out a cognition once it's learned by the base.
-

Baby GPT as Finite State Markov Chain Visualization
By
–
This is a baby GPT with two tokens 0/1 and context length of 3, viewing it as a finite state markov chain. It was trained on the sequence "111101111011110" for 50 iterations. The parameters and the architecture of the Transformer modifies the probabilities on the arrows. E.g. we
-
Gatekeeping the Gatekeepers of Deep Learning
By
–
I am not trying to gatekeep modern deep learning. The relevant truths of DL are open to anyone who's even slightly good at math! I'm trying to gatekeep the gatekeepers who shouldn't be allowed to gatekeep something with such a low gate.
-
Knowledge of Positional Embeddings and Attention Mechanism Improvements
By
–
Like I said, I already knew about positional embeddings being trained rather than being straight from the 2018 paper and had heard of axial transformers and various other attempts to defeat quadratic attention. (I think I can predict your next gatekeep, but go ahead and do it.)
-

Data Lakes and Industrial IoT Key Takeaways Explained
By
–
Want to dive deep into the world of #datalakes and #IIoT? Check out this article summarizing the key takeaways from the recent All Things IIoT Day session. http://
ow.ly/TpzV50NsCmV #sponsored #hitachivantara_iiot #datascience #industrialIoT @BigDataMinded @YvesMulkers via @fogoros -
AI Model Next-Token Distribution and Sampling Explained
By
–
Even for pre-trained models without tuning/RLHF, what’s actually modeled is the next-token distribution. A completion made by repeated temperature / top-p sampling has free parameters and isn’t directly optimized as a prediction of anything during training
-
Pythia: Suite for Analyzing LLMs Across Training and Scaling
By
–
9/ Pythia – a suite for analyzing LLMs across training and scaling; includes 16 LLMs trained on public data and ranging in size from 70M to 12B parameters.