‘MrsFormer’ Employs a Nove Multiresolution-Head Attention Mechanism to Cut Transformers’ Compute and Memory Costs https://
syncedreview.com/2022/11/14/mrs
former-employs-a-nove-multiresolution-head-attention-mechanism-to-cut-transformers-compute-and-memory-costs/
…
LLMS
-
MrsFormer: Multiresolution Attention Cuts Transformer Costs
By
–
-
LangChain 0.0.13 Release: Vector DB QA and Documentation Updates
By
–
LangChain Version 0.0.13 Question/Answering w/ a vector DB chain (demo coming tmrw) Loading a prompt from a text file (
@edmarferreira first commit!) Misc cleanup (Eugene x4!!!) Big Documentation overhaul (w/ Eugene again) -
CRINGE Loss: Learning What Language Not to Model
By
–
The CRINGE Loss: Learning what language not to model Adolphs et al.: https://
arxiv.org/abs/2211.05826 #ArtificialIntelligence #DeepLearning #MachineLearning -
GPT-2’s Inscrutable Internal Matrices Remain Poorly Understood
By
–
Nobody's ever even going to understand how GPT-2 worked, except that there sure were a lot of inscrutable matrices in there.
-
Generative engines and the indexing of human thought
By
–
The global indexing of human thought + a dash of GPT3 = generative engines.
-
Symmetry Reduces Parameters Through Parameter Tying
By
–
Symmetry can also exponentially reduce the number of parameters by parameter tying.
-
Analysis of GPT-4 rumors and model capabilities
By
–
GPT4 > ? https://
thealgorithmicbridge.substack.com/p/gpt-4-rumors
-from-silicon-valley
… -
How Artificial Intelligence Perceives The Mission District
By
–
What does the Mission look like to an artificial intelligence? – Mission Local Read more here: https://
ift.tt/PBG1jz8 #ArtificialIntelligence #AI #DataScience #100DaysOfCode #Python #MachineLearning #BigData #DeepLearning #NLP #Robots #IoT -

Efficiently Scaling Transformer Inference Techniques
By
–
Efficiently Scaling Transformer Inference Pope et al.: https://
arxiv.org/abs/2211.05102 #ArtificialIntelligence #DeepLearning #MachineLearning -
Hybrid AI Systems More Practical Than Full LLM Retraining
By
–
I love the dramatic futurism here! But it's more likely they will be hybrid systems. You don't want to retrain a LLM everytime a Wikipedia page is changed, or a news item is published. (Plus, Google would be in the best place to deliver this new system, incrementally.)