One thing we know is that if future AI systems are built on the same blueprint as current Auto-Regressive LLMs, they may become highly knowledgeable but they will still be dumb.
They will still hallucinate, they will still be difficult to control, and they will still merely
@ylecun
-
Why Current Auto-Regressive LLMs Will Remain Fundamentally Limited
By
–
-
Autoregressive LLMs limitations: knowledge without intelligence
By
–
One thing we know is that if future AI systems are built on the same blueprint as current Auto-Regressive LLMs, they may become highly knowledgeable but they will still be dumb.
They will still hallucinate, they will still be difficult to control, and they will still merely -
CNN History: From Convolutional Networks to Computer Vision
By
–
I started using the phrase Convolutional Network (without the Neural) in the late 90s when "neural" had become taboo in the ML community.
The acronym CNN was originally used by Rama Chellappa and later adopted by the Computer Vision community around 2013.
I never liked this -
DETR: Detection Transformer Architecture Overview
By
–
DETR (and others) https://
arxiv.org/abs/2005.12872 -
Getting Started with Llama-2: Comprehensive Tutorial Guide
By
–
How to get started with Llama-2 ?
Here is a comprehensive tutorial. -
Convolution Equivariance vs Self-Attention Permutation Properties
By
–
Convolution is equivariant to translations.
Self-attention is equivariant to permutations.
They both have a role to play.
Conv is efficient for signals with strong local correlations and motifs that can appear anywhere.
SelfAtt is good for "object-based" representations where -
Leading Mathematicians and Computer Scientists Petition for Kidnapped Children
By
–
Petition "freedom for kidnapped children"
Signed by laureates of the Fields Medal, Abel Prize, Nevanlinna Prize, Breakthrough Prize, ACM Turing Award, and ACM Prize in Computing. -
Equivariant operators and local connections in neural architectures
By
–
I'd say taking out local connections and shared parameters over locations is like taking out bread from pizza. Lots of operators are local and equivariant to translations besides convolutions.
-
ViT and ConvNets Achieve Equal Performance at Same Compute
By
–
Compute is all you need.
For a given amount of compute, ViT and ConvNets perform the same. Quote from this DeepMind article: "Although the success of ViTs in computer vision is extremely impressive, in our view there is no strong evidence to suggest that pre-trained ViTs -
AI as Solution for Social Media Content Moderation and Safety
By
–
When people talk about issues with social networks, they don't realize that AI is not the cause, AI is actually an essential piece of the solution. To take down propaganda, hate speech, attacks on democracy, child exploitation, calls to violence, dangerous misinformation, etc in