LENS, a zero trainable params multimodal framework is shown below. It's frozen CLIP, BLIP, and LLM doing all the works.
@jeande_d
-

LENS: Vision Language Models Using Frozen Modules
By
–
Towards Language Models That Can See:
Computer Vision Through the LENS of Natural Language LENS is a modular approach that can solve computer vision and vision language tasks using frozen vision modules and LLM. LENS can achieve reasonably good zero-shot/few-shot -

DeepRob: Deep Learning for Robot Perception Course Michigan 2023
By
–
DeepRob: Deep Learning for Robot Perception – Michigan, 2023 This is a great course that covers fundamentals of neural networks with application in robot perception. Also covers techniques for training and debugging neural nets and emerging topics in deep learning for robotics
-

Guide pratique sur l’utilisation des LLMs : modèles, données et applications
By
–
Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond This survey paper provides a comprehensive and practical guide on working with LLMs for various settings(in-context learning & fine-tuning). I like how it approach LLMs usage in lens of models, data, and
-

LLM in Production: Essential Playlist by MLOps Community
By
–
LLM in Production – By MLOps Community This is an excellent playlist of LLMs in Production. Contains great talks and discussions around LLMs, their use-cases and how to make value with them. https://
youtube.com/playlist?list=
PL3vkEKxWd-us5YvvuvYkjP_QGlgUq3tpA
… -
Multimodal Technology Emerges as Top Prediction for Future
By
–
Hard to predict what’s next these days but multimodal is a top candidate for sure.
-

Awesome Multimodal Large Language Models: Papers and Datasets
By
–
Awesome Multimodal Large Language Models A curated list of the latest papers and datasets in multimodal large models. Provides an excellent listing of multimodal papers, codes, and demos. Papers about topics like instruction tuning, in-context learning, chain-of-thought(CoT),
-
Normalization Placement in Transformer Architectures: Text vs Vision
By
–
Thanks! The relative position of normalization is one of the few things that changed about the original transformer architecture. I think it's not exactly clear where it should be placed. Most transformers for texts use post-norm(one above) whereas vision transformers tends to
-

Transformer Neural Network Graphics as Therapeutic Practice
By
–
The highest form of therapy is reproducing the transformer neural network graphic 😀
-

Comprehensive Survey on Transformer Applications Across Multiple Deep Learning Domains
By
–
A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks A great survey paper that provides a comprehensive analysis of highly influential transformer-based models in top five application domains: NLP, Computer Vision, Multi-Modality, Audio and Speech
