M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning paper page: https://
huggingface.co/papers/2306.04
387
… Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks.
@_akhaliq
-

M3IT: Large-Scale Multi-Modal Multilingual Instruction Tuning Dataset
By
–
-

Trending AI News Stories and Papers
By
–
Trending AI news stories + papers https://
open.substack.com/pub/akhaliq/p/
trending-ai-news-stories-papers-07a
… -
Vision Pro and AI/VR Tools: The Next Big Opportunities
By
–
vision pro is the next bug thing ChatGPT is dead I found 1 million AI/VR tools released this week to help you make a bajillion dollars Don’t get left behind
-

InternLM: Multilingual 104B Parameter Language Model Released
By
–
InternLM: A Multilingual Language Model with Progressively Enhanced Capabilities github: https://
github.com/InternLM/Inter
nLM-techreport
… present InternLM, a multilingual foundational language model with 104B parameters. InternLM is pre-trained on a large corpora with 1.6T tokens with a multi-phase -

Recognize Anything: Strong Foundation Model for Image Tagging
By
–
Recognize Anything: A Strong Image Tagging Model paper page: https://
huggingface.co/papers/2306.03
514
…
demo: https://
huggingface.co/spaces/xinyu12
05/Tag2Text
… present the Recognize Anything Model (RAM): a strong foundation model for image tagging. RAM can recognize any common category with high accuracy. RAM introduces a -

Grounding Instructional Steps in Narrated How-To Videos
By
–
Learning to Ground Instructional Articles in Videos through Narrations paper page: https://
huggingface.co/papers/2306.03
802
… present an approach for localizing steps of procedural activities in narrated how-to videos. To deal with the scarcity of labeled data at scale, we source the step -

Emergent Correspondence Discovery in Image Diffusion Models
By
–
Emergent Correspondence from Image Diffusion paper page: https://
huggingface.co/papers/2306.03
881
… Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We -

Mega-TTS: Zero-Shot Text-to-Speech Scaling with Inductive Bias
By
–
Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias paper page: https://
huggingface.co/papers/2306.03
509
… Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot -

LEACE: Perfect Linear Concept Erasure in Closed Form
By
–
LEACE: Perfect linear concept erasure in closed form paper page: https://
huggingface.co/papers/2306.03
819
… introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while inflicting the least possible damage to
