Current multimodal LLMs excel in English and Western contexts but struggle with cultural knowledge from underrepresented regions and languages. How can we build truly globally inclusive vision-language models? We are introducing CulturalGround, a large-scale dataset with 22M
CulturalGround: Building Inclusive Multimodal LLMs for Global Contexts
By
–
