To demonstrate the effectiveness of CulturalGround, we fine-tune an existing multimodel model Pangea on a subset of the dataset. The resulting model achieves state-of-the-art performance for its model size on multiple cultural benchmarks in PangeaBench and other benchmarks. Below
@jeande_d
-

CulturalGround: Building Inclusive Multimodal LLMs for Global Contexts
By
–
Current multimodal LLMs excel in English and Western contexts but struggle with cultural knowledge from underrepresented regions and languages. How can we build truly globally inclusive vision-language models? We are introducing CulturalGround, a large-scale dataset with 22M
-

CulturalGround: Scalable Multilingual VQA Data Pipeline from Wikidata
By
–
CulturalGround data construction: We designed a scalable pipeline to create culturally grounded multilingual VQA data from Wikidata, a structured knowledge base. Our data curation pipeline:
– Cultural Entity Selection: Extract 3M+ culturally relevant entities from Wikidata -
Computers: Few Loops Over Trillions of Transistors
By
–
A computer? a few for loops over zillion transistors
-
Revisiting Incomplete Computer Vision Project After 3 Years
By
–
I wanted to make this happen 3yrs ago but couldn't finish it, and closed the case. 15%, part1/part2 were for vision fundamentals. https://
github.com/Nyandwi/deep-c
omputer-vision
… Richard's CV algorithms and applications(
https://
szeliski.org/Book/) strikes a nice balance and I agree with @giffmana
, -
Reinforcement Learning of Large Language Models Course
By
–
Reinforcement Learning of Large Language Models Youtube playlist: https://
youtube.com/playlist?list=
PLir0BWtR5vRp5dqaouyMU-oTSzaU5LK9r
… Website: https://
ernestryu.com/courses/RL-LLM
.html
…
Tweet: -

UCLA Spring 2025: Reinforcement Learning of Large Language Models Course
By
–
Reinforcement Learning of Large Language Models, Spring 2025(UCLA) Great set of new lectures on reinforcement learning of LLMs. Covers a wide range of topics related to RLxLLMs such as basics/foundations, test-time compute, RLHF, and RL with verifiable rewards(RLVR).
-

Data Scarcity and Pretraining: The New AI Exploration Era
By
–
Good blog on "era of exploration" – Data scarcity is the new bottleneck. LLMs consume data far faster than humans can produce it. We're running out of high-quality training data. – Pretraining solved exploration by accident. Pretraining effectively pays a massive, upfront
-

Stanford CS336: Complete LLM Development from Data to Deployment
By
–
Stanford CS336 – Language Modeling from Scratch Excellent new lectures on "language modelling from scratch" class from Stanford just dropped. Covers a whole LLM stack from data collection, cleaning, transformer modelling and training, evals and deployment.
