BREAKING: Elon Musk’s xAI just launched the API for Grok 3. It’s fast, expensive, and ready for real-world use. Here’s everything you need to know
GENERATIVE AI
-
Seed-Thinking-v1.5: Advanced AI Model Capabilities Showcase
By
–
Now, can you imagine how strong Seed-Thinking-v1.5 will be?
Seed-Thinking-v1.5: https://
github.com/ByteDance-Seed
/Seed-Thinking-v1.5/blob/main/seed-thinking-v1.5.pdf
…
DAPO: https://
arxiv.org/abs/2503.14476
VAPO: -

Huawei Pangu Ultra 135B Matches DeepSeek-R1 Performance
By
–
Huawei is gearing up to release Pangu Ultra—a 135B dense Transformer trained on 13.2T tokens with 8,192 Ascend NPUs. Despite fewer parameters, it matches DeepSeek-R1 in performance.
No full tech report or model yet. Check it out: https://
arxiv.org/abs/2504.07866 -

Google begins limited rollout of Veo 2 on AI Studio
By
–
BREAKING : Seeing multiple reports that Veo 2 is rolling out on Google AI Studio. Seems like it is a very limited rollout for now. h/t @HarisHiew
-
GPT 4.5 Podcast Watch Party with Sam Altman
By
–
ok apparently there is TOO MUCH ALPHA in @sama
's GPT 4.5 podcast today https://
discord.gg/kTARCma3?event
=1360137758089281567
… we are hosting a live watch party in 5 mins join -
Secret Algorithm Behind Unbeatable Digital Masterpiece
By
–
One day, my digital masterpiece caught the eye of Alok, our soft-spoken instructor. He was quite fascinated but, being a man of few words, all I could get out of him was "Hmm, interesting." When he asked how I made it unbeatable, I explained that I had a secret algorithm, which
-

DDT: Decoupled Diffusion Transformer for Image Generation
By
–
DDT: Decoupled Diffusion Transformer
Hugging Face:
https://
huggingface.co/papers/2504.05
741
…
Paper:
https://
arxiv.org/pdf/2504.05741
Code:
https://
github.com/MCG-NJU/DDT -

Decoupled Transformer Improves Diffusion Model Convergence
By
–
Can a decoupled encoder-decoder Transformer bring faster convergence and better sample quality to diffusion models?
Check out Decoupled Diffusion Transformer, today's #1 paper on Hugging Face!
From Nanjing University & ByteDance—worth a deep dive. -

Scaling Laws for Native Multimodal Models
By
–
Paper:Scaling Laws for Native Multimodal Models
Link: https://
arxiv.org/pdf/2504.07951 -

Apple Researchers Discover Scaling Laws Native Multimodal Models
By
–
Researchers from Apple just uncovered scaling laws for native multimodal models. What did they find?
– No clear advantage for late-fusion over early-fusion.
– Early-fusion performs better at small scales.
– Add MoEs, and you unlock modality-specific weights→big performance gains
