WebExplorer Explore and Evolve for Training Long-Horizon Web Agents
@_akhaliq
-

Sonoma Sky Alpha and Sonoma Dusk Alpha Stealth Models Released
By
–
Sonoma Sky Alpha and Sonoma Dusk Alpha two new stealth models are now available in anycoder via @openrouter Context: 2 million tokens
-

Transition Models: Rethinking the Generative Learning Objective
By
–
Transition Models Rethinking the Generative Learning Objective
-
Planning with Reasoning using Vision Language World Model
By
–
Planning with Reasoning using Vision Language World Model pic.twitter.com/RZF8SVpV17
— AK (@_akhaliq) 4 septembre 2025Planning with Reasoning using Vision Language World Model
-

EmbeddingGemma: Multilingual On-Device Embedding Model in Browser
By
–
model carrot vibe coding another Google EmbeddingGemma app, a sota multilingual embedding model perfect for on-device use cases using transformers.js and onnx-community/embeddinggemma-300m-ONNX At only 308M params, the model can run 100% locally in your browser app:
-
MiniCPM-V 4.5 Achieves 77.0 Score Surpassing GPT-4o and Gemini
By
–
MiniCPM-V 4.5
— AK (@_akhaliq) 4 septembre 2025
achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro, and strong open-source models like Qwen2.5-VL 72B
powered… pic.twitter.com/H7s4zanvSKMiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro, and strong open-source models like Qwen2.5-VL 72B powered
-
Vibe Coding: Deploying Chat App with Prompts
By
–
vibe coding and deployed chat app for it in anycoder in a few prompts
-

T2R-bench: Industrial Table-to-Report Generation Benchmark
By
–
T2R-bench A Benchmark for Generating Article-Level Reports from Real World Industrial Tables
-

PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
By
–
PVPO Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning

