SOTA sweep:
HealthBench 65.1 / Hard 44.4 Hallucination 3.5% (lower than ChatGPT) ScanBench all-stations #1: 74.9 / 72.1 / 74.4
@baichuanai
-
Baichuan AI Achieves State-of-the-Art Results Across Multiple Benchmarks
By
–
-
SPAR, Fact-Aware RL, and Rubric Evolution in AI Training
By
–
Key takeaways:
SPAR: align RL credit to where decisions happen — optimize stage-wise, not via one noisy end reward. Fact-Aware RL: verify atomic claims with retrieval → make hallucination measurable & optimizable
Rubric Evolution: auto-mine & patch adversarial reward hacks. -
Baichuan-M3 Technical Report Released for Clinical Decision Support
By
–
Baichuan-M3 Technical Report is live. Report:
https://
arxiv.org/abs/2602.06570
Models: https://
hf.co/collections/ba
ichuan-inc/baichuan-m3
…
Try: https://
ying.ai
Built for clinical decision support, not trivia QA, optimized on the real outpatient workflow: Inquiry → Lab Testing → Diagnosis →→ -

Baichuan-M3 Medical LLM Open-Sourced, Outperforms GPT-5.2
By
–
We have officially open-sourced the new-generation medical LLM Baichuan-M3, which scored 65.1 on HealthBench and claimed the championship with 44.4 on HealthBench Hard. It has comprehensively outperformed GPT-5.2 across the board in the medical field. @OpenAI @claudeai
-

M2 Plus Shows Significantly Lower Medical Hallucination Rate Than DeepSeek
By
–
Evaluation shows that the medical hallucination rate of M2 Plus is significantly lower than that of general LLMs, about three times lower than DeepSeek. Now it’s available on web (
http://
ying.ai) and in mobile app stores. -

Baichuan Launches Enhanced Medical Model Baichuan-M2 Plus with API
By
–
We have just launched our enhanced evidence-based medical model, Baichuan-M2 Plus, while simultaneously upgrading its supporting application Baixiaoying and opening the API.
-
Baichuan-M2: Medical Reasoning Model Surpasses Open-Source Competitors
By
–
Today we introduced Baichuan-M2, a medically-enhanced reasoning model, which surpasses all open-source models including gpt-oss-120b on the #HealthBench Benchmark. Try it here: https://
huggingface.co/baichuan-inc/B
aichuan-M2-32B
…
#AIHealth -

Baichuan-Omni-1.5 Surpasses GPT-4o Mini in Multimodal Processing
By
–
In terms of visual, speech, and multi-modal streaming processing, Baichuan-Omni-1.5 supasses GPT-4o mini, and its leading edge is even more pronounced in the field of multi-modal medical applications.
-

Baichuan Launches Omni-1.5 Multimodal Open-Source Model
By
–
We launched Baichuan-Omni-1.5 open-source omni-modal large model. It not only supports omni-modal understanding of text, image, audio, and video, but also has dual modalities generation capabilities for both text and audio.
-

Baichuan-M1-14B: Medical AI Model for Healthcare Evidence
By
–
It has unlocked a medical evidence-based model, enabling a complete end-to-end service from medical evidence retrieval to deep reasoning. To better promote AI healthcare ecosystem, we also launched the industry's first open-source medical enhancement large model, Baichuan-M1-14B.