Stanford testing all 11 major models and finding sycophancy across every single one means this isn't a ChatGPT problem or an OpenAI problem. It's a training methodology problem that the entire industry shares.
LLMS
-
Microsoft Announces 3 New World-Class MAI Models in Foundry
By
–
More on the blog today: microsoft.ai/news/today-were…
-
Microsoft Releases MAI-Transcribe-1 Transcription Model
By
–
Three models. Three top-tier results. All shipped within just a few months by the @MicrosoftAI team.
— Mustafa Suleyman (@mustafasuleyman) 2 avril 2026
– MAI-Transcribe-1 dropped today, the most accurate transcription model in the world across 25 languages according to FLEURS WER benchmark.
– MAI-Voice-1 sets a new standard for… pic.twitter.com/rfwqcEaxR8Three models. Three top-tier results. All shipped within just a few months by the @MicrosoftAI team.
– MAI-Transcribe-1 dropped today, the most accurate transcription model in the world across 25 languages according to FLEURS WER benchmark.
– MAI-Voice-1 sets a new standard for -

LLaMA-Factory: Fine-Tune 100+ LLMs Without Coding
By
–
If you found it useful, reshare it with your network Follow me → @Sumanth_077 for more insights and tutorials on AI Engineering! nitter.net/Sumanth_077/status/203… Sumanth (@Sumanth_077) Fine-Tune 100+ LLMs without writing a single line of code! LLaMA-Factory lets you train and fine-tune open-source LLMs and VLMs without writing any code. Here's why it's a game changer for fine-tuning: • Fine-tune 100+ LLMs/VLMs with built-in templates (LLaMA, Gemma, Qwen, Mistral, DeepSeek, and more). • Zero-code CLI & Web UI for training, inference, merging, and evaluation. • Supports full-tuning, LoRA, QLoRA, freeze-tuning, PPO/DPO, OFT, reward modeling, and multi-modal fine-tuning. • Speeds up training/inference with FlashAttention-2, RoPE scaling, Liger Kernel, and vLLM backend. • Integrates experiment tracking via LlamaBoard, TensorBoard, Weights & Biases, MLflow, and SwanLab. It's 100% Open Source Link to the Github Repo in the comments! — https://nitter.net/Sumanth_077/status/2039701710659272775#m
→ View original post on X — @sumanth_077, 2026-04-02 13:50 UTC
-

LLaMA-Factory: Fine-Tune 100+ LLMs Without Code
By
–
Fine-Tune 100+ LLMs without writing a single line of code! LLaMA-Factory lets you train and fine-tune open-source LLMs and VLMs without writing any code. Here's why it's a game changer for fine-tuning: • Fine-tune 100+ LLMs/VLMs with built-in templates (LLaMA, Gemma, Qwen, Mistral, DeepSeek, and more). • Zero-code CLI & Web UI for training, inference, merging, and evaluation. • Supports full-tuning, LoRA, QLoRA, freeze-tuning, PPO/DPO, OFT, reward modeling, and multi-modal fine-tuning. • Speeds up training/inference with FlashAttention-2, RoPE scaling, Liger Kernel, and vLLM backend. • Integrates experiment tracking via LlamaBoard, TensorBoard, Weights & Biases, MLflow, and SwanLab. It's 100% Open Source Link to the Github Repo in the comments!
→ View original post on X — @sumanth_077, 2026-04-02 13:49 UTC
-
Github Repo: LlamaFactory – Advanced Language Model Fine-tuning Framework
By
–
Github Repo: github.com/hiyouga/LlamaFact…
→ View original post on X — @sumanth_077, 2026-04-02 13:49 UTC
-

Alibaba releases Qwen 3.6 Plus with coding and vision
By
–
Alibaba released Qwen 3.6 Plus, an upgraded agentic model with coding and vision capabilities. Qwen 3.6 Plus comes with a 1M context window and is already available on Qwen Chat. https://
x.com/Alibaba_Qwen/s
/Alibaba_Qwen/status/2039697007489765727
… -
Claude Opus outperforms Sonnet on complex coding tasks
By
–
I keep pushing Claude Code the hardest and Opus has been amazing for complex tasks. The difference between Sonnet and Opus for following large skill files is quite noticeable on my end.
-
Claude vs GPT: Performance Comparison for Daily Work
By
–
The pace of releases is getting hard to keep up with haha. On my end, Claude's models usually hold up quite nicely vs. GPT's for actual day to day work. We'll see if that changes it!
