Bi'an: A Bilingual Benchmark and Model for Hallucination Detection in Retrieval-Augmented Generation Bi’an introduces a bilingual benchmark dataset (Bi’anBench) and lightweight judge models for hallucination detection in Retrieval-Augmented Generation (RAG). The dataset spans
GENERATIVE AI
-

SWE-RL: Enhancing LLM Reasoning Through Software Evolution
By
–
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution SWE-RL is a reinforcement learning (RL) approach that enhances LLM reasoning for software engineering by learning from open-source software evolution data. It trains Llama3-SWE-RL-70B on GitHub
-

AI Co-Scientist: Multi-Agent System for Scientific Hypothesis Generation
By
–
Towards an AI co-scientist This paper introduces an AI co-scientist, a multi-agent system built on Gemini 2.0, designed to assist researchers by generating and refining novel scientific hypotheses. The system employs a "generate, debate, and evolve" framework to iteratively
-

AI Co-Scientists and LLM Reasoning Advances This Week
By
–
This week, AI is stepping into new dimensions—becoming co-scientists, sculpting 3D avatars, and blending cloud and on-device models into a seamless dance of creativity and efficiency. – Towards an AI co-scientist
– SWE-RL: Advancing LLM Reasoning via Reinforcement Learning -

Building Production-Ready RAG Systems with Qdrant FastAPI
By
–
RAG Pipeline Guide Learn to build production-ready RAG systems using Qdrant vector DB, FastAPI, and LangChain's LCEL for sophisticated retrieval chains. Key features:
– Async vector search
– Type-safe FastAPI endpoints
– LCEL RAG pipelines Start building better RAG systems! -
Claude 4.5 handles topic changes naturally
By
–
You can feel when a topic change breaks a model. 4.5 handles it very naturally.
-
Extended AI Conversations Demonstrate Coherence and Topic Versatility
By
–
I have a two-day long chat spanning so many topics. It's so good (and still completely coherent).
-
o3-mini-high excels in its specialized niche
By
–
Yes, o3-mini-high is definitely good at its own niche.
-
AI Standards Rising: From Coherence to Physics Problem-Solving
By
–
just a coherent sentence was impressive back then. now it needs to solve physics to get any attention
-
GPT-4.5 Performance Assessment Beyond Benchmarks
By
–
GPT-4.5 is the best model anywhere. Talk to it long enough and you will agree. Fuck the benchmarks.
