Robust benchmarking tools are crucial for developing complex AI systems. The #MosaicAI team created an evaluation suite to stress-test RAG workflows with long contexts. See the benchmark results for new state-of-the-art models from OpenAI & Google Gemini:
MosaicAI Benchmark Suite Evaluates RAG Workflows Performance
By
–