Today we’re announcing the recipients of our LLM Evaluation research grants. These four projects will each receive $200K in grant funding from Meta to further their work in novel new work in evaluations over the next year. MMRLU: Massive Multitask Regional Language
LLMS
-

Gemini Hallucinations on AI Existential Risk Documentation
By
–
and if I try to feed it the URL of the docs page that it’s on (maybe that way Gemini can read the page content), it hallucinates a response about climate change when the piece is about ai existential risk vibe shifts… @DynamicWebPaige @DaveCitron
-

Large Language Model Generates Protein Equivalent to 500 Million Years Evolution
By
–
Whoa! When a large language of life model generates a protein equivalent to ~500 million years of evolution. @ScienceMagazine http://
science.org/doi/10.1126/sc
ience.ads0018
… @THayes427 @proteinrosh @EvoscaleAI @arcinstitute @UCBerkeley -

ChatGPT subjects PerePChabrier and SylvainLyve to a lie detector
By
–
ChatGPT puts @PerePChabrier and @SylvainLyve to the lie detector → https://youtu.be/VTI818S-csU So, who is telling the truth in the #Vilebrequin drama? It's… surprising! PS: the video is kind and educational to illustrate everything AI can do in 2025!
-

SwiftKV reduces Llama inference costs by up to 75%
By
–
In December, @SnowflakeDB AI Research announced SwiftKV, a new approach that reduces inference computation during prompt processing. Today they're making SwiftKV-optimized Llama models available on Cortex AI that reduce inference costs by up to 75%!
-

FACTS Grounding: New Benchmark for LLM Factuality Evaluation
By
–
FACTS Grounding: A new benchmark for evaluating the factuality of large language models https://
bit.ly/3DuVWV6
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
Prompt Caching Documentation for Claude API
By
–
Check out the docs here: https://
docs.anthropic.com/en/docs/build-
with-claude/prompt-caching
… -
Anthropic Improves Prompt Caching with Automatic Cache Hit Detection
By
–
Quality-of-life upgrade for @AnthropicAI devs: We've adjusted prompt caching so that you now only need to specify cache write points in your prompts – we'll automatically check for cache hits at previous positions. No more manual tracking of read locations needed.
