It's "raw" lm-evaluation-harness, I didn't add the recommended system prompt.
GENERATIVE AI
-

Reflection technique between AI models by Fernando Neto
By
–
Example from the great @FernandoNetoAi of reflection with another model.
-

Reflection CoT Evaluation Methods Comparability Issues
By
–
This is super cool but I have a lot of questions. First, reflection = CoT on steroids. It means you can't compare these scores at all. Remember when people made fun of Gemini for providing CoT results for MMLU? This is a lot worse. Secondly, if you don't parse the output and
-
Populating AI People Game with Artificial Intelligence Characters
By
–
Not if we populate the with AI People @aipeoplegame
-

Flux LoRAs Eliminates Need for Model Training
By
–
Thank god for Flux LoRAs i don't need to train anymore
-
French PM Must Balance Daily Concerns With Tomorrow’s Structural Challenges
By
–
Le défi du nouveau PM : répondre aux préoccupations du quotidien des Français (pouvoir d’achat, sécurité), tout en prenant à bras le corps les défis de demain, certes plus abstraits, mais structurants pour l’avenir : IA, dette, énergie, éducation… Difficile, mais pas impossible.
-

Google DeepMind AlphaProteo and top AI stories today
By
–
Top stories in AI today: -Google DeepMind reveals ‘AlphaProteo’
-New AI agent builds apps from prompts
-Find top prompts with Google’s Prompt Gallery
-AI creates infinite Super Mario Bros game
-6 new AI tools & 4 new AI jobs Read more: http://
therundown.ai/p/google-alpha
proteo-closer-to-curing-cancer
… -

New 70B Model Available on Hugging Face When Main Demo Busy
By
–
if the main demo is busy looks like you can try the new 70B model here on @huggingface : https://
huggingface.co/spaces/feather
less-ai/try-this-model
… -
Link to the Reflection-Llama-3.1-70B model on the Hub
By
–
Check it out on the Hub -> https://
huggingface.co/mattshumer/Ref
lection-Llama-3.1-70B
… -

New 70B LLM Outperforms Claude-3.5-Sonnet and GPT-4o with Innovative Fine-Tuning
By
–
> A new 70B open-source LLM beats Claude-3.5-Sonnet and GPT-4o! Matt Schumer, CEO from Hyperwrite AI, had an idea he wanted to try out: why not fine-tune LLMs to always output their thoughts in specific parts, delineated by tags? Even better: inside of that, you