@upstageai fine-tuned Solar-Proofread w/ Predibase, achieving 79% accuracy vs. fine-tuned GPT-4o mini’s 71% for a major international media company. Discover the power of fine-tuned SLMs: https://
pbase.ai/4e1DQaE #AI #LLM #Upstage #Predibase #SmallLanguageModels
LLMS
-
Fine-tuned Solar-Proofread Outperforms GPT-4o Mini
By
–
-

LLMs Outpace Human Researchers in Novelty Study
By
–
Are LLMs Outpacing Human Researchers? A Groundbreaking Study Says Yes! In a recent study, researchers compared the novelty of ideas generated by 100 NLP experts with those from Large Language Models (LLMs) —and the results are eye-opening. LLMs outperformed human
-

Require secret, third‑party LLM benchmarks before trusting models
By
–
“I would not trust any claims of a superior model until I see [both] ELO points on LMSys [and] private LLM evaluation from a trusted 3rd party, such as Scale AI's benchmark. The test set must be well-curated and held secret, otherwise it quickly loses potency.”
-

Fast API Llama 405B: 4X Speed Single Rack Efficiency
By
–
We had to hop on the trend Some very demure, very mindful facts about our fast API: 4X the speed of other Llama 405B providers Runs on a single rack at less than 19kW Full precision with the highest accuracy Composition of Experts model to support customer
-
New cookbook: Multi-agent web browsing system with Transformers Agents
By
–
New cookbook: Multi-agent web browsing system! To support the launch of multi-agent systems in our Transformers Agents library, I highlight how to build one of these hierarchies to make a system that browses the web and leverages code to efficiently answer user questions!
-
Community-driven transparency reshaping open-source AI research
By
–
regardless of the outcome and results of the reflection-70b dive-in, the amazing thing about open-weights and open-source is how the whole community can dive in these questions together and study the model transparency in AI is becoming so crucial as the field grow
-
Using ChatGPT conversations for collaborative social experiences
By
–
Social aspect is there too. It is like using same ChatGPT conversation with friends
-

Build and sell GPTs with this creator mega-prompt
By
–
I'm shocked why people aren't building and selling GPTs. I just made a GPT creator mega-prompt you can use to build GPTs. Build, sell, and make money! Like + comment "GPT" and I'll DM you the file. (Must be following me to receive it)
-
AI Evaluation Ecosystem Needs Better Benchmark Overfitting Detection
By
–
The whole Reflection-70B debacle points the the desperate need for a better AI evaluation ecosystem. It needs to be extremely easy to adjudicate:
(1) is the model overfit to benchmarks
(2) is the model truly unique (i.e. not a wrapper or thin fine-tune) https://
x.com/shinboson/stat
/shinboson/status/1832933747529834747
… -
Compact coding assistant challenges AI industry’s model size obsession
By
–
Try this "a powerful yet surprisingly compact coding assistant that threatens to disrupt the AI industry’s fixation on ever-larger models"