We're super excited to be a part of this amazing #LLM Conf from the genius minds of @HamelHusain + @dan_s_becker alongside so many other thought leaders in AI. Today is your last chance to register and get all of the free credit goodness! Not to mention you will learn a ton!
LLMS
-

Perplexity updates iOS app with manual interrupt controls for voice interaction
By
–

Perplexity got a slightly updated push-to-talk UI on its iOS app. Now it has Stop and Pause buttons which allow you to interrupt the model manually It seems like Perplexity is shipping what Open AI has been promised but with more reliable controls.
-

Fine-tuning LLMs for 2x Faster Inference with Speculative Decoding
By
–
Fine-tuning #LLMs isn't just for customization You can also #finetune LLMs to increase throughput using speculative decoding Check out the replay of our recent tech talk to learn how we fine-tuned an LLM using to increase inference by 2x https://
pbase.ai/4bV6rNc. -

Scale launches SEAL Leaderboards for frontier LLM evaluation
By
–


New from Scale: SEAL Leaderboards — a new benchmark arena for frontier LLMs – Private, novel assessments that models can’t train on
– ELO-scale rankings (via Bradley-Terry)
– Domain leaderboards (today: coding, math, instruct, Spanish — more soon!) (Links in reply) -
GSM1k Model Released: Harder Benchmark Coming Soon
By
–
Ya. It’s GSM1k under the hood. We’ll have a harder one soon.
-
GSM-1k Benchmark Limitations and Upcoming Harder Math Evaluation
By
–
GSM-1k wasn’t really designed to distinguish between top models, more to detect overfitting. we will fix this for the next round with a harder math eval!
-
Flash Praised as Excellent Small Language Model
By
–
Yes, I’ve said this before, but Flash is a crazy good small model
-
LocalLlama community feedback as crucial evaluation metric
By
–
r/LocalLlama comments section remains a very important evals cross-check no matter what 🙂
-

LLM Evaluation Improvements and the Challenge of Creating Good Benchmarks
By
–
Nice, a serious contender to @lmsysorg in evaluating LLMs has entered the chat. LLM evals are improving, but not so long ago their state was very bleak, with qualitative experience very often disagreeing with quantitative rankings. This is because good evals are very difficult
-
Optimizing batch size and throughput for GPT-3 training
By
–
Okay that's good to know 🙂
I was just following the GPT-3 paper numbers in the table, but like I mentioned it's possible the settings are way too conservative.
Few more things to try: we want to increase the batch size as much as possible to get higher tok/s. Are you using -r 1