This is entirely speculative but… Tomorrow's LS paper club with @picocreator is going to be extremely lit! come learn about the state of the art in Speculative Decoding!
LLMS
-

Phi-3.5 MoE: Official Mixture of Experts Model Performance
By
–
Super funny to see an official Phi MoE! Phixtral is a little project I made 8 months ago by combining 2-4 finetunes with MergeKit. It worked better than expected. The Phi-3.5 MoE looks great on benchmark, curious to see how it performs in practice. Model:
-

RNNs with Expressive Hidden States Match Strongest Transformers
By
–
Excited to feature Learning to (Learn at Test Time): RNNs with Expressive Hidden States! This new architecture replaces RNN hidden states with a machine model, matching or exceeding the strongest transformers and Mamba. The author @karansdalal is here to answer your questions!
-
OpenAI Lets Customers Customize GPT-4o Model
By
–
An @openai exclusive from me: the company will let customers customize its flagship model, GPT-4o. this comes as competition grows for AI products for business, and businesses face growing pressure to demonstrate that investments in AI are worth it.
-

Genie achieves 43.8% score on SWE-bench Verified benchmark
By
–
Genie case study benchmarks With a fine-tuned GPT-4o model, Genie achieves a SOTA score of 43.8% on the new SWE-bench(opens in a new window) Verified benchmark, announced last Tuesday.
-
Domino Simplifies LLM Training with Ray Support
By
–
Domino simplifies and accelerates training complex LLMs. Using native Ray support and pre-built reference projects, teams can fine-tune Llama 70B without understanding the specifics or topology of hardware or infrastructure! See for yourself in this demo! https://
domino.buzz/4dMwAib -
China Approves 190 LLMs and Strengthens Generative AI Governance
By
–
[#Article] GenAI Adoption: China Approves 190 LLMs, Strengthens Generative AI and Cyberspace Governance https://actuia.com/actualite/adoption-de-la-genai-la-chine-approuve-190-llms-renforce-la-gouvernance-de-lia-generative-et-du-cyberespace/
… #AI #ArtificialIntelligence -

LongVILA: Scaling Long-Context Visual Language Models for Videos
By
–
LongVILA Scaling Long-Context Visual Language Models for Long Videos discuss: https://
huggingface.co/papers/2408.10
188
… Long-context capability is critical for multi-modal foundation models. We introduce LongVILA, a full-stack solution for long-context vision-language models, including system, -

Cybench: Framework for Evaluating Language Models Cybersecurity Capabilities
By
–
Cybench A Framework for Evaluating Cybersecurity Capabilities and Risk of Language Models discuss: https://
huggingface.co/papers/2408.08
926
… Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have the potential to -
xGen-MM BLIP-3 Live Presentation by Manli Shu and Le Xue
By
–
.
@ManliShu and @Le_Xue01 presented xGen-MM (BLIP-3) live on X today if you missed the live session see the recording here: https://
x.com/i/broadcasts/1
vOxwrzOBpqJB
…
