10). Granite Code Models – introduce Granite, a series of code models trained with code written in 116 programming languages; it consists of models ranging in size from 3 to 34 billion parameters, suitable for applications ranging from application modernization tasks to on-device
LLMS
-
MAmmoTH2: Harvesting Web Data to Enhance LLM Reasoning
By
–
9). MAmmoTH2 – harvest 10 million naturally existing instruction data from the pre-training web corpus to enhance LLM reasoning; the approach first recalls relevant documents, extracts instruction-response pairs, and then refines the extracted pairs using open-source LLMs;
-

Consistency LLMs: Parallel Decoders Reduce Inference Latency
By
–
6). Consistency LLMs – uses efficient parallel decoders that reduce inference latency by decoding n-token sequence per inference step; inspired by he human's ability to form complete sentences before articulating word by word…
-

Flash Attention Stability: Numeric Deviation Effects Analysis
By
–
7). Is Flash Attention Stable? – develops an approach to understanding the effects of numeric deviation and applies it to the widely-adopted Flash Attention optimization…
-
DrEureka Automates Sim-to-Real Robot Design Using LLMs
By
–
5). DrEureka – uses LLMs to automate and accelerate sim-to-real design; it requires the physics simulation for the target task and automatically constructs reward functions and domain randomization distributions to support real-world transfer…https://t.co/4k5bHESjbD
— DAIR.AI (@dair_ai) 12 mai 20245). DrEureka – uses LLMs to automate and accelerate sim-to-real design; it requires the physics simulation for the target task and automatically constructs reward functions and domain randomization distributions to support real-world transfer…
-

xLSTM Scales LSTMs to Billions Parameters with Modern LLM Techniques
By
–
2). xLSTM – attempts to scale LSTMs to billions of parameters using techniques from modern LLMs; to enable LSTMs the ability to revise storage decisions, they introduce exponential gating and a new memory mixing mechanism…
-
DeepSeek-V2: 236B MoE Model with Efficient Latent Attention
By
–
3). DeepSeek-V2 – a strong MoE 236B parameter model, of which 21B are activated for each token; supports a context length of 128K tokens and uses Multi-head Latent Attention (MLA) for efficient inference by compressing the Key-Value (KV) cache into a latent vector…
-

AlphaMath Zero: MCTS Enhances LLM Mathematical Reasoning
By
–
4). AlphaMath Almost Zero – enhances LLMs with Monte Carlo Tree Search (MCTS) to improve mathematical reasoning capabilities; the MCTS framework extends the LLM to achieve a more effective balance between exploration and exploitation…
-
Top ML Papers Week May 6-12 xLSTM AlphaFold DeepSeek
By
–
The Top ML Papers of the Week (May 6 – May 12): – xLSTM
– DrEureka
– AlphaFold 3
– DeepSeek-V2
– Consistency LLMs
– AlphaMath Almost Zero
… -

Analysis of ChatGPT Conversational Mode UI and Functionality
By
–
Something new to add regarding phone calls I like this theory but these strings are likely related to an existing functionality that allows ChatGPT to be connected with a car headset. If you are in a conversational mode, it would be shown as an actual phone call