A Survey of Large Language Model-Brained GUI Agents: https://arxiv.org/abs/2411.18279 (download 78-page PDF) #LLMs #GenAI #GenerativeAI #AI #MachineLearning #DataScience #DataScientist
LLMS
-
Stanford CS224N: Natural Language Processing with Deep Learning Course
By
–
I like this one. Introducing Stanford CS224N: Natural Language Processing with Deep Learning https://
web.stanford.edu/class/cs224n/ -
Tokenization vs arithmetic errors in LLMs
By
–
Less in the “profound” category but inspired by ppl saying this about “David Mayer” which seems to just be a string filter In general I think arithmetic and indexing failures are over-attributed to tokenization; e.g. digit tokenization is more consistent than most realize
-
On LLM behavior and the ‘tokenization issue’ excuse
By
–
No LLM behavior is so complex, profound, or absurd that someone, somewhere, will not suggest “tokenization issue” as the full explanation for why it happens.
-

Workaround for ChatGPT string-filtering in prompts
By
–

This seems to work because ChatGPT checks for the exact string “David Mayer” — any trick that avoids that string in both input and output, like asking it to say “DavidMayer” or “d4v1d m4y3r” or “David Meyer” (with two e’s), should work:
-
Removing 10 Minute Session Limit for Local AI NPC
By
–
Yeah, we removed the 10 minute session limit. Plus when running AI NPC (LLM) on your local PC, there's no cost, and no time limit to it. Transparency notice: if you will continue using TTS (text to speech), this is still running in cloud and there's cost associated to it. We
-
ChatGPT Language Games: Measuring Bullshit in LLM Outputs
By
–
10). Measuring Bullshit in Language Games Played by ChatGPT – proposes that LLM-based chatbots play the ‘language game of bullshit’; by asking ChatGPT to generate scientific articles on topics where it has no knowledge or competence, the authors were able to provide a reference
-

Survey on LLM-as-a-Judge: Building Reliable Evaluation Systems
By
–
7). Survey on LLM-as-a-Judge – provides a comprehensive survey of LLM-as-a-Judge, including a deeper discussion on how to build reliable LLM-as-a-Judge systems.
-

TÜLU 3 releases state-of-the-art open post-trained models
By
–
8). TÜLU 3 – releases a family of fully-open state-of-the-art post-trained models, alongside its data, code, and training recipes, serving as a comprehensive guide for modern post-training techniques.
-

Qwen2.5 Achieves State-of-the-Art Math Reasoning Surpassing GPT-4o
By
–
5). High-Level Automated Reasoning – extends in-context learning through high-level automated reasoning; achieves state-of-the-art accuracy (79.6%) on the MATH benchmark with Qwen2.5-7B-Instruct, surpassing GPT-4o (76.6%) and Claude 3.5 (71.1%).
