The tokenizer is an architectural prior disguised as preprocessing. And almost everyone has been treating it like plumbing. A new paper by Jan Tempus, Philip Whittington, Craig W. Schmidt, Dennis Komm, and Tiago Pimentel changes the frame: Tokenisation via Convex Relaxations
MACHINE LEARNING
-

LLM vs RAG vs AI Agent vs MCP Comparison
By
–
#LLM vs. RAG vs. #AIAgent vs. MCP
by @Python_Dv #GenerativeAI #ArtificialIntelligence #MachineLearning #ML -

Tokenization as Architectural Prior: New Research Paper
By
–
The tokenizer is an architectural prior disguised as preprocessing. And almost everyone has been treating it like plumbing. A new paper by Jan Tempus, Philip Whittington, Craig W. Schmidt, Dennis Komm, and Tiago Pimentel changes the frame: Tokenisation via Convex Relaxations
-
Lightning Indexer DSA Top-k Inference Path Demo
By
–
It’s just demoing the inference-time DSA top-k selection path at this point. The Lightning Indexer has trainable parameters, but this version does not train them (yet)
-
Validating Data Integrity for AI Models
By
–
How do you validate the integrity of data feeding your AI models?
-

DeepSeek Sparse Attention Implementation in LLMs Repository
By
–
Added a DeepSeek Sparse Attention (DSA) from-scratch implementation to my LLMs-from-scratch repo thanks to an awesome new reader contrib. With motivation, overview, and GPT-style model reference implementation as standalone example code: https://
github.com/rasbt/LLMs-fro
m-scratch/tree/main/ch04/09_dsa
… -

OpenAI reasoning model reportedly solves decades-old math problem
By
–
80 years. Every top mathematician alive tried this problem. Nobody solved it. An internal OpenAI reasoning model did it in a single attempt. 9 of the world's top mathematicians verified the proof. A Fields Medalist said he'd recommend it for publication "without any
-

A Quick Cheat Sheet to Master Agentic AI
By
–
A Quick Cheat Sheet to Master #AgenticAI
by @genamind #GenAI #LLM #ArtificialIntelligence #ML #MachineLearning #GenerativeAI -
Model tries half context to avoid incompatible downloads
By
–
Yeah and it tries half context if nothing fits at full. Saves you from downloading models that won't work
-
AI in education: human versus statistical approximation debate
By
–
The recent NYT opinion piece by @tab_delete is so fascinating, beyond just the way AI is upending education. "What really is the distinction between a nameless, faceless human you’ll never meet except over the internet and a statistical approximation of the same thing?".
