From the replies, some people appear to think that LLMs read books in order. They don’t; they take interleaves batches samples from the entire training set.
LLMS
-

Code as Agentic Actions for LLMs
By
–
This is a really important part, thank you for highlighting this @sloppenheimer
! Code (indeed many other languages than python could work) is just the better, overlooked version of writing agentic actions for LLMs -
Tracking Per-Token Loss to Measure Document Contribution in LLM Training
By
–
It would be interesting if LLM training tracked the per-token loss back to the source material — it would be an objective measure of how much each specific book / document contributed to the training. Might say something useful for human learning!
-
QwQ 32B Draft Model Usage Discussion
By
–
I wonder if it’s more geared toward being used as a draft model for QwQ 32B
-
Qwen2.5 7B vs Ministral 8B: Detailed Model Testing and Comparison
By
–
right now qwen2.5 7b or ministral 8b. I’m running a more detailed test right now actually
-
Discussion of image generator and LLM prompt behavior
By
–
The image generator and the LLM are two models and the text shown in the first reply isn’t what it’s using as the prompt (which is probably just “portrait of Gary Marcus” or similar). Also note it doesn’t draw “Gary” or “Marcus” as majority black, but it does for “Marcus Gary.”
-

Grok generating hybrid likenesses
By
–



I think I figured this out: Grok is drawing a hybrid of Gary Marcus and Marcus Garvey
-

Deep Learning Job Interview Questions Fully Solved
By
–
Deep Learning job interview questions, fully solved & covering a wide range of key AI topics: https://
bit.ly/4bv6FL9 credit: @papers_daily
