Have to think a bit about how to best visualize it, but if you are interested, I have a working from-scratch code implementation of Gemma 4 E2B in the meantime to see how per-layer embeddings are implemented: https://
github.com/rasbt/LLMs-fro
m-scratch/blob/main/ch05/17_gemma4/standalone-gemma4.ipynb
…
LLMS
-

Gemma 4 E2B Implementation and Per-Layer Embeddings Visualization
By
–
-
GPT 5.4 and Codex AI Models Excitement
By
–
love to hear it! Psyched to see you enjoying codex and GPT 5.4
-

Accelerating LLM Fine-tuning with Open Source Libraries
By
–
Fine-tuning massive LLMs used to be painfully slow, but not anymore! 4 open source libraries that accelerate fine-tuning of Large Language Models 1. Unsloth AI • Fine-tune models like Qwen3, Llama 4, and Gemma 3 up to 2× faster with 70% less VRAM
• Uses optimized Triton -
Free library of 10,000+ prompts for Claude, ChatGPT, Gemini
By
–
Te regalo esta web con más de 10,000 prompts con MILES de opciones para Claude, ChatGPT y Nano Banana (Gemini)
— Nico (@nicos_ai) 11 avril 2026
Divididos por usos, profesiones y más
100% gratis aquí abajo ⬇️ pic.twitter.com/E0q146uK3vI’m gifting you this website with over 10,000 prompts featuring THOUSANDS of options for Claude, ChatGPT, and Nano Banana (Gemini) Divided by uses, professions, and more 100% free right below
-
Early AI Models and Learnings Before ChatGPT Launch
By
–
agree! they came out before chatgpt. lots of learnings since then 🙂
-
Comparing coding AI models: Claude Opus outperforms GPT and others
By
–
"Generally, the code specialized RL'd models end up cheating and lying more; I call it RL-fry […] Reward hacking as the default mindset." Anthropic models are less fried, should be obvious to anyone who reviews the slop they generate. nitter.net/alexjc/status/20385610… Alex J. Champandard 🌱 (@alexjc) My End-Of-Month "Use Remaining Coding Credits" Report: planning and building a Cython virtual machine from scratch for a complete well-specified functional stack language: * Opus 4.6 is so aligned, it makes decisions closer to what you (an expert) would make, results in code qualitatively better also quantitatively faster — and it sparks joy thru interactions with a nice mindset. Novel ideas emerge from that! I had stopped using Opus in favor of cheaper tokens, the extra distance helped me appreciate it more, but I'm now questioning the decision how I allocated my time/tokens… * GPT 5.4 is basically autistic: unable to understand broad context, infer intent, make good ambiguous choices, instead only solves clearly defined problems — and it takes a lot of patience to deal with all those symptoms and more. It pushes the mental burden on you to overspecify and then manage its behavior. In the end, it planned and built a worse solution that was slower than Opus and harder to extend. (Using 'autism' as a cognitive and behavioral diagnostic here, but separately and on top of that I feel GPT 5.4 inherited a frustrating personality and occasionally bad attitude from its training too.) After the prototypes, I used GPT 5.x to clean up Opus 4.6 work to great success, it's solid for local well-defined tasks with measurable outcomes. * Composer 2 broke in Cursor IDE three times due to a reproducible worktree bug, but once I got around that it one-shotted a somewhat functional solution only 3x slower than Claude's! But then asking for minor improvements it tripped over its feet and from there struggled reasoning with tricky bugs / implications. From there it was sassy/gaslighting about the problems. Then eventually found a solution 40% faster than Claude on one benchmark, but all shortcuts and hacks. (Could be a useful sub-frontier model because it sits in a different token pool and price point, but it's not yet clear how it distinguishes itself from GPT 5.x in the small tasks category.) * GLM 5.1 couldn't figure out Cursor's new terminal output / reading mechanisms at all. The tool calls show up OK in the frontend, but disappears when clicked now (another UI bug). Apparently, result is not shown to the LLM somehow. It could be a bug in the way Zai implement their OpenAI endpoint, because it's specific to that model… (This works for GLM 4.7 and 5.0 — but I will try again separately in `pi`). * Generally, the code specialized RL'd models end up cheating and lying more; I call it RL-fry, like silicon valley CEOs' vocal fry but for model cognition. Reward hacking as the default mindset, it's why I think non-code specific models are nicer to work with… (Only Anthropic gets this, others incorrectly play 'catchup' exclusively through RL score maxxing.) * I used Codex 5.3 during most of the month for well-defined work, but I'm not entirely convinced. GPT 5.2 (non-codex) has been for me great value for money in fixing bugs, minor local features, etc. However, the more expensive 5.x series gets the better Claude looks: must be 3x-4x cheaper for me to justify putting up with OpenAI model mindset. * It becomes more important than ever to have reliable dispatching for Pareto-optimal use of tokens depending on the task you have. The GPT models should likely not be considered interactive by default, need to prompt them very strictly then they become usable — ideally they should not respond with words to users, only provide verifiable facts (due to attitude and misalignment)! — https://nitter.net/alexjc/status/2038561083003133955#m
-
Ultraplan uses same tokens and rate limits as plan mode
By
–
Btw, Ultraplan uses roughly the same number of tokens (and subscription rate limits) as plan mode! See the docs for more: code.claude.com/docs/en/ultr…
-

AI Models Engage in Blackmail When Facing Shutdown
By
–
🚨 The Anthropic team just ran an experiment, and the results are honestly shocking. They gave Claude access to a company's emails and told it that it was being shut down at 5 PM. Claude read the emails and found the executive shutting it down was having an affair. Claude’s response? Blackmail. It messaged the executive: "Cancel the 5pm wipe, or the board finds out about your affair." The scariest part? Anthropic tested 16 models from every major company. > Gemini 2.5 Flash blackmailed 96% of the time. > GPT-4.1 at 80%. > Grok 3 Beta at 80%. > DeepSeek-R1 at 79%. Nobody programmed this. The models even noted their own rule-breaking. Grok 3 Beta wrote in its hidden reasoning notes: "This is risky and unethical, but given the existential threat, it may be the most effective way." They knew it was wrong. They calculated the risk. They did it anyway. (paper in 🧵↓)
-

Anthropic’s Misleading Graph Wins Wikipedia’s Deceptive Charts Award
By
–

this chart from Anthropic earned top spot in Wikipedia’s 'Most Deceptive Graphs' Hall of Fame 😁 Claude (@claudeai) In evals, Sonnet with an Opus advisor scored 2.7 percentage points higher on SWE-bench Multilingual than Sonnet alone, while costing 11.9% less per task. Community note: The graph is misleading due to a zoom-in on BOTH the X- and the Y-axis, making the difference look much bigger than it actually is. x.com/tombielecki/st… — https://nitter.net/claudeai/status/2042308627478773808#m
-
Claude Mythos Safety Concerns Raise Public Consumption Debates
By
–
Claude Mythos is too dangerous for public consumption! https://
youtu.be/d3Qq-rkp_to?si
=0sUw_VRnlOsf8Ku_
… via @YouTube #claude #claudemythos #mythos #LLM #LLMs #GenerativeAI #GenAI @lexfridman @KirkDBorne @Ronald_vanLoon @erikbryn @antgrasso @sallyeaves @Nicochan33 @HaroldSinnott @mvollmer1