Give me long context benches for decode and prefill
@theahmadosman
-

Lab will take care of small specialized models as future
By
–

My lab will take care of this I had a stream in October 2024 saying small and specialized models are the future that I need to find
-
Nemotron 3 Ultra among top 5 open-source AI models
By
–
I now rank Nemotron 3 Ultra among the top 5 Opensource models out there Frontier intelligence at home
-
Love stumbling upon tears advising to buy GPU for local models
By
–
I just love randomly stumbling upon tears that say “Buy a GPU and run your own models locally! “
-

Thanos leaves only Qwen3.5-2B: Ahmad’s hypothetical scenario
By
–
Fun LLM Question from Mike today > Let’s say Thanos snaps away every model from existence except Qwen3.5-2B > Ahmad, what would you do? Assumptions & Clarifications 1. All papers, tech reports, synthetic datasets, HF repos, checkpoints, eval harnesses, and research artifacts
-
Harvey Legal AI model launches, confirms 2024 prediction about future
By
–
Enterprise focused, Harvey Legal AI model just came out last week for example
— Ahmad (@TheAhmadOsman) 7 juin 2026
I had a stream in October 2024 saying this is the future, and other times as well like 👇https://t.co/NRg9QzsBvkEnterprise focused, Harvey Legal AI model just came out last week for example I had a stream in October 2024 saying this is the future, and other times as well like
-

Prediction: unnerfed Opus 4.5 quality, NVFP4 optimized fits 20GB full context
By
–
I am actually sticking by prediction that we’ll have the unnerfed Opus 4.5 quality not the current shit one And NVFP4 optimized of that model would fit on 20GB with full context
-

Open source AI is the definite future, doubling down again
By
–
I got a lot of backlash on this post in January but I knew where we were heading with full conviction so I doubled down like I usually do I am doubling down again and saying that unless things go catastrophically wrong, Opensource AI is the definite future Cheers
-

LLM Decoding Simplified – upcoming article
By
–
LLM Decoding Simplified From the upcoming article on X
-
LLM inference: mostly about avoiding unnecessary movements
By
–
LLM inference is mostly about avoiding unnecessary movements btw