Quick question, what's your expected minimum token per second when using a local LLM?
LLMS
-
AI Prompt Development Requires Expertise and Extensive Iteration
By
–
In all fairness, this prompt was like a 60+ hour project that required me to also have some expertise in a couple of other fields to nail it over at least 25 iterations. I agree there will be a shift, but I don't think, even with advances in AI in the next 1-2 years, that this
-
MPC as API Bridge Between LLMs and Software Tools
By
–
Yes thats the beauty of MPC, its like an API between the LLM and Tools (apps, services, etc) This is the future of how we will interact with software
-
Sesame Team Releases CSM-1B Model on Hugging Face
By
–
Amazing work by the @sesame team Try it out on hugging face, BURN THEM GPU's https://
huggingface.co/spaces/sesame/
csm-1b
… -
Sesame Labs Releases Advanced Open Source Conversational Speech Model
By
–
I'm speechless Sesame Labs has done the seemingly impossible with their new CSM (conversational speech model) → 1M hours of training data → ULTRA FAST → Voice cloning / watermarking
→ Apache 2.0 licenced
→ Based on llama & mimi Link below -

Chitu Matches vLLM Performance on H20 DeepSeek-R1
By
–
On an H20 (96GB) cluster, Chitu performs comparably to vLLM when deploying DeepSeek-R1-671B.
-
Development Team Optimizes GeMM and MoE for FP8 Processing
By
–
A member of the development team told us: "We've optimized a series of key operators, such as GeMM and MoE, at the instruction level to achieve native processing capabilities for FP8 data."
-

DeepSeek-R1-671B Outperforms vLLM on A800 Cluster
By
–
And it outperforms vLLM when deploying DeepSeek-R1-671B on an A800 (40GB) cluster.
-
Chitu: High-Performance Open-Source LLM Inference Framework
By
–
Yet another open-source project from China: Chitu! A high-performance inference framework for LLMs, designed for efficiency, flexibility, and availability. Chitu supports various mainstream models, including DeepSeek, the LLaMA series, Mixtral, and more.
-
Einstein’s Quote Questions Language-Only Approach to AGI
By
–
Einstein: "The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought." Are we really on the right path to AGI if we base it solely on language?