A team from MIT built a model that scores 61.9% on ARC-AGI-PUB using an 8B LLM plus Test-Time-Training (TTT). Previous record was 42%. (via Localllama)
LLMS
-
Teaching ChatGPT Pronunciation: A Reciprocal Learning Experience
By
–
https://
x.com/lgj170/status/
1857072165700473130/video/1
… ChatGPT not only teaches people, you can also act as a teacher to ChatGPT and teach the AI pronunciation 🙂 -

Google’s Benchmark Tricks and Performance Loss to Competitors
By
–
Looks like Google was using some tricks again to push its benchmarks on chatbot arena. However, it's loosing on LiveBench against Sonnet 3.5 and o1-preview. Somehow I am not surprised. I somehow didn't expect it to beat o1 and Sonnet either. It seems to be at least a little bit
-
ChatGPT as an Infinite Patient AI Tutor for Every Child
By
–
This is why ChatGPT is such a transformative technology. Teaching and learning will never be the same again. Every child born today has an AI partner who has infinite patience to answer the child's every question on every subject. The tutor is always there and never gets tired.
-

Anthropic as OpenAI’s Key Competitor Driving Innovation
By
–
Anthropic is the most important competitor we have to OpenAI. They keep pushing them to invent better and especially more efficient models.
-
MIT Lab Demonstrates Test-Time Training Boosts LLM Abstract Reasoning
By
–
MIT Lab publishes "The Surprising Effectiveness of Test-Time Training for Abstract Reasoning": Test-Time Training (TTT) produces a 61.9% score on the AGI-ARC benchmark. MIT researchers have demonstrated how Test-Time Training (TTT) significantly boosts large language models’
-
01.ai Trains #6 World Model for $3M with $0.14/M Token Inference
By
–
01.ai trained the #6 model in the world for $3M pre-train cost. And the inference price is $0.14/million tokens! tomshardware.com/tech-indust…
-

Yi Models Integration with CAMEL Framework for Multi-Agent AI
By
–

🎉Love seeing Yi models in @CamelAIOrg! Powerful Yi models join forces with this awesome multi-agent framework. Can't wait to see what AI agents you'll build! #YiLightning #LLM #AI CAMEL-AI.org (@CamelAIOrg) 📢 We've just added support for the Yi-series of LLM models in the 🐫 CAMEL framework! This enhancement allows users to leverage various performance tiers with models like yi-lightning, yi-large, yi-medium, and yi-large-turbo, providing greater flexibility in language processing tasks. Thanks to our contributor MuggleJinx for this significant contribution! 🤝 Explore more here: github.com/camel-ai/camel/pu…. — https://nitter.net/CamelAIOrg/status/1857118366730776719#m
-
ChatGPT least biased AI in evaluations, customization emphasized
By
–
we are proud of how consistently chatgpt scores as the least biased ai in evals. that is an important default (and then users should have lots of choice to customize).
-

01.ai Trains GPT-4 Competitor with 95% Fewer Resources
By
–
Chinese startup 01 .ai trains competitive LLM using 95% fewer resources through innovative engineering optimization. 01 .ai trained a GPT-4 competitor using just 2,000 GPUs and $3M, while achieving competitive performance. Through innovative engineering and optimization techniques, they achieved what OpenAI did with $80-100M, demonstrating remarkable cost efficiency in LLM training. → Training Resource Optimization at 01 .ai Using only 2,000 GPUs versus OpenAI's estimated 10,000+ GPUs for GPT-3. The company achieved competitive performance despite severe hardware constraints due to US regulations. → Cost Efficiency Breakthrough $3M total training cost compared to OpenAI's $80-100M for GPT-4. Model ranked sixth in performance according to UC Berkeley's LMSIS benchmark. → Technical Innovation in Inference Transformed computational problems into memory-oriented tasks. Built multi-layer caching system and specialized inference engine. Achieved inference costs of 10 cents per million tokens – 1/30th of industry standard. → Engineering Focus Areas Prioritized GPU resource allocation. Optimized both training speed and inference efficiency. Developed custom inference architecture for maximum hardware utilization.
