Thinking token support coming soon.
@localghost
-
Tiny 1.5B Model Outperforms GPT-4o, Runs Offline
By
–
“outperforms GPT-4o and Claude-3.5-Sonnet on math benchmarks”
— Aaron Ng (@localghost) 22 janvier 2025
Here it is running on a phone, completely offline.
Tiny 1.5B AI model that will run on almost any hardware, even without internet.
Are you getting it yet? https://t.co/RjXTKvsRPd pic.twitter.com/6GimUSBKOo“outperforms GPT-4o and Claude-3.5-Sonnet on math benchmarks” Here it is running on a phone, completely offline. Tiny 1.5B AI model that will run on almost any hardware, even without internet. Are you getting it yet?
-
R1 Raw CoT Update Replaces o1-Style Thinking UI
By
–
This update doesn’t support the o1-style thinking UI btw For now you can enjoy r1’s raw CoT
-
Apollo 1.0.23 Brings Compact DeepSeek r1 Models to Mobile
By
–
Apollo 1.0.23 is In Review and has the smaller DeepSeek r1 models ready to try. I suspect these will be the best models you can run on a phone.
-
o1 Canvas Feature Preferred Over Claude 4o
By
–
o1 with canvas would be nice too. can’t bring myself to use 4o anymore
-
Q2 Model Performance Limited by RAM Requirements
By
–
q2 yes. But needs more ram right now for any higher than that
-
o1 Pro Speed Trade-offs: Quality vs Waiting Time
By
–
feeling the same way. o1 pro is great, but not always worth waiting 10x longer
-
Microsoft Phi-4 Streaming: Powerful AI on Personal Devices
By
–
Here’s Microsoft’s Phi-4 streaming from my MacBook to my iPhone at 26 tok/s.
— Aaron Ng (@localghost) 18 janvier 2025
In 2025 you can serve powerful AIs to your whole network with just a few apps. Offline, free, and private. pic.twitter.com/AskeoW0FdQHere’s Microsoft’s Phi-4 streaming from my MacBook to my iPhone at 26 tok/s. In 2025 you can serve powerful AIs to your whole network with just a few apps. Offline, free, and private.
-
4-bit Quantization: RAM Limitations on iPhone Devices
By
–
it's the 4bit quantized version. you won't be able to fit the biggest one on an iPhone yet due to RAM limits
-
Q2 AI Model Deployment Limitations Discussion
By
–
I plan on getting q2 on there but it's not likely going to be anywhere as smart