For that, you would need something more optimized for CoreML, but you could potentially run this on an iPad with M2? I don't have an iPad Pro, so I can't really try. I am also excited about the work of @argmax
; they are the ones to check out for device optimization!
COMPUTING
-
CoreML Optimization for iPad M2 Device Deployment
By
–
-

Phi-2-DPO-7K: Microsoft’s Fast Fine-Tuned Language Model
By
–
Introducing phi-2-dpo-7k! A Microsoft Phi-2, fine-tuned on a diverse cocktail of 7k chat interactions from @argilla_io
's latest DPO dataset, including orca pairs, ultra-feedback ratings, and capybara-dpo Runs ultra-fast on Apple Silicon, thanks to MLX -

Chiplets Summit: Disrupting AI Architecture for Compute Demands
By
–
Join @UntetherAI
's @beach66 at the #ChipletSummit next week! He'll give a deeper look at how #chiplets disrupt conventional architectures to enable innovative designs that satisfy #AI's compute demands. Seats are filling fast! Register now: https://
chipletsummit.com/registration/ -

SAS Code Spring Challenge: Repetitive Programming Task
By
–
Celebrate the predicted early Spring with some SAS code, but make sure you have a lot of time on your hands…and don't say we didn't warn you. #FeelsLikeWeveDoneThisBefore #GroundhogDay
-
Progression from GPU Poor to GPU Middle Class Status
By
–
Upgraded from “GPU Poor” to “GPU Middle Class”
-

Sakana AI Receives Japanese Government GPU Cluster Grant
By
–
We finally have our own GPU cluster! https://
sakana.ai/nedo-grant/ Excited announce that @SakanaAILabs is 1 of 7 institutions in Japan chosen by the Japanese government to receive the NEDO supercomputer grant, for developing foundation models to strengthen Japan’s AI ecosystem. -

Meta Deploys Artemis AI Chip to Reduce NVIDIA GPU Reliance
By
–
Meta released plans to deploy its 2nd-gen custom AI chip, codenamed "Artemis". The chip will be integrated into data centers as soon as this year to reduce reliance on costly NVIDIA GPUs. Meta is making big bets on AI.
-
Exponential Technology Acceleration: 15 Years Compressed to Months
By
–
The technology is exponential. What used to take 15-years will likely now take 15-months (or less).
-
National AI Research Resource NAIRR Launch Explained
By
–
This video aged well: Published in 2021, here we explained why we were calling for the creation of a national research cloud—now materialized as the National AI Research Resource or NAIRR.
-
Mixtral 8×7 Achieves 488 Tokens Per Second on Groq Hardware
By
–
wow, mixtral 8×7 can hit ~488 tok/s using groqchat's custom chips. realtime analysis or nested llm calls become way more feasible at these speeds
— Aaron Ng (@localghost) 1 février 2024
whole responses in about a second: pic.twitter.com/wv8jX4Uj93wow, mixtral 8×7 can hit ~488 tok/s using groqchat's custom chips. realtime analysis or nested llm calls become way more feasible at these speeds whole responses in about a second: