Anyone got a working recipe for running the 128,000 token variant on a Mac?
HARDWARE
-

NVIDIA Hopper leads generative AI performance in MLPerf benchmarks
By
–
NVIDIA Hopper takes lead in Generative AI on MLPerf! See https://
nvda.ws/3J1knIW In the latest (4th) round of #MLPerf performance benchmarking – the 'gold standard' for #AI workload #testing – the formidable Llama 2 70B and Stable Diffusion XL are center stage -
70B Model Context Window Limitations and Hardware Requirements
By
–
Ooof. I don't know if 128K is possible on the 70B, unless someone can get me a rig with like 16 or more A/H100s. 32K is likely doable though.
-
Running AI Applications on Budget Laptops: RAM Requirements
By
–
Presumably this one but I'm surprised it runs at all on a cheap laptop, @damnkittyworks how much RAM do you have?
-

USB Key Hardware Device Generates Strong User Demand Interest
By
–
i'd pay $$ to be shipped this USB key right now, to be honest
-

Running Llama-3-70B-Instruct locally on M4 Mac
By
–
Pleased to report that the Llamafile release of Llama-3-70B-Instruct works on my 64GB M4 – I had to download the 37GB file, quit a whole bunch of apps to free up RAM, then chmod 755 it, run ./Meta-Llama-3-70B-Instruct.Q4_0.llamafile and visit it on port 8080 on localhost
-

Meta-Llama-3-70B runs on M2 Mac with 64GB RAM
By
–
OK, can confirm that it works on an M2 with 64GB of RAM I had to quit Firefox and VS Code to free up enough RAM for it to run – Meta-Llama-3-70B-Instruct.Q4_0.llamafile now uses 38GB of RAM and runs at about 7.5 tokens a second
-
Meta’s AR Assistant and Neural Interface Dataset Development
By
–
Also this dataset reminds me a lot of what Meta is doing to build a real world AR assistant and neural interfaces https://t.co/LX9O7bWH5E
— Bilawal Sidhu (@bilawalsidhu) 21 avril 2024Also this dataset reminds me a lot of what Meta is doing to build a real world AR assistant and neural interfaces
-

Edge AI Adoption: Real-Time Decision-Making Through Model Optimization
By
–
The Promise of Edge AI and Approaches for Effective Adoption: Organizations are adopting edge AI for real-time decision-making using efficient and cost-effective methods such as model quantization, multimodal databases, and distributed inferencing. https://
kdnuggets.com/the-promise-of
-edge-ai-and-approaches-for-effective-adoption?utm_source=dlvr.it&utm_medium=twitter&utm_campaign=the-promise-of-edge-ai-and-approaches-for-effective-adoption
…