Note: we’re aware of crashes and navigation issues by some users and are actively fixing them. Please update to the latest version of MacOS for the best experience.
SOFTWARE
-
Perplexity launches macOS app with keyboard shortcut access
By
–
Perplexity is now on MacOS. Ask anything with ⌘ + ⇧ + P.
— Perplexity (@perplexity_ai) 24 octobre 2024
Download now: https://t.co/Z9UF7og614 pic.twitter.com/eFAJfAVPZsPerplexity is now on MacOS. Ask anything with ⌘ + ⇧ + P. Download now: http://
pplx.ai/mac -
Lutra AI Extends Support for Excel Files and Large CSVs
By
–
Love it! We've extended this to also support Excel files, large CSVs, and more – check out @Lutra_AI
-
torch.compile CUDAGraph overhead optimization reinforcement learning
By
–
there's more CPU overhead, but it practically doesn't matter IMO unless you are doing tiny reinforcement-like workloads — and even there `torch.compile` will CUDAGraph it to near-zero overhead.
-
Claude Adds Data Analysis Feature Preview
By
–
BREAKING 🚨: Claude got Data Analysis as a feature preview.
— 🚨 AI News | TestingCatalog (@testingcatalog) 24 octobre 2024
With this feature, it can execute the code to run complex data visualisations and render them via Artifacts. https://t.co/99xovpNVqO pic.twitter.com/yxeHS3BgudBREAKING : Claude got Data Analysis as a feature preview. With this feature, it can execute the code to run complex data visualisations and render them via Artifacts.
-

Two Quantization Techniques for Model Optimization Released
By
–
We used two different techniques for quantizing these models.
Quantization-Aware Training with LoRA adaptorsprioritizing accuracy.
SpinQuant, a post-training quantization method which prioritizes portability.
Both versions are available for download as part of this release. -

Llama 3.2 Quantized Versions Speed Up Inference 2-4x
By
–
We want to make it easier for more people to build with Llama — so today we’re releasing new quantized versions of Llama 3.2 1B & 3B that deliver up to 2-4x increases in inference speed and, on average, 56% reduction in model size, and 41% reduction in memory footprint.
Details -
Cerebras Powers Instant Inference for Major AI Applications
By
–
Numerous companies are using Cerebras to make their inference run at instant speed. These include:
– @GSK for drug discovery
– @Livekit for voice AI
– @tavus for digital twins
– @vellum_ai for testing & iteration -
Cerebras Announces 3x Faster Inference Speed Update
By
–
After this huge speed update we will be focusing on supporting additional customer models, context, and capacity. Stay tuned for more updates!
Chat: http://
Inference.cerebras.ai
API key: http://
cloud.cerebras.ai
Blog: https://
cerebras.ai/blog/cerebras-
inference-3x-faster
… -
Cerebras Inference 3x Faster: Llama 70B Reaches 2,100 Tokens/Second
By
–
🚨 Cerebras Inference is now 3x faster:
— Cerebras (@cerebras) 24 octobre 2024
Llama3.1-70B just broke 2,100 tokens/s
– 16x faster than the fastest GPU solution
– 8x faster than GPUs running Llama *3B*
– It's like the perf of a new hardware generation in a single software release
Available now at… pic.twitter.com/9VgGWGO6qYCerebras Inference is now 3x faster: Llama3.1-70B just broke 2,100 tokens/s
– 16x faster than the fastest GPU solution
– 8x faster than GPUs running Llama *3B*
– It's like the perf of a new hardware generation in a single software release
Available now at