Okay I did a first quick pass of naive CUDA kernels for the forward pass of GPT-2 and pushed everything to one file in llm.c, Still only ~1000 lines of code: https://
github.com/karpathy/llm.c
/blob/master/train_gpt2.cu
… Current per iteration timings on my Lambda box <3 A100 40GB PCIe, B=4, T=1024:
– llm.c: 111ms
–
COMPUTING
-

CUDA kernels for GPT-2 forward pass implementation in llm.c
By
–
-
Weight-Only Quantization and Security Wrapper Integration for LLMs
By
–
PenghuiCheng for integrating weight-only quantization to accelerate transformer-based models on @Intel platforms
https://
python.langchain.com/docs/integrati
ons/llms/weight_only_quantization/
… And JamsheedMistri for adding the [
@layerup_
](
https://
x.com/layerup_) security wrapper for LLMs
https://
python.langchain.com/docs/integrati
ons/llms/layerup_security/
… (5/n) -
Meta Aligns PyTorch Compiler Flow on Custom AI Hardware
By
–
+1 to same compiler flow on our chips as mainline PyTorch. At Meta, we don't have a separate copy of PyTorch or Dynamo/Inductor. The source of truth is Github and everything is upstreamed/mainline. Triton's bugs have been going over time, but we liked it as a starting point, and
-
Unified Software Stack Approach for AI Inference Optimization
By
–
we are truly playing out the "uniform software stack" story. As long as you build a triton backend, and you drive down driver/firmware bugs, that is all that's needed. 99% of the user-level papercuts are taken care of.
It also helps that these are targeted for Inference right -
Triton Backend Enables AMD Code Generation Support
By
–
Triton can generate anything, as long as you write a Triton backend. As of today, upstream Triton can generate AMD code too.
-
Meta’s Chip Strategy: Not a Vendor Business Model
By
–
Meta is not in the chip-selling or chip-renting business
-

Advanced AI Techniques for Route Optimization and Delivery
By
–
Discover advanced AI techniques to optimize routes for deliveries, pickups, dispatching jobs, and more that can significantly save time, resources, and money. Watch this on-demand hands-on lab > https://
nvda.ws/3TPXlKo #cuOpt #RouteOptimization #GTC24 -
Simplified Management of Serverless and Dedicated LLM Deployments
By
–
ICYI: Managing #serverelss and dedicated #LLM deployments has never been easier! 🚀
— Predibase by Rubrik (@predibase) 10 avril 2024
😎 see all #deployments and status in one place
🆕 create new dedicated deployments with just a few clicks
🛠️ select the right #GPU for the job and customize autoscalinghttps://t.co/W9hCTKaySr pic.twitter.com/zImeI45nNFICYI: Managing #serverelss and dedicated #LLM deployments has never been easier! see all #deployments and status in one place create new dedicated deployments with just a few clicks select the right #GPU for the job and customize autoscaling https://
pbase.ai/449UUXT -
Meta Launches MTIAv2 Inference Chip for AI Acceleration
By
–
Meta announces 2nd-gen inference chip MTIAv2.
* 708TF/s Int8 / 353TF/s BF16
* 256MB SRAM, 128GB memory
* 90W TDP. 24 chips per node, 3 nodes per rack.
* standard PyTorch stack (Dynamo, Inductor, Triton) for flexibility Fabbed on TSMC's 5nm process, its fully programmable via the -

Intel Modernizes Future Networks with Ecosystem Partners at MWC24
By
–
At #MWC24 @intel showed how they are working alongside their world-class ecosystem to modernize and monetize the #network of the future, today. @IntelEdge #AI #IoT #IntelAmbassador @pierrepinna @Hal_Good @gvalan @enilev @Analytics_699 @AlexMachicado