The code hasn't changed so there are padding tokens in there still… I could be using that padding though!
@alexjc
-
Lambda Labs GPU pricing comparison A100 availability
By
–
Thanks. I've been using Lambda Labs but they only had A100 when I checked! Seems the prices are similar…
-
TokenMonster vocabulary files released for AI models
By
–
I also uploaded the binary files for the tokens to this repository, which you can use as drop-in replacement. The TokenMonster vocabulary file is alongside the tokens and on GitHub too:
-

NanoGPT Achieves 40% Training Efficiency Boost Via Tokenizer
By
–
New Submission On NanoGPT Speed-Run! Improved training token efficiency by ~40% to reach target HellaSwag score by switching the vocabulary and tokenizer only. It now takes 1050 steps instead of 1750 steps; current record on 11/24/24 labelled in the graph as "TikToken
-
Government Influence on Intellectual Property Court Cases Revealed
By
–
Q: What makes you think there are policymakers or officials beyond the court involved in this? Basically, they told us! Already two years ago, policymakers involved with WIPO wrote with glee how governments will try to influence the court cases their way, and if that doesn't
-

2D Convolution Stride Optimization for Compute Efficiency
By
–
This analysis suggests that a 2D convolution filter of stride=2 would be the most compute-efficient for this particular workload… Also, ideally batch up all the grey data and do the calculation all at once.
-
Court Ruling Could Enable Technology Piracy Legal Concerns
By
–
It's a weird verdict, could be used to justify any kind of piracy…
-
Court Dismisses AI Copyright Case for Lack of Standing Evidence
By
–
Plaintiff could not provide evidence of injury-in-fact, they didn't have the quality ML forensics as NYT. The judge threw it out for lack of standing only. They could return with better plead case.
-
Triton vs Taichi: PyTorch Integration Framework Comparison
By
–
I went for Triton because it was already closely integrated into PyTorch. Taichi looks really polished in comparison, but not tried it yet!
-
GPU Kernel Memory Optimization Challenges in AI Development
By
–
I know what you mean. For me there was nothing between 1x and 70x, it either worked or it didn't — had to redesign it each time. Now running out of shared memory for the GPU kernel, just a tad too small to be useful!