My favorite part about @honicky
's Paper Club session this week on the 1-bit LLMs paper – relating it to @jefrankle
's Beyond Chinchilla laws and adjusting the equations for the memory/latency characteristics of 1-bit LLMs to derive an optimal param count/data size to aim for. no
@swyx
-

1-bit LLMs: Optimizing Parameters and Data with Chinchilla Laws
By
–
-
Devin Coding Agent: Overclaim Concerns but Genuine Advanced Capabilities
By
–
i do think @lulumeservey may want to damage control this one. it was an obvious overclaim that lost a lot of fidelity in the final message. but still does not diminish the core utility that devin provides in that it really is the most advanced coding agent ive ever seen. not
-
Suggestion to try Gemini with Twitter integration magic
By
–
throw it into gemini maybe 🙂 cc @altryne who has some twitter ripping magic
-

LMSYS Performance Gap Widens Between SOTA and Other Models
By
–
the lmsys gap between sota and the rest is enormous and just widening
-
Devin vs Open Devins Live Stream – Submit Questions Now
By
–
for those on youtube we are streaming Devin vs all the open devins now, submit question pls
-
AI Influencers Exaggerating Minor Language Model Issues
By
–
tbf i was only joking about it but theres a very real set of ai influencoooors who build their identity around making mountains out of language model esoterica molehills
-

GPT4T Reclaims SOTA with Open Source Lightweight Evals
By
–
GPT4T retakes the SOTA LLM spot (probably) but also i like this way of announcing things: open source your evals, make it reproduceable and comparable. and **LIGHTWEIGHT**: eleuther’s harness and oai/evals always felt far too heavy of a lift and not flexible. this thing
-
InfiniAttention Paper Performance Review for Model Size
By
–
first thing i looked for in the infiniattention paper. theres limits ofc but they do pretty well for the size
-

Google’s Linear Attention Breakthrough: Griffin and Infini-Attention
By
–
with Griffin and Infini-attention, it increasingly feels like Google leapfrogged Together and RWKV in the race for scaling up linear attention, and they shared a watercooler conversation with Anthropic or something papers are trickling out now but it's hard to know what's
-
Emerging Consensus: 8x22B Model is Mixtral Medium
By
–
think emerging consensus is 8x22B -is- mixtral medium