I've doubled LLaMA 3's context window to 16K tokens. Fully open-source. Link in thread:
@mattshumer_
-

LLaMA 3 Context Window Extended to 16K Tokens Open-Source
By
–
I've doubled LLaMA 3's context window to 16K tokens. Fully open-source. Link in thread:
-

LLaMA 3 Context Window Doubled to 16K Tokens
By
–
I've doubled LLaMA 3's context window to 16K tokens. Fully open-source. Link in thread:
-
Context Length Extension Testing and Performance Benchmarks
By
–
I haven't run real benchmarks on it (if anyone wants to, that'd be awesome!), but from my homemade haystack tests, the context length extension worked well.
-

LLaMA 3 Context Window Doubled to 16K Tokens
By
–
I've doubled LLaMA 3's context window to 16K tokens. Fully open-source. Link in thread:
-
Evaluation Loss in Model Training: Understanding Its Role
By
–
eval loss. i don't always use it, as it's not a great measure, but it's good for things like this
-
Faster AI Inference Enables More Thinking Same Time
By
–
The model will output the same result, no matter the speed. What faster inference enables is more thinking in the same amount of time.
-

Autocomplete for Design: Revolutionary AI Innovation Arrives
By
–
This is utterly insane.
— Matt Shumer (@mattshumer_) 22 avril 2024
Autocomplete for design is here. https://t.co/SVbJ6JbtM4This is utterly insane. Autocomplete for design is here.
-
Getting into AI: From GPT-2 fine-tuning to ChatGPT
By
–
This was probably around late 2020, 2021. Before that I was doing a bit of fine-tuning over GPT-2, but chargpt was when I really got into it.
-
Groq’s Time to First Token Performance Needs Improvement
By
–
If Groq gets their time to first token down a bit, probably not. They still take maybe a half second to a second to get you that first token.
