honestly find it a bit troubling that for a lot of models we don't quite know what the actual System prompt is.. I'd advocate for API providers for making them public.
@reach_vb
-

Gemini 2.0 Flash Introduces Full CoT Traces
By
–
Gemini 2.0 Flash Thinking: Open AI has been Noam-ed! 👑
— Vaibhav (VB) Srivastav (@reach_vb) 19 décembre 2024
with FULL CoT Traces! pic.twitter.com/7ngPko3ZqvGemini 2.0 Flash Thinking: Open AI has been Noam-ed! with FULL CoT Traces!
-
ChatGPT Free Alternative Raises Pricing Concerns
By
–
oofff steep price, chatgpt can do it for free..
-

ModernBERT: New Open-Source Language Model from Answer.ai
By
–
ModernBERT ftw! @answerdotai & @LightOnIO killing it!! > ModernBERT-base: 22 layers, 149M params
> ModernBERT-large: 28 layers, 395M params > 2 trillion tokens of English and code data.
> Up to 8,192 tokens, ideal for processing long documents
> RoPE for long-context support -

New Hardware Gains Enable Faster LLM and VLM Deployment
By
–
This is going to be so, so fun to plug LLMs & VLMs with
— Vaibhav (VB) Srivastav (@reach_vb) 17 décembre 2024
> 67 INT8 TOPS (1.7x increase)
> 102GB/s memory bandwidth (2x increase)
pic.twitter.com/ByRYOjdLcoThis is going to be so, so fun to plug LLMs & VLMs with > 67 INT8 TOPS (1.7x increase)
> 102GB/s memory bandwidth (2x increase) -

Mosaic ML Achieves Major Milestone Recognition
By
–
Massive congrats to Mosaic ML team! – you deserve this and more!
-

Falcon 3 Language Models Released: 1B to 10B Parameters
By
–
Falcon 3 is out! 1B, 3B, 7B, 10B (Base + Instruct) & 7B Mamba, trained on 14 Trillion tokens and apache 2.0 licensed! > 1B-Base surpasses SmolLM2-1.7B and matches gemma-2-2b
> 3B-Base outperforms larger models like Llama-3.1-8B and Minitron-4B-Base
> 7B-Base is on par with
