Their AI-focused ARR tripled year over year to surpass $500 million. Very few enterprise software companies report this level of direct, paid AI adoption. The engine of this breakthrough is Firefly, which now reaches $300 million.
COMPUTING
-
Ollama slower, slop, code thieves; better alternatives listed
By
–
ollama > slower than llama.cpp on windows
> slower than mlx on mac
> slop useless wrapper
> literal code thieves alternatives? > lmstudio
> llama.cpp
> exllamav2/v3
> vllm
> sglang
> trt-llm literally anythingʼs better than ollama -

Why focus on inference engines: performance gains with vLLM and Sglang
By
–

Why do I focus on Inference Engines/Software Stacks for your hardware? – 2x RTX 3090s: ~14.5 tok/s → ~64 tok/s moving to vLLM w/ TP=2 – RTX PRO 6000: ~32 tok/s → ~110 tok/s moving to Sglang So: – CUDA/2+ GPUs: ExLlamaV3/vLLM/Sglang > llama.cpp – Edge: llama.cpp > Ollama
-
Software-brained approach limits code tools for knowledge work
By
–
A fundamental problem with extending Codex/Cowork/Code to all knowledge work is that they remain very "software-brained" where the end result (the software) is what is important & that code serves as a source of truth. For a lot of other knowledge work, the process is at least
-
Why AI Data Centres Use So Much Water and Its Water Crisis Impact
By
–
Why Do AI Data Centres Use So Much Water ? Is it creating a water crisis? https://
youtu.be/DpfffbzEcno?si
=6KZ-h4-T9VaTYgVo
… via @YouTube #AI #artificialintelligence #datacentre @lexfridman @KirkDBorne @Ronald_vanLoon @erikbryn @antgrasso @sallyeaves @Nicochan33 @HaroldSinnott @mvollmer1 @marcusborba -

Meta AI unveils Artifacts tab to store presentations and documents
By
–
Meta AI gets a new Artifacts tab on the web. All presentations, documents, web pages and other creations would be stored there. Bridging the gap.
-
Kurzweil’s AGI predictions mixed with flawed connectome immortality
By
–
Kurzweil mixed bold but solid predictions (given enough compute, a neocognitron like architecture can achieve AGI before 2030) with dogshit (digitizing the connectome is a near term way to human immortality). But both looked like the same kind of scifi to non experts.
-
Stop hardware cost to token calculations, models improve, prices rise
By
–
Can we stop doing hardware cost to token generation calculations on the timeline please? If you haven't noticed, models keep getting better & more efficient, and hardware prices keep going up
-
ColBERT outperforms on CPU with low latency for embeddings
By
–
I'm talking about individual descriptions used for embeddings. It doesn't need to be particularly long for late interaction to perform better. The tradeoff really depends on the use case. In this case, even on a cheap CPU, the latency is so low that ColBERT just works better!
-

Luke Alonso uploaded NVFP4 of GLM 5.2, 467GB on 4 DGX Sparks
By
–
Luke Alonso has uploaded an NVFP4 of GLM 5.2 467GB, would fit on 4x DGX Sparks (~$20k)