2/5 The setup: JLMs are used only during the reward phase of Online RL, sitting idle the rest of each step. To avoid GPU waste, we shared JLM deployments across multiple training jobs. This improved utilization but exposed JLMs to unpredictable traffic bursts from multiple
AI HARDWARE
-

Scaling vLLM: Doubling Throughput and Halving Latency
By
–
1/5 Go Big or Go OOM: The Art of Scaling vLLM .
We doubled throughput and cut latency in half-same GPUs, just better vLLM config then added smart autoscaling to handle traffic bursts. Here's what we learned optimizing LLM-as-a-Judge for GRPO training. -

Boston Dynamics Atlas demonstrates rapid progress in embodied AI
By
–
Boston Dynamics just dropped a new Atlas video and it’s a reminder of how fast embodied AI is progressing. What stands out isn’t just agility or balance, but full-body control under real-world constraints. Every movement reflects years of work in perception, coordination and
-
MRI-Guided Cryoablation Destroys Spinal Cancer Tumor Without Surgery
By
–
😳🩻🤷🏽♀️ A spinal cancer tumour. Frozen solid. Gone the same day.
— Catherine Adenle (@CatherineAdenle) 8 février 2026
According to ABC News Australia, doctors at Liverpool Hospital used Australia’s first MRI-guided cryoablation to destroy a cancerous spinal tumour without open surgery.
Here’s the wild part 👇
➡️ A needle guided… pic.twitter.com/VvYjXk7W4CA spinal cancer tumour. Frozen solid. Gone the same day. According to ABC News Australia, doctors at Liverpool Hospital used Australia’s first MRI-guided cryoablation to destroy a cancerous spinal tumour without open surgery. Here’s the wild part A needle guided
-
Codex not recognizing CUDA on user’s machines
By
–
I still can’t get codex to recognize that CUDA exists on my machines.
-
256 Tb/s Fiber Optic Data Rates Demonstrated Over 200 km
By
–
256 Tb/s data rates over 200 km distance have been demonstrated on single mode fiber optic, which works out to 32 GB of data in flight, “stored” in the fiber, with 32 TB/s bandwidth. Neural network inference and training can have deterministic weight reference patterns, so it is
-
NVIDIA ecosystem vs Mac or 5090 for large models
By
–
If you want ALL of the native support for the entire NVIDIA software ecosystem, AND you anticipate working with models that are 100 GB or larger, then it is a pretty solid choice. Otherwise a decently speced Mac would work, or a workstation with a 5090.
-
FP16 vs INT8: Comparing Model Quantization Trade-offs
By
–
FP16 vs. INT8: Speed vs. Efficiency ⚡
— Satya Mallick (@LearnOpenCV) 5 février 2026
Both make models faster, but the choice depends on your hardware. 🛠️
💎 FP16 (Half Precision): The "safe" bet. Fast on GPUs, retains high accuracy, and requires almost no extra work.
🔋 INT8 (8-Bit Integer): The "efficiency" king. Uses… pic.twitter.com/Y4wgZ4gT2vFP16 vs. INT8: Speed vs. Efficiency Both make models faster, but the choice depends on your hardware. FP16 (Half Precision): The "safe" bet. Fast on GPUs, retains high accuracy, and requires almost no extra work. INT8 (8-Bit Integer): The "efficiency" king. Uses
-
vLLM Kernel Optimizations Boost GB200 Inference Performance
By
–
Impressive deep dive! It’s great to see the vLLM team maximizing the GB200’s potential. These kinds of kernel-level optimizations are exactly why the PyTorch ecosystem continues to be the foundation for next-gen inference performance.
-

Yocto BSP Layer Enables Custom Linux for Axelera Metis Accelerators
By
–
Building custom embedded Linux images with Yocto? There's now an official BSP layer for Axelera Metis hardware. meta-axelera includes recipes for the PCIe driver and udev rules, so your custom Linux will recognize Metis accelerators straight out of the build. Supports Scarthgap,
