running Qwen3.5 397B MoE (17B active/token)
— Ahmad (@TheAhmadOsman) 9 avril 2026
on 4x DGX Sparks in FP8 (~400GB)
> OpenCode driving
> agent exploring its own config
> probing all 4 Sparks (via ssh) + reporting thermals
> inspecting how vLLM is serving it
> collecting + analyzing its own stats
local AI is awesome https://t.co/KU9u30GgXk pic.twitter.com/yPWSbSKto8
running Qwen3.5 397B MoE (17B active/token) on 4x DGX Sparks in FP8 (~400GB) > OpenCode driving
> agent exploring its own config
> probing all 4 Sparks (via ssh) + reporting thermals
> inspecting how vLLM is serving it
> collecting + analyzing its own stats local AI is awesome