Also, I don’t know how OP is getting Qwen3.5 27B @ 30 tps on the DGX Spark Doesn’t make sense for a Dense model on DGX Spark’s Unified Memory (273 Gbps) (My personal experiments showed it’s 4 tokens/sec) Very curious how you got that @TeksEdge
DGX Spark Qwen3 27B Inference Speed Discrepancy Questioned
By
–