Yeah that is definitely faster because you’re fully offloading to the GPU which runs at ~1800Gbps in comparison to DGX Spark’s 273Gbps
AI HARDWARE
-

New Website to Benchmark Hardware Tokens per Second Performance
By
–
Update: @AlpinDale and I agreed to collab on a website that does more accurate calculations for your hardware tokens/sec performance Expect more on this after I am done with GTC this week
-
Qwen 3.5 MoE Model Recommendations for Spark and Mac Studio
By
–
For a single Spark my first recommendation would be Qwen 3.5 122B MoE (Int4) For the Mac Studio I would recommend the 397B in 4-bit (and I think @ivanfioravanti would agree with me here)
-

Use MoE Models on Unified Memory Hardware Like DGX Spark
By
–
As I have mentioned before, stop trying to get Dense models running on the DGX Spark/Mac Studios Unified Memory is best fit for MoEs because you only make each token go through a small subset of the numbers of parameters in the model Optimize for your hardware
-
Qwen 3.5 27B Matches Sonnet 4.5 on Single RTX 5090
By
–
That model, Qwen 3.5 27B Dense is equal to Sonnet 4.5 Runs great on a single RTX 5090 w/ full context But we are not anywhere near Opus 4.5 even with Qwen 3.5 397B MoE
-
DGX Spark Qwen3 27B Inference Speed Discrepancy Questioned
By
–
Also, I don’t know how OP is getting Qwen3.5 27B @ 30 tps on the DGX Spark Doesn’t make sense for a Dense model on DGX Spark’s Unified Memory (273 Gbps) (My personal experiments showed it’s 4 tokens/sec) Very curious how you got that @TeksEdge
-
Qwen3 27B Performance Claims Disputed on DGX Spark Hardware
By
–
Also, I don’t know how OP is getting Qwen3.5 27B @ 30 tps on DGX Spark That number is impossible for a Dense model on DGX Spark’s Unified Memory (273 Gbps) My personal experiments showed it’s 4 tokens/sec for that model on the Spark
-
Unified Memory Faster for Loading Large MoE Models
By
–
the issue is that unified memory would still be faster for loading MoEs that are larger than the largest single GPU in terms of Memory
-

RTX PRO 6000 GPU Inference Faster Than Unified Memory After Loading
By
–
This will probably be great for Large single GPUs (e.g. RTX PRO 6000) You’re limited to 40Gpbs initially (during model loading) but then once the model is fully loaded on the GPU it should be extremely faster than Unified Memory speeds for inference
-
Drone modification adds precise impact and proximity sound features
By
–
brother come arrest me omg i'm so dangerous!!!!