NVIDIA RTX Spark PCs Launch in October With 128GB Memory

NVIDIA RTX Spark laptops bring Blackwell AI performance and shared memory to compact systems launching in October.

Hardware by Nahe Yan on  Sep 06, 2026

AI hardware is moving toward systems capable of running larger models locally, while software updates are improving performance on existing RTX hardware. This week brings developments around NVIDIA's Spark platform, local AI software, and a new open model that can now run on a single Spark system.

The main story this week is NVIDIA's RTX Spark PCs. NVIDIA confirmed that the RTX Spark PCs will ship in October. It is the DGX Spark chip, now available in Windows laptops and small desktops. It has 20 CPU cores, a Blackwell GPU, up to 1petaflop of AI compute, and up to 128GB of shared memory.

NVIDIA RTX Spark Launch 128GB Memory

NVIDIA RTX Spark PCs Arrive in October

Lenovo has two laptops so far: the Yoga Pro 9i and a 2-in-1. There are no prices yet. As with the Sparks, memory bandwidth really decides how fast these machines will write text day to day. Petaflop numbers can be misleading on their own. As a reference point, the DGX Spark has 273GB/s of memory bandwidth.

On two of them, a 27-billion-parameter model writes at about 40tokens/s. Lenovo's spec sheet puts the laptop memory a touch faster than desktop boxes. That suggests similar speeds, around 30 to 60 tokens/s, on a model that fits one person at a time. 

Testing with the Sparks over a 200GB link showed that going from two boxes to four gave 1.46x the speed, well short of double on a far faster link than any home network. Pooling memory across machines is now potentially possible, and testing will show how well the speed holds up when people start using P.A.I.R. routers after release.

RTX Spark Performance Still Needs Independent Testing

A few honest caveats remain. Nobody outside NVIDIA has benchmarked an RTX Spark yet. The 1petaflop figure is a 4-bit figure that assumes sparsity, which is the most generous way of counting, and it runs Windows on ARM. So it is worth checking that the applications you use will run. The quieter announcement probably matters more today.

NVIDIA has added new optimizations to llama. cpp and vLLM, the engines underneath LM Studio and Ollama, the two apps most people run models with, and says local models run up to 1.9x faster on RTX cards. For existing RTX owners, the speedup arrives with the next update. The third story is the Qwen 3.8 Flash model. It came out on Wednesday the 26th, the same day as GLM 5.3 Flash.

Two days later came the giants: Tencent Hunyuan preview at 770 billion parameters and the full GLM 5.3 at a similar size, both open-weight, and neither fits in one box. Comparing models on the Artificial Analysis database, the two Flash models score 46 each, and the big GLM scores 49. What is new this week is that the smaller one now runs at home.

NVIDIA DGX Cluster Runs Model

Updating LM Studio or Ollama is the first step to take advantage of the available performance.

But there are two catches. That build squeezes the model down to about 2 bits per weight, and the 64 tokens/s figure is for structured output, with ordinary prose closer to 25 tokens/s. For comparison, a four-DGX Spark cluster runs the same model as a 4-bit MXFP4 build spread across all four boxes at about 46 tokens/s for a single user.

A 2-bit build on one box is a real achievement, but you should check the quality numbers before relying on it. Nobody has posted a Hunyuan recipe yet. In a separate development, llama.cpp version 0.4 came out on Thursday and has the first support for it. One post also shows GLM 5.3 Flash running at 64 tokens/s on a single Spark.

Anyone considering an RTX Spark should wait for the first independent benchmarks in October, although the memory bandwidth already indicates what to expect. The main question is whether it makes more sense to connect two ordinary PCs with P.A.I.R. or buy one machine with enough memory.

Nahe Yan

Editor, NoobFeed

Gaming Hardware Updates

No Data.