Surface Laptop Ultra and RTX Spark: 128 GB of unified memory for local AI on Windows
On October 7, Microsoft and Nvidia unveiled the Surface Laptop Ultra, the first Surface laptop with the Arm-based Nvidia RTX Spark chip, built for AI that runs locally.
What's announced
| Item | Detail | |---|---| | Price | from $2,599 (entry) to $5,900 (128 GB memory, 1 TB) | | Memory | up to 128 GB unified memory | | Local AI | models over 120 billion parameters, per Microsoft | | Shipping | from October 16 | | RTX Spark Dev Box | from $6,000 |
RTX Spark PCs from Lenovo and Acer are also expected in October. Nvidia also says new llama.cpp and vLLM optimizations are available, directly and through LM Studio and Ollama.
Why unified memory matters
For local inference, the limit is the memory available for weights, not just compute. 128 GB shared between CPU and GPU can load models that regular graphics cards (24 to 32 GB of VRAM) can't hold. It's the approach unified-memory Macs already popularized.
Caveats
- The announced capabilities (120B+ locally) are vendor figures: real tokens/s depends on quantization and memory bandwidth, and hasn't been independently measured yet.
- A quantized 120B model already takes tens of GB, leaving little headroom for long context.
- At $2,600-5,900 the laptop targets developers and professionals. For occasional use, an API at $0.10 per million tokens is hard to beat (see our small-model price war article).
- Ecosystem: Windows on Arm and the CUDA/Arm stack are newer than the usual x86 or Mac setups; check your tools' compatibility (Ollama, LM Studio, vLLM) before buying.
The real appeal is privacy: keeping data and prompts on the machine.
Sources: Nvidia (IFA 2026), Memesita, Tech Startups, explainx.ai.