The short version
Microsoft’s Surface Laptop Ultra, announced on 7 October 2026, puts Nvidia’s RTX Spark and up to 128 GB of unified memory in a laptop, so large models can run on the device. For developers the useful question is which requests should stay local. That depends on memory math, latency, privacy and having one interface in your code that can talk to both a local model and a cloud one.
What Microsoft and Nvidia announced
At its 7 October event in San Francisco, Microsoft introduced the Surface Laptop Ultra, built around Nvidia’s RTX Spark. Coverage from BGR, Notebookcheck and Engadget agrees on the main points.
- RTX Spark combines a Grace CPU with a Blackwell GPU that has 6,144 cores, and the CPU and GPU share one pool of unified memory, up to 128 GB.
- The Surface Laptop Ultra starts at $2,599 with 24 GB of memory, with preorders open and availability from 16 October. A Surface RTX Spark Dev Box is listed at $5,999.
- Laptops from Lenovo, Asus, Dell, MSI and HP with RTX Spark were also announced, shipping from 16 October.
- Windows gets hybrid behaviour that decides between local and cloud models, Copilot access to local files with permission, and Microsoft Execution Containers to contain what agents can do.
A few details differ between outlets, such as which specific open models were named and how much RAM they need, so I am relying only on the hardware and platform points that several sources agree on. Microsoft’s own product page is the place to confirm final specs and prices, and independent benchmarks will fill in the performance picture.
Why unified memory is the headline
Running a language model locally is mostly a question of whether the weights fit in memory the GPU can reach. On a typical laptop with a discrete GPU, that memory is small, and the model has to be squeezed to fit. With unified memory, the CPU and GPU draw from the same pool, so a 128 GB machine can hold models that used to need a workstation.
Capacity is only half the story. How fast the memory can feed the chip decides how many tokens per second you get, and I have not seen verified bandwidth or throughput numbers yet. That is why independent benchmarks are worth waiting for.
The memory math
You can estimate whether a model fits with one line of arithmetic. The weights take roughly parameters times bits per weight, divided by eight.
// Weights only, in gigabytes, for a model with paramsBillions parameters.
export const weightsGB = (paramsBillions, bitsPerWeight) =>
(paramsBillions * bitsPerWeight) / 8
weightsGB(8, 4) // 4
weightsGB(70, 4) // 35
weightsGB(70, 16) // 140
weightsGB(675, 4) // 337.5- An 8 billion parameter model at 4 bits per weight needs about 4 GB. It fits almost anywhere.
- A 70 billion parameter model at 4 bits needs about 35 GB. It fits on a 128 GB machine with room to spare, and not on a 24 GB one.
- The same 70 billion model at 16 bits needs about 140 GB. It does not fit on 128 GB.
- A sparse mixture of experts model is the trap. Mistral Large 3 has 675 billion parameters in total but only 41 billion active per token. Active parameters set the speed, while total parameters set the memory, so at 4 bits it needs about 338 GB resident. It will not run on a laptop however few experts fire.
Treat the result as a floor. The attention cache grows with context length and with each concurrent session, and the runtime needs working space. As a starting guess I leave about 20 percent of memory free, then measure with the real context length my application uses.
Local or cloud: write the rule down
Having the hardware does not mean everything should run on it. I decide per request type, using five questions.
- Privacy. Does the input contain data that should not leave the device?
- Latency and offline use. Does the feature need to respond instantly or work without a connection?
- Cost. Is it called so often that per token pricing adds up?
- Quality ceiling. Does the task need the strongest model available, or is a good enough one fine?
- Context and load. Does it need a very long context, or will traffic spike beyond one machine?
In general, private documents, autocomplete, classification and first drafts are good local candidates. The hardest reasoning, very long contexts and bursty workloads still belong in the cloud. Most real products will use both, which is why the next part matters.
One interface, two backends
Most local runtimes expose an OpenAI compatible HTTP API. The llama.cpp server, for example, serves chat completions on a local port, so the client code is the same and only the base URL changes. Put that behind one function and the rest of your app never needs to know where a request ran.
const backends = {
local: { baseUrl: 'http://localhost:8080/v1', model: 'local' },
cloud: { baseUrl: process.env.CLOUD_BASE_URL, model: process.env.CLOUD_MODEL, key: process.env.CLOUD_KEY },
}
async function chat(backend, messages, timeoutMs) {
const response = await fetch(backend.baseUrl + '/chat/completions', {
method: 'POST',
headers: {
'content-type': 'application/json',
...(backend.key ? { authorization: 'Bearer ' + backend.key } : {}),
},
body: JSON.stringify({ model: backend.model, messages }),
signal: AbortSignal.timeout(timeoutMs),
})
if (!response.ok) throw Object.assign(new Error('request failed'), { status: response.status })
const data = await response.json()
return data.choices[0].message.content
}
export async function complete(messages, { private: isPrivate = false } = {}) {
if (isPrivate) return chat(backends.local, messages, 60000) // never leaves the device
try {
return await chat(backends.local, messages, 8000)
} catch {
return chat(backends.cloud, messages, 30000)
}
}Notice the private flag. Falling back to the cloud when the local model is slow is a convenience for ordinary requests. For sensitive input it would defeat the point, so those calls wait for the local model and fail instead of leaving the machine.
Agents on a laptop need walls
Microsoft says Copilot can now read local files with permission and take actions on the machine, and that Microsoft Execution Containers will contain what agents are allowed to do. That is the right instinct. An agent with access to your files is a program that follows instructions from text it reads, and some of that text will be hostile.
Whatever platform you build on, give each tool the least access it needs, validate every argument in code, and require confirmation for anything that changes something. The same discipline applies to cloud agents, as in An AI Booking Assistant That Cannot Double Book.
What I would do this week
- Run the memory math for the models you actually want to use, with your real context length.
- Put local and cloud behind one function so you can move requests either way.
- Decide which request types must never leave the device.
- Look for independent tokens per second numbers before choosing hardware.