Two machines built for local AI inference—NVIDIA's DGX Spark and AMD's Ryzen AI Max+ 395—arrive at a crossroads where the decisive question is not which is faster, but faster at what. The Spark processes prompts at five times the speed of its rival, yet the two are nearly indistinguishable in generating output, a distinction that maps almost perfectly onto how different people actually work. Beneath the benchmarks lies a further humbling truth: software backend choices can shift apparent performance as much as the hardware gap itself, and the price difference, once framed as two-to-one, has qu
DGX Spark vs Ryzen AI Max+: The Gap Isn't What Benchmarks Say
Software choices can move the apparent gap as much as the machines differ
So the Spark is five times faster at prompts but only 13 percent faster at generation. How do I know which one matters for what I do?
You measure it. Log your actual token counts for a week—how many tokens go into prompts versus how many come out as completions. That ratio tells you everything. If you are mostly generating, the AMD box wins on price. If you are mostly processing long prompts, the Spark's advantage becomes real.
But the article says that five-times number is not stable. It moved depending on the software version and backend. How confident should someone be in that?
That is exactly right. The five-times is directional—it shows the Spark has an architectural advantage at prompt processing. But the actual number you get depends on which version of llama.cpp you run, whether you use ROCm or Vulkan, what quantization you pick. One person measured a 36 percent swing just by changing backends on the same hardware.
So benchmarks are not useless, they are just not predictive of what I will actually see?
Correct. They show you the shape of the difference, not the magnitude. The Spark will be faster at prompts than the AMD box, but whether it is 2x or 5x depends on your exact setup.
The price comparison also shifted. The article says it is not 2x anymore, it is 1.3x to 1.8x. But that is still a lot of money. What do you get for that premium besides prompt speed?
Clustering capability, if you need it. The Spark has ConnectX-7 networking that lets you pool two boxes into 256GB. The AMD boxes can network over Ethernet or Thunderbolt, but the bandwidth is orders of magnitude lower. And CUDA support, if your software stack depends on it.
What about the memory situation? Both say 128GB, but the article suggests that is not the whole story.
On the Spark, 128GB is coherent across CPU and GPU, but the OS and runtime take their share. You get roughly 87GB free at launch. On the AMD side, Windows caps VRAM at 96GB, but Linux goes higher. One user reported 101GB in use. So it is not a clean comparison.
And if your model is 100GB, that matters a lot.
Exactly. If your model is 95 to 110GB, the two machines are not substitutes. One person had to stream part of their model off SSD to make it fit on the AMD box, when it would load cleanly on the Spark.
So the real decision is: what do I actually do with this machine, and what do I actually need to run?
Yes. Measure your token ratio. Check your model sizes. Know whether you need CUDA. Those three things decide it more than any benchmark.
The Pulse
- A single benchmark split the comparison in two: the Spark is five times faster at reading your input, yet only 13 percent faster at producing output—a gap that is everything or nothing depending on your workflow.
- Software variability threatens to swallow the hardware story entirely, with backend switches on a single machine moving generation speed by 36 percent—more than the entire measured gap between the two devices.
- The price narrative has shifted; AMD machines that once seemed half the cost now land between $2,600 and $3,600 against the Spark's fixed $4,699, compressing the premium to a 1.3–1.8x multiple rather than the 2x ratio that shaped early debate.
- Hidden memory ceilings complicate both sides: the Spark's 128GB shrinks to roughly 87GB free at launch, while AMD's ceiling rises or falls with operating system and configuration—making models in the 95–110GB band the true stress test.
- The resolution being offered is deceptively simple: log your own prompt and completion token counts for a week, take the ratio, and let that single number decide which machine is worth its price.
Two machines built for local AI inference—NVIDIA's DGX Spark and AMD's Ryzen AI Max+ 395—arrive at a crossroads where the decisive question is not which is faster, but faster at what. The Spark processes prompts at five times the speed of its rival, yet the two are nearly indistinguishable in generating output, a distinction that maps almost perfectly onto how different people actually work. Beneath the benchmarks lies a further humbling truth: software backend choices can shift apparent performance as much as the hardware gap itself, and the price difference, once framed as two-to-one, has quietly narrowed. The purchase decision, it turns out, depends on a measurement most buyers have never made about themselves.
The comparison between NVIDIA's DGX Spark and AMD's Ryzen AI Max+ 395 appears clean until the numbers arrive—and then it fractures along a single fault line. On prompt processing, the Spark is five times faster. On token generation, it leads by 13 percent. Both figures come from the same benchmark, the same model, the same run. Which one matters is entirely a function of what you do with the machine, and most people have never measured that about themselves.
The deeper problem is that neither figure is stable. Software backend choices—Vulkan versus ROCm, different llama.cpp builds, batch settings—can move performance by as much as 36 percent on a single machine, which is larger than the entire measured gap between the two devices. One AMD maintainer demonstrated this on his own hardware: a backend flag shifted generation speed more than the hardware difference between the two competing products. The five-times prompt advantage is better understood as an architectural statement than a specification to budget against.
The price story has also quietly changed. The Spark now costs $4,699 following a list price increase. AMD machines that launched near $1,999 have drifted upward with memory shortages, with recent listings ranging from roughly $2,600 to well past $3,500. The multiple is now 1.3 to 1.8, not the 2-to-1 ratio that defined the early debate. NVIDIA's pricing discipline and the Spark's bundled ConnectX-7 networking—which allows two units to pool memory for larger models—further compress the practical gap for buyers with expansion in mind.
Memory availability on both sides is softer than the spec sheet implies. The Spark's 128GB coherent pool shrinks to roughly 87GB free once the OS and runtime take their share. AMD's ceiling moves with operating system and configuration, reaching over 100GB on Linux in documented cases. For most models this is irrelevant. For models in the 95–110GB band, it decides everything.
The honest purchase framework is one neither vendor can supply: measure your own token ratio. A week of logging prompt and completion counts will reveal whether your workload is generation-heavy—in which case AMD does the job for less—or prompt-heavy, in which case the Spark's architectural advantage becomes decisive and its premium smaller than it first appeared. That single self-measurement is worth more than any benchmark in the comparison.
The comparison between NVIDIA's DGX Spark and AMD's Ryzen AI Max+ 395 looks straightforward until you run the numbers. Then it splits in half, and each half points a different direction.
On token generation—the speed at which a model produces output once it has read your prompt—the Spark leads by 13 percent. That is the kind of margin you notice in a spreadsheet and forget in practice. On prompt processing, the speed at which the machine reads what you just typed or pasted, the Spark is five times faster. That is the kind of margin that changes how you work. Both measurements come from the same benchmark run on the same model, GPT-OSS 120B at MXFP4 quantization through llama.cpp, measured by HardwareCorner and collected by IntuitionLabs. The Spark delivered 1,723 tokens per second of prompt processing against 340 on the AMD box, and 38.6 tokens per second of generation against 34.1. The gap you should care about depends entirely on what you actually do with the machine, and almost nobody has measured that about themselves.
But there is a deeper problem hiding inside these numbers, one that most comparisons skip entirely. The five-times prompt advantage is not a hardware constant. It is a software snapshot. In llama.cpp's own development threads, contributors have documented variance that makes any single benchmark run a shaky foundation for a purchase decision. One user reported a sixfold collapse in prompt processing speed after moving to a newer Vulkan build, which vanished entirely when switching to ROCm. Another reported close to a tenfold drop on the same model across different builds on an RTX 3090. The Spark side moves too. NVIDIA's own official benchmark discussion shows the same model at 2,046 tokens per second of prompt processing and 45.3 of generation with larger batch settings, against the 1,723 and 38.6 that HardwareCorner measured. A public Strix Halo benchmark repository reports the AMD machine generating 53.4 tokens per second on the same model, using a different quantization and serving method. Line those up and the AMD machine generates somewhere between 34 and 53 tokens per second depending on how you run it, and the NVIDIA machine somewhere between 38 and 45. The honest version of this comparison is that software choices can move the apparent gap between these two machines about as much as the machines differ from each other.
The clearest evidence comes from an AMD maintainer who rebuilt llama.cpp both ways on a single machine with an idle GPU and measured ROCm at roughly 59 tokens per second of generation against Vulkan at roughly 80. A backend flag moved generation by about 36 percent on one machine. The measured generation gap between the two machines was 13 percent. The same person measured vLLM at around 20 tokens per second for single requests and 162 to 181 for batched offline benchmarks, a gap that reflects the difference between how synthetic tests run and how real serving works. This is not a reason to ignore benchmarks. It is a reason to treat the five-times prompt figure as a directional statement about architecture rather than a specification you can budget against.
The price gap that made this comparison famous has also shifted. The Spark costs $4,699 after NVIDIA's 18 percent list price increase in February 2026. The 128GB AMD machines launched around $1,999 and have drifted upward with memory shortages, with a July survey of Amazon listings putting the cheapest at roughly $2,600 and the range running well past $3,500. That is a multiple between 1.3 and 1.8, not the 2 to 1 ratio that circulated in 2025. The gap is real and still material, but it moves in NVIDIA's favor for a reason worth noticing: NVIDIA sets a list price and holds it, while the AMD boxes are sold by a dozen OEMs whose pricing tracks the DRAM spot market. In a shortage, the vendor with pricing discipline looks cheaper over time even when it started more expensive. One more line item that rarely gets counted: the Spark ships with ConnectX-7 networking, which lets two boxes pool 256GB for a 405B model. Bought separately that card is not a rounding error, and there is no equivalent path on the AMD side.
The real difference between these machines is not what the specs say. The Spark's 128GB is coherent across the CPU and GPU with no fixed allocation ceiling, but that is not the same as 128GB being available to your model. DGX OS, the runtime, and everything else resident take their share. NVIDIA's own benchmark output shows the GPU reporting roughly 119GB total with about 87GB free at launch. On the AMD side, AMD's product documentation states that up to 96GB of the 128GB can be converted to VRAM on Windows through Variable Graphics Memory. On Linux, GTT allocation reaches considerably higher, with one engineer at Red Hat reporting 101GB in use on a DeepSeek V4 run. The real story is not a clean 96 against 128. It is a configured ceiling that moves with your operating system on one machine and a soft ceiling set by whatever else is running on the other. For most models this never comes up. For models in the 95GB to 110GB band it decides everything.
What actually decides the purchase is a measurement neither vendor can make for you: what fraction of the tokens you process are prompt tokens? If you chat, generate, and write code in short exchanges, most of your tokens come out of generation, generation is a near tie, and the AMD box does the job for less. If you paste 30,000-word documents, run retrieval over a corpus, or operate an agent that resends its context every turn, most of your tokens go into prompt processing, and that is where the Spark stops being a 13 percent machine and starts being a several-times machine. This is the same split that decides every machine in the broader AI hardware landscape, and almost nobody knows the number about themselves. It is cheaper to find out than to buy the wrong box twice: log your prompt and completion token counts for a week of real work and take the ratio. That single measurement is worth more than every benchmark in this comparison. If your workload is generation-heavy and your models fit comfortably in what AMD will assign to the GPU, and you do not depend on CUDA, NVIDIA is difficult to justify on generation performance alone. If your workload is prompt-heavy, if you need CUDA and TensorRT-LLM, if your models sit in the 95GB to 110GB band, or if a second box is anywhere on the roadmap, NVIDIA is worth it and the premium is smaller than the headline suggests.
Notable Quotes
The honest version of this comparison is that software choices can move the apparent gap between these two machines about as much as the machines differ from each other— Analysis of llama.cpp benchmark variance
What fraction of the tokens you process are prompt tokens? If you chat, generate, and write code in short exchanges, most of your tokens come out of decode, decode is a near tie, and the AMD box does the job for less— Workload-based purchasing guidance