Speculative Decoding in vLLM on AMD GPUs
Posted by ankitg12 1 day ago
Comments
Comment by intothemild 1 day ago
Stock vLLM runs so slowly on these cards compared with vLLM forks like Radiance. Going from say 20-30t/s gen, to 150-200t/s
Most of AMD/vLLM work seems to be around their data centre cards, or the AMD AI Halo/Ryzen and ignores the R9700 AI Pro.
Really wish this would change.
Comment by roenxi 1 day ago
It makes a huge amount of sense after considering AMD's approach to graphics cards from around 2010 to 2025. They just didn't see graphics cards as viable compute platform and many who made the mistake of believing that good specs would translate into in-practice performance got badly burned. I'd have been involved in the AI boom but for an expensive AMD graphics card, I'm not going to forget that for a while.
George Hotz was interesting as a public example, but I think his story probably repeated a few times outside the public eye. People tried to make AMD work and ended up the worse for it.
People who had an interest in using AMD cards to get things done are probably by and large waiting for a new generation of hopefuls to prove this time is different. The mutterings out of AMD are promising, but that isn't persuasive enough given the scale of the failures.
Comment by hgoel 21 hours ago
I agree with your assessment that the story of supporting the competition, only to get burned, has repeated many times with AMD outside the public eye. It's why I don't put much stock in claims that things work great as long as specific flags are used.
Comment by lrvick 19 hours ago
Comment by pyrolistical 16 hours ago
For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp
Now if you meant 4xr9700 tensor parallelism with mtp, 80 tg/s starts to make sense
Comment by latentsea 11 hours ago
Comment by naasking 15 hours ago
Comment by SomeHacker44 16 hours ago
Comment by lrvick 11 hours ago
Obviously only on a system you do not trust at all.
Comment by androiddrew 1 day ago
Comment by intothemild 1 day ago
The MXFP4 fork is excellent too. Its my daily driver right now. https://codeberg.org/ggz14/radiance-vllm-mxfp4
Also has PARO quant support there too (early stage)
Also speedups in both repos for 4x R9700s
Comment by karmakaze 19 hours ago
What kind of performance are you getting with 4x R9700s--what do you do with all the VRAM (batching, concurrent requests, etc)?
Comment by intothemild 18 hours ago
Yeah the R4D Kernel rules imho.
Comment by karmakaze 8 hours ago
I just got DeepSeek Harness (DSH) set up with 2x R9700 and it's rather mind blowing that these can do actual work and quickly. Up until now I've always been evaluating and searching for better hardware/model/tweaks. This is much more than I even hoped for and considered getting extra 3090/4090. Now I can stop looking/tweaking and start using it for all the different things I've yet to discover it's good for. I do plan to also try/use Hermes and Pi. DSH is annoying that every plugin install/remove requires a restart--given that "everything's a plugin".
Comment by karmakaze 22 hours ago
Comment by nicce 22 hours ago
Comment by intothemild 22 hours ago
Comment by minraws 1 day ago
They might in the future but future is in the future ofc
Edit: to be clear I think it's ridiculous they don't but from a company's stand point it doesn't make much sense
Comment by websap 1 day ago
Comment by _factor 1 day ago
Comment by brookst 1 day ago
Comment by mrhenio 23 hours ago
Comment by Roark66 1 day ago
Only if AMD made a card like this with 48G+ I'd consider it.
Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.
I believe Intel has claimed something like 1400tok/s (generation! Not prefill) of Qwen3.6-moe on Arc b70.
I was actually very interested in this so I checked the details. Turns out it was 200 simultaneous users running the same 1024 token prompt :D so all the experts got maximum parallelism.
How often are you going to run 200 parallel sessions with a tiny context and same prompt running at 7tok/s.
Based on how much my rtx3090 is getting on a single user (150tok/s) I'm estimating b70 to probably get less than that.
Sadly nvidia is king now.
Also, most of us already have nvidia cards and no inference software supports mixing let's say nvidia, Intel and amd cards in inference of one model.
Comment by nicce 22 hours ago
You would need 3x RTX 3090 - but then the PCIe bandwidth comes an issue if you really want tensor parallelism with three cards. 3x PCIe5 x16 is not cheap with direct CPU access.
So surprisingly, 2x r9700 starts be a nice deal.
> Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.
Well, luckily these are not vendor numbers. Prefill also scales almost linearly with the amount of GPUs.
Comment by latentsea 11 hours ago
I don't see how. Even on 32GB you can run Q6_K_XL quant with MTP at 200k context k=q8_0, v=q5_1. So 48GB VRAM is good enough to run Q8 at long context. Also with tings like ninfer and it's various forks I'm seeing people get very good performance out of Qwen3.8 models on all sorts of NVIDIA cards.
Comment by nicce 4 hours ago
Comment by formerly_proven 21 hours ago
Comment by nicce 21 hours ago
Comment by snovv_crash 20 hours ago
Comment by nicce 19 hours ago
Comment by dist-epoch 1 day ago
> I have had direct contact with members of the AMD RTG team and I was disgusted to find that AMD doesn't even provide them with hardware to work on. The developer I was working with had to buy the GPU he was writing drivers for.
Comment by da-x 1 day ago
Comment by dist-epoch 1 day ago
An NVIDIA consumer GPU sells for 50+% or more than an equivalent AMD GPU. Because people are buying NVIDIA GPUs to run local models instead of AMD ones.
I did the same thing, I paid 50% more to get an 5070 Ti instead of the equivalent AMD.
This is probably good for gamers, AMD GPUs are not price inflating to the same degree as NVIDIA, because they are bad at LLMs.
> That was the reason for comparing them in the first place: based on performance, they are direct competitors, or at least they are meant to be. However, as things stand today, there is a massive price divide between the two, with the RTX Ti GPU now commanding a premium of more than 50%.
https://www.techspot.com/review/3168-geforce-rtx-5070-vs-rad...
Comment by sznio 1 day ago
Comment by prymitive 1 day ago
Comment by esseph 15 hours ago
The AMD drivers are mainline kernel, while Nvidia is still handing out binary blobs.
Comment by numpad0 23 hours ago
Comment by esseph 15 hours ago
It would have been sold off to grey market or destroyed 3-4yr ago in most companies.
Newer, much cheaper stuff ironically would run better. That card isn't even really designed for AI workloads at all.
Comment by lrvick 19 hours ago
They are actually great at LLMs but you need to invest in keeping up with community tuning efforts. But I am fine with most thinking they are bad at LLMs, because I keep buying more of them!
Comment by esseph 1 day ago
9060 XT 16GB seems to have some great performance with gpt oss 20B and others, and works great with their lemonade-server.
What exactly do you think is running on a Strix Halo?
Comment by flufluflufluffy 1 day ago
Also, what is the difference between “target model” and “target-model,” if any? I feel like half the instances of that phrase included the hyphen and half didn’t.
Comment by daemonologist 22 hours ago
If you have some other source of parallel data (lots of users, many separate tasks) then speculative decoding might not provide any benefit.
Comment by zackangelo 20 hours ago
I think the piece of information that might make this click for you is the model outputs the probability distribution for all intermediate tokens even during prefill.
So for example, let's say you prefill the prompt "The quick brown fox" (and for the sake of simplicity, let's say each word is a single token). The model outputs a tensor that is [4, $vocabulary_size]. The first dimension is a token index into the input and the 2nd dimension assigns a probability to each token in the vocabulary. So even during prefill, we can look at the prediction logits for all of the intermediate tokens. That is, we can look at what the model would have predicted after "quick" and "brown", not just the tail token "fox".
In the single token autoregressive case, we just look at the next token prediction for "fox". But in the speculative decoding case we can use this information to compare the distribution of the draft model against the target model. In the greedy decoding case (ie, no sampling) we just make sure the highest probability token matches in draft and target. If we have sampling params like temperature and top-P, we have to apply something called Leviathan rejection sampling to the distribution. This basically allows us make sure the distribution is the same even if the exact probabilities are not and accept or reject draft tokens on that.
Comment by lucrbvi 1 day ago
Comment by ahepp 18 hours ago
If the draft was wrong you can kill the speculative decode, and you haven’t lost anything except for idle time
It’s true we can’t show the user the token until we have the true decode finished, but we can launch more work internally before we’re certain
Comment by cesarb 22 hours ago
Yes, but you can do it in parallel.
Suppose you predicted the tokens "D E F" in the sequence "A B C D E F". To "generate" the last token (F), it must know all preceding tokens (A B C D E). To "generate" the next-to-last token (E), it must know all preceding tokens (A B C D). And so on.
Assuming the prediction is correct, it can then run the "generation" for tokens D, E, and F at the same time. At the end, after all these tokens were "generated", it compares each token with the prediction; if the "generation" result was "D H F" it knows it has to discard the last two predicted tokens (and output "D H"), if the "generation" was "D E H" it knows it has to discard the last predicted token (and output "D E H"), etc.
And the most important part is that you can do it in parallel for each layer of the model. That is, you run "A B C D E F" through the first layer, then through the second layer, and so on; you only have to load the model weights from memory once for each layer. Instead of reading the full weights for all layers once for D, then once for E, then once for F, you only read them once for "D E F", and if the prediction was correct, you output three tokens by the (memory read) price of one (you still had to do the same amount of compute, but AFAIK LLMs tend to be more memory-bound than compute-bound).
Comment by porridgeraisin 22 hours ago
Correct except for the word "autoregressive". When you have to verify a sequence of tokens (which were autoregressively generated by the cheap model), you can do each token in parallel. This amortizes the cost of loading the weights from vram to the processors (the primary cost in LLM serving) across those tokens. Cost here is wall clock time, as well as power.
The autoregressive decoding that generates this batch of tokens is delegated to the cheaper model where the cost of loading the weights is lower and so not amortizing it is fine.
Verification means, how close is each token in this sequence to the one I would have output. You keep the longest prefix that is close enough for your liking.
Comment by quietraster 19 hours ago
Comment by ThiraSoft 21 hours ago
Comment by hn45e7pbij 1 day ago
Comment by myuzio 22 hours ago
Comment by foota 21 hours ago
Comment by jeanmichelselli 22 hours ago
Comment by suprjami 16 hours ago
"If you remove the hype about how transformative they are, they really are transformative".
No. If you remove the hype about how transformative they are, you're left with what they actually are: a sometimes mildly useful tinker toy.