Contents
- RTX 4090 Specs
- Real-World Inference Speed on Llama 3.3 70B
- The VRAM Reality Check
- Buying Used: Risk Assessment
- Is a Used 4090 Still Worth It in 2026?
- Comparing to New Alternatives
- FAQ
RTX 4090 Used Market Is Pricey — But Still Worth Knowing
The RTX 4090's used market in June 2026 is not the $900–$1,200 bargain bin you might have heard about. It's reality-check time: used 4090s are selling for $2,100–$2,400 on eBay right now, depending on condition and model variant. That's roughly 30–40% under new retail ($2,755 for most AIB models), but it's not cheap.
Here's why you'd still consider one: it delivers 50–60 tokens/second on Llama 3.1 70B at Q4_K_M quantization — the fastest single-GPU experience for running 70B models. The 24GB VRAM is the bottleneck that matters. Yes, newer cards exist. But for today, if you want local 70B inference without multi-GPU complexity, the used 4090 is still the simplest play.
Warning
The used GPU market moves fast. Prices were higher in early 2026; they've stabilized somewhat by April. Verify current eBay listings and bestvaluegpu.com before assuming any specific price — these are snapshots, not floor prices.
RTX 4090 Specs: Built for Heavy Lifting
Let's nail down what you're actually buying:
| Spec | Value |
|---|---|
| VRAM | 24GB GDDR6X |
| Memory Bandwidth | 1,008 GB/s (384-bit bus @ 21 Gbps) |
| CUDA Cores | 16,384 |
| Tensor Cores | 512 |
| TDP | 450W |
| PCIe | Gen 4 x16 |
| Launch | October 2022 |
| Original MSRP | $1,599 |
The standout spec is that 24GB paired with 1,008 GB/s bandwidth. Bandwidth is what lets you move model weights on and off the GPU fast — critical for inference where you're constantly reading parameters. CUDA cores are plenty; it's the VRAM that's the real asset here.
Real-World Inference Speed: Llama 3.3 70B at Q4_K_M
Here's where it matters. Llama 3.3 70B Q4_K_M is 42.5 GB — larger than the 4090's 24GB VRAM. llama.cpp places as many layers as fit on the GPU and offloads the rest to system RAM. With 24GB, approximately 45 of 80 layers run on GPU; the remaining 35 layers run from CPU RAM.
Where these numbers come from:
- RTX 4090 bandwidth: 1,008 GB/s (GPU portion: fast)
- CPU RAM bottleneck: system RAM bandwidth (typically 40–100 GB/s for DDR4/DDR5) for ~18 GB of CPU-side layers
- Combined throughput is CPU-bottlenecked
Community-reported generation speeds:
- Generation speed: approximately 10–15 tok/s (varies significantly by system RAM speed and DDR5 vs DDR4)
- Prompt processing (prefill): 80–120 tok/s (GPU-limited, faster because GPU bandwidth dominates prefill)
- Power draw: 370–420W sustained
- Temps: 76–82°C (depends on cooling solution)
That's usable for 70B inference. You can maintain a real conversation without the minutes-per-response frustration. For faster speeds, increase the GPU layer count and use faster DDR5 RAM; for more context, reduce layers per GPU. See the llama.cpp advanced guide for flag reference.
Note
At Q3_K_M, the 70B model shrinks to ~32 GB — still too large for a single 4090, but fewer layers offload to CPU, improving speed to ~15–20 tok/s with trade-offs in reasoning quality.
Tip
Q4_K_M is the Goldilocks quantization for 70B inference — it preserves reasoning quality while keeping speed above 50 tok/s. For interactive coding or writing, this is the standard.
The VRAM Reality Check
Llama 3.1 70B at Q4_K_M takes about 42–45GB total when you account for model weights plus KV cache (the temporary memory for context). The RTX 4090 only has 24GB. How does this work?
llama.cpp splits the load: the 24GB on the GPU holds as many transformer layers as fit, and the remaining layers run from CPU RAM. The CPU RAM bandwidth is the bottleneck — typically 40–100 GB/s for DDR4/DDR5 versus 1,008 GB/s on the GPU. You're looking at ~10–15 tok/s instead of 1–3 tok/s on CPU-only. The speed hit from partial offload is real but acceptable for interactive use.
Bottom line: 24GB is enough to run 70B models at workable speeds. True full-GPU 70B inference requires 48GB+ (dual RTX 3090, Mac M4 Ultra, or H100).
Buying Used: Risk Assessment and Verification
This is the part that matters most. You're trusting a stranger's hardware.
Red Flags to Avoid
- Cryptominer stock with no history — Mining hammers the memory subsystem. Ask where the card came from. Recent consumer sales beat old miner lots.
- Coil whine under load — Test it before you fully commit. A card that whines under full utilization will drive you insane.
- Loose cooling solutions — Check if fans spin freely, heatsink is secure. Replacement thermal paste is $15; a shot cooling solution is a pain.
- No warranty or return guarantee — eBay's 30-day return policy is your safety net. Use it. Insist on Seller Guarantee coverage.
How to Verify Before Buying
Before closing the deal, ask the seller to run this stability test and share a screenshot:
stress-ng --gpu 1 --gpu-ops 0 --timeout 600s
Or if they have llama.cpp installed:
./main -m llama-7b-q4.gguf -n 128 -ngl 99 -t 4 --verbose
10 minutes of clean generation with no crashes = good sign. If the seller balks, buy elsewhere.
After you receive it, repeat the test locally. If VRAM degradation occurs mid-run (prompt tokens process fine, but output gets garbage after 2–3 minutes), that's a failing card.
Where to Buy Used 4090s
- eBay — Most selection, but 30-day returns. Prices $2,150–$2,350.
- Swappa — Verified sellers, slightly higher prices ($2,200–$2,400), solid guarantees.
- Local tech communities — meetup.com has local AI/hardware groups in most metro areas. Higher trust, easier in-person testing.
- B-stock retailers — Newegg B-stock occasionally has returned/open-box 4090s at $1,900–$2,100. Often carries factory warranty.
Is Used 4090 Still the Move in April 2026?
Here's the honest take.
Buy a used 4090 if:
- You need 70B inference right now and can't wait
- You found one under $1,400 (rare, but possible in local markets)
- You run 70B models multiple times per week and want the simplest single-GPU setup
- You're comfortable with the verification steps and 30-day return window
Wait if:
- You can run Qwen 2.5 14B or smaller for your use case — even an RTX 5060 Ti handles those well
- You're willing to wait 6 months for new 16GB+ consumer cards to stabilize in price
- You want new-card peace of mind (warranties, driver support, NVIDIA support forums)
- You're building for future-proofing over immediate performance
Comparing to New Alternatives
RTX 5090 (new): 32GB GDDR7 at $1,999 MSRP, but scarce as of June 2026. With 32GB it fits more of the 70B model on-GPU than the 4090, resulting in somewhat faster generation (~12–18 tok/s vs 10–15 tok/s). The 1,792 GB/s bandwidth also means models that fit entirely on-GPU are substantially faster. Worth considering if you can find one at MSRP.
RTX 5080 (new): 16GB at $999 retail. It cannot run Llama 3.3 70B at Q4_K_M without heavy offloading (model is 42.5 GB, card has 16GB). Fine for 8B–14B models at excellent speeds; not the right card for 70B inference.
Mac Mini M4 Pro 48GB ($1,999): Unified memory means all 48GB is available without PCIe overhead. Runs 70B at ~4–6 tok/s — slower decode than a 4090 setup, but fits the entire model cleanly without any offload. See the Mac Mini M4 Pro local LLM review.
For a single-GPU 70B setup, the used 4090 remains the simplest play. For smaller models at speed, a new RTX 5080 is better (warranty, lower power draw).
FAQ
How long will the RTX 4090 stay relevant for local AI?
Probably through 2027, maybe into 2028. New models get more efficient every quarter, so a 4090 in 2026 will handle the 200B models we see in 2027 — just with more quantization. The 24GB VRAM is the limiting factor, not raw performance.
Is a used 4090 good for anything besides 70B inference?
Absolutely. It crushes gaming at 4K, handles video encoding, fine-tunes models, runs Stable Diffusion without breaking a sweat. You're not locked into AI. That's a bonus if you get one.
What's the power cost difference between RTX 4090 and RTX 5080?
4090 pulls 370–420W sustained on inference (up to 450W peak). RTX 5080 pulls 200–250W sustained. At $0.15/kWh, the 4090 costs ~$100/month if you run it 12 hours a day. 5080 costs ~$35/month. Over a year, that's $780 difference. Factor that into your ROI if you're doing heavy daily inference.
Can I use a used 4090 in a gaming PC?
Yes. Two caveats: make sure your PSU is 1000W+ (4090 + CPU can spike to 550W+), and you'll want a full-size case. The 4090 is a chonky card. But performance-wise, it's overkill for modern 1440p gaming — you're paying for 70B capabilities you're not using. If you're gaming and running AI half-and-half, it's fine.
Should I buy now or wait for prices to drop further?
Used 4090 prices have been trending upward slightly as new cards like the 5090 stay scarce. They've stabilized around $2,100–$2,400 since January 2026. If you need one now, get it. If you can wait until Q4 2026 when more 5090s ship and 4090s flood the secondhand market, prices might dip $300–$500. But that's speculative.
Final Verdict
The used RTX 4090 is still the single-GPU king for running 70B models locally. At $2,100–$2,400, it's expensive — but it's the 24GB single-card option for 70B inference, delivering approximately 10–15 tok/s with partial CPU offload. That's workable for daily interactive use on serious 70B models. Verify stability before purchase.
For smaller models (8B–14B), a new RTX 5080 or RTX 5070 is faster and comes with a warranty. For 70B inference from a single card, the used 4090 remains the practical choice in 2026.
Take the time to verify. Don't skip the stability test. Use eBay's return guarantee. Then run your favorite 70B model at 50+ tok/s and enjoy the silence of not having to wait for your GPU to think.