CraftRigs

News

Hardware releases, driver updates, and industry developments that matter for local AI builders.

59 articles
Sort:
DeepSeek V4 2026: What Fits Your GPU? — diagram
Technical Report

DeepSeek V4 2026: What Fits Your GPU?

DeepSeek V4's 400GB VRAM demand breaks most builds—32B distilled fits RTX 4090, 70B needs dual-GPU, 671B demands 8× setup. Match model to hardware tier before you buy.

DeepSeeklocal LLM
Llama 4 Scout: 70B Quality on a 24 GB GPU? — diagram
Technical Report

Llama 4 Scout: 70B Quality on a 24 GB GPU?

Dense 70B models need 42 GB VRAM and fail on consumer GPUs—Llama 4 Scout's 109B/17B MoE design runs at 28 tok/s in 18.5 GB. Here's the catch: 4K context ceiling, routing opacity, and a 5-point quantization hit you can't ignore.

MoELlama70B models
M5 Max vs RTX 4090 for Local LLMs: The Memory Verdict — diagram
Technical Report

M5 Max vs RTX 4090 for Local LLMs: The Memory Verdict

Desktop GPUs win on tok/s but choke at 24 GB—M5 Max's 128 GB unified memory runs 70B+ models and 405B MoE desktops can't touch. 18–22 tok/s, zero throttling, one machine for code and inference.

GPUM5 Maxlocal LLM
Local LLM Benchmarks 2026: Best Models by VRAM Tier — diagram
Technical Report

Local LLM Benchmarks 2026: Best Models by VRAM Tier

Chasing leaderboard ELO wastes VRAM—Qwen3-235B hits 82.1 MMLU-Pro at 48 GB, but 24 GB buyers win with 32B Q4. Match model to hardware tier and task, not raw rank. Here's the benchmark-grounded map.

benchmarksopen source LLMlocal LLM
Qwen 3.6 27B Beats 70B Models on 24GB VRAM — diagram
Technical Report

Qwen 3.6 27B Beats 70B Models on 24GB VRAM

70B models choke your GPU—Qwen 3.6 27B hits 73.2% HumanEval at 18–22 tok/s on a single RTX 3090, no cloud bill. Run it locally in one Ollama command.

Qwen70B modelslocal LLM
RTX 5060 vs RX 9060 XT: 8GB Trap or 16GB AI Win? — diagram
Technical Report

RTX 5060 vs RX 9060 XT: 8GB Trap or 16GB AI Win?

8GB GPUs choke on 13B models—RX 9060 XT 16GB unlocks them for $349, beats RTX 5060's VRAM wall, but a $220 used RTX 3090 changes everything. Pick right, not new.

RTX 5060RX 9060 XTlocal LLM
RTX 5060 vs RX 9060 XT vs Arc B580 for Local LLMs 2026 — diagram
Technical Report

RTX 5060 vs RX 9060 XT vs Arc B580 for Local LLMs 2026

RX 9060 XT doubles VRAM for $299 but needs ROCm 6.3 flags; RTX 5060 runs one-click at 8 GB; Arc B580 is beta-only. Pick the GPU that matches your tolerance for compile flags, not just your budget.

nvidia amd intel 2026 local ailocal LLM
llama.cpp TurboQuant vs vLLM 2-bit: 24GB Card Winner — diagram
Technical Report

llama.cpp TurboQuant vs vLLM 2-bit: 24GB Card Winner

TurboQuant vs vLLM 2-bit KV on 24GB: 64K context, 38 tok/s vs. 128K, 18 tok/s. Which Llama 70B quantization actually wins? April 2026 head-to-head benchmark.

kv-cache-quantizationlocal-llmllama-benchmarks
RDNA4 Windows ROCm Broken: RX 9070 XT Workarounds — diagram
Technical Report

RDNA4 Windows ROCm Broken: RX 9070 XT Workarounds

ROCm broken on RDNA4 Windows. Vulkan workaround: 28–32 tok/s on Llama 70B Q4. Setup guide, throughput benchmarks vs. Linux ROCm, timeline for Windows RDNA4 fix.

rdna4windowsrocm
RTX 5080 $1,249: 50-Series Pricing Breaks Open — diagram
Technical Report

RTX 5080 $1,249: 50-Series Pricing Breaks Open

RTX 5080 dropped to $1,249 in April 2026—$250 under MSRP. GDDR7 yield pressure signals deeper cuts ahead. Buy now or wait for $999? Full TCO analysis inside.

rtx-5080gddr7-supplygpu-pricing
GPU Price Alert: MSI Is Warning of 15-30% Hikes
Technical Report

GPU Price Alert: MSI Is Warning of 15-30% Hikes

MSI's GM warned investors of 15-30% GPU price hikes in 2026. Here's what to buy before prices move — and why the window is closing fast.

gpu-pricesmsirtx-5060-ti
DLSS 5 and What It Means for AI GPU Buyers
Technical Report

DLSS 5 and What It Means for AI GPU Buyers

DLSS 5 is exclusive to RTX 50-series Blackwell GPUs and arrives Fall 2026. Here's how it changes the buying calculus for dual-use AI and gaming builds.

dlss-5rtx-5090rtx-5070-ti
The Xiaomi Hunter Alpha Mystery
Technical Report

The Xiaomi Hunter Alpha Mystery

A nameless 1T-parameter model appeared on OpenRouter, everyone assumed it was DeepSeek V4, and they were wrong. Here's what Hunter Alpha actually was — and what it signals.

xiaomimimo-v2hunter-alpha
Technical Report

GPU Price Tracker: Best Deals This Month

Current best-value GPU deals for local LLM builds in March 2026. Where prices stand, what's overpriced, and exactly which cards to buy right now.

gpu-priceslocal-llmnvidia
Technical Report

5 LLM Milestones That Changed What Hardware You Need

Five moments in LLM development that directly shifted what GPU, RAM, and compute you need to run local models. Understanding these shifts explains the hardware landscape in 2026.

llm-historyhardwarelocal-llm
RTX 5070 Ti for Local LLMs: 896 GB/s at $749 — Worth It? — news diagram
Technical Report

RTX 5070 Ti for Local LLMs: 896 GB/s at $749 — Worth It?

The RTX 5070 Ti delivers 89% of RTX 4090 bandwidth at roughly 35% of its street price. Here's who should buy it for local LLM inference — and the 16GB VRAM ceiling to watch.

rtx-5070-tinvidiablackwell