CraftRigs
GT

Georgia Thomas

Setup Guides · Troubleshooting · Workflows · Benchmarks Austin, TX

Most local AI setup guides assume things just work. They don't. ROCm refuses to detect your GPU, Ollama throws cryptic CUDA errors, and the GitHub issue thread is 200 comments deep with no resolution.

Georgia writes the guides that end the debugging spiral, step-by-step, tested on the hardware CraftRigs readers actually own. She tests on the same used RTX 3090s and budget builds her readers run, because a guide that only works on a $6,000 workstation isn't a guide.

Editorial disclosure: Georgia is an editorial persona of the CraftRigs AI-assisted editorial team — a consistent beat and methodology, not an individual human reviewer. How our research and sourcing works: How CraftRigs Works.
Setup Guides Troubleshooting Workflows Benchmarks
267 Articles Published
227 Setup Guides
Nov 2025 Member Since

Latest from Georgia

267 articles
70B Model Alternatives 2026: Qwen vs DeepSeek vs Scout — diagram
Guide

70B Model Alternatives 2026: Qwen vs DeepSeek vs Scout

Dense 70B needs 48 GB VRAM and loses to MoE alternatives—Qwen 3.6-27B hits 86.2 MMLU-Pro at 12 GB, Scout runs 128K context at 20 GB. Match model to GPU tier, not parameter count.

Jul 11, 2026
Apple Silicon LLM Speed: M5 Max vs M4 Max vs Mac Studio — diagram
Guide

Apple Silicon LLM Speed: M5 Max vs M4 Max vs Mac Studio

Wrong Mac kills your 70B model—M5 Max and M4 Max tie at 75 tok/s for 7B, but only M3 Ultra Mac Studio hits 10–12 tok/s on 70B. Memory bandwidth, not GPU cores, decides what you can run. Pick chip by model size, not benchmark hype.

Jul 11, 2026
$500 Local LLM PC: 24 GB VRAM for $320 Used — diagram
Guide

$500 Local LLM PC: 24 GB VRAM for $320 Used

New 8 GB GPUs can't run 70B models—used RTX 3090 24 GB hits 75 MB/VRAM per dollar at $320, beats RX 9060 XT 2:1, and resells at 65%. Build the $500 system that outperforms $550 new.

Jul 11, 2026
GPU Market 2026: $299 vs $349 LLM Buy — diagram
Guide

GPU Market 2026: $299 vs $349 LLM Buy

RTX 5060's 8GB traps local LLM buyers at 13B models—RX 9060 XT 24GB at $349 unlocks 70B inference for $14.54/GB. Used 3090s hit $240, but 350W TDP burns savings. Pick your tier, know the CUDA trade-off.

Jul 11, 2026
Wrong Quant Format Won't Load—Here's the Fix — diagram
Guide

Wrong Quant Format Won't Load—Here's the Fix

34% of model load failures are format mismatches, not broken files. Match GGUF/GPTQ/AWQ/EXL2/MLX to your runtime before you download—Apple Silicon hits 6.8 tok/s with MLX, RTX 4090 pushes 28 tok/s with EXL2, but wrong format fails entirely.

Jul 11, 2026
Qwen 3.6-27B vs Llama 70B: Smarter, Not Bigger — diagram
Guide

Qwen 3.6-27B vs Llama 70B: Smarter, Not Bigger

Bigger isn't better anymore—Qwen 3.6-27B hits 78.4 MMLU-Pro at 16.2GB VRAM, runs 3.5× faster on a single RTX 4090, and costs 2.1× less than Llama 3.3 70B. Match model to use case, not parameter count.

Jul 11, 2026
RTX 5060 vs RX 9060 XT: Which $299 GPU Runs LLMs? — diagram
Guide

RTX 5060 vs RX 9060 XT: Which $299 GPU Runs LLMs?

8GB VRAM hard-caps at 7B parameters—RX 9060 XT's 16GB runs 13B fully in VRAM for $349, but CUDA's zero-config beats ROCm's setup friction. Pick the stack you can live with.

Jul 11, 2026
LLM TPS Benchmarking: Measure Real Speed, Not Marketing — diagram
Guide

LLM TPS Benchmarking: Measure Real Speed, Not Marketing

Published tokens-per-second figures lie by 40–60% from warm-cache hero runs. Fix prompt length to 512 tokens, measure cold-start TTFT, and use llama-bench -p 512 -n 128 to get reproducible numbers that actually match your hardware. Stop guessing, start measuring.

Jul 11, 2026
VRAM Too Small? Fit 70B Models With the Right Quant — diagram
Guide

VRAM Too Small? Fit 70B Models With the Right Quant

8GB GPU hitting OOM on 7B models? Q4_K_M fits 70B on 48GB, MoE cuts active params 75%, and the wrong quant costs 40% accuracy. Match model, quant, and budget tier—here's the exact table.

Jul 11, 2026
70B CPU Inference: 2–7 tok/s Without Buying GPUs — diagram
Guide

70B CPU Inference: 2–7 tok/s Without Buying GPUs

Your Threadripper already owns the 70B path—CPU-only hits 2.1–7.1 tok/s on DDR5 bandwidth, beats dual RTX 3090 cost for batch jobs, and leaves GPU free. Honest benchmarks, no GPU required.

May 22, 2026
ROCm 7.2 GPU Matrix 2026: Windows vs Linux — diagram
Guide

ROCm 7.2 GPU Matrix 2026: Windows vs Linux

Wrong driver kills AMD ROCm on Windows—ROCDXG not Adrenalin, clean install required. Full consumer GPU matrix with Linux native, WSL2 grades, and honest tok/s numbers. Install right, skip the forum archaeology.

May 22, 2026
Mac Studio Axed, M5 Ultra Delayed: Buy M4 Now? — diagram
Guide

Mac Studio Axed, M5 Ultra Delayed: Buy M4 Now?

Mac Studio 128GB killed in May 2026, M5 Ultra pushed to Q4—M4 Max 96GB is your only 70B Q8_0 option until 2027. We map every tier, benchmark AMD's Strix Halo rival, and tell you whether to buy or wait.

May 22, 2026
RTX 5060–5090 Street Prices Exposed: Buy or Wait? — diagram
Guide

RTX 5060–5090 Street Prices Exposed: Buy or Wait?

NVIDIA cut RTX 50-series production 40%—street prices run 18–67% over MSRP. RTX 5060 Ti 16GB at $485 beats the stack, but used RTX 3090 at $500 undercuts everything. Match card to budget before Computex hype resets the board.

May 22, 2026
Arc B580 Stuck at 13 tok/s? IPEX-LLM Unlocks 70 — diagram
Guide

Arc B580 Stuck at 13 tok/s? IPEX-LLM Unlocks 70

Your $250 Arc B580 runs Mistral 7B at 13 tok/s via Vulkan—leaving XMX acceleration idle. IPEX-LLM Docker hits 70 tok/s with full setup steps for Linux and Windows. Stop running at 20% speed.

May 22, 2026
LM Studio 0 GPUs: Fix for CUDA, ROCm, WSL2 (2026) — diagram
Guide

LM Studio 0 GPUs: Fix for CUDA, ROCm, WSL2 (2026)

LM Studio shows 0 GPUs detected? NVIDIA CUDA 12.2, AMD ROCDXG May 2026 driver, and WSL2 passthrough each have distinct fixes—most are driver mismatches, not hardware failure. Run the 60-second pre-flight, then follow your platform path.

May 22, 2026
Open WebUI Pipelines: Wire Any LLM Backend (2026) — diagram
Guide

Open WebUI Pipelines: Wire Any LLM Backend (2026)

Native Ollama locks you to one machine—Pipelines unlocks llama.cpp, remote Ollama, and custom APIs in the same chat UI. New April 2026 Desktop App auto-recovers GPU crashes. Wire your backend once, switch models freely.

May 22, 2026
VRAM Shortage 2026: Buy Used, Not New — diagram
Guide

VRAM Shortage 2026: Buy Used, Not New

HBM and GDDR7 shortages keep RTX 50-series prices 40% above MSRP. Used RTX 3090 24GB at $500 beats new cards on $/GB-VRAM. Here's how to navigate the chaos and buy smart.

May 22, 2026
GMKtec EVO-X2 Memory Bandwidth: 256 GB/s, Qwen 3.6 Speed — diagram
Guide

GMKtec EVO-X2 Memory Bandwidth: 256 GB/s, Qwen 3.6 Speed

EVO-X2: 256 GB/s unified memory (273 GB/s observed). Delivers 14–18 tok/s on Qwen 3.6 35B Q4_K_M. Compare Mac Mini M4: 120 GB/s. See specs, thermals, BIOS quirks, $1,500–$2,200 pricing, and when to buy.

May 8, 2026