Model Inference
Long Context Inference: CPU Memory Bandwidth Bottleneck
--- March 2025. We're running a document-embedding pipeline for a mid-size legal firm. 200K-token contexts. 70B parameter model. The GPUs are 40%% utilization...
How to Benchmark Long Context CPU Inference
So you've got a model that reads a 500-page document and you're wondering if your CPU server can handle it. I've been there. In March of this year, a fintech...
CPU vs GPU for Long Context Inference: Throughput Benchmarks That Actually Matter
It's September 2026. Context windows are no longer a talking point. They're 1M tokens on commodity hardware, and 10M-100M token experiments are happening dai...
Best CPU for Long Context LLM Inference in 2026
You're staring at a 100,000-token context window crawling at 4 tokens per second, wondering why your $4,000 GPU rig feels like a 1998 dial-up modem. I've bee...
How to Reduce Inference Cost with Model Architecture (2026 Buyer's Guide)
Last month, a fintech client sent me their AWS bill. They were spending $84,000 a month on inference for a single fraud-detection model. Their CTO looked at ...