Topic Cluster // 2 Articles
Model Optimization
01
Why Is LLM Inference Slow? A Practitioner's Guide to Fixing It
I spent three weeks in early 2024 trying to get a single 70B parameter model to respond in under two seconds. My team at SIVARO had built what we thought was...
02
Why Is LLM Inference Slow? A Practitioner’s Guide to What’s Actually Going On
The first time I deployed a large language model in production — a 7B parameter LLaMA variant, back in early 2024 — I sat staring at the latency dashboar...