I remember the exact moment I stopped paying full price for GPU inference. It was January 2025. We were running a customer-facing document extraction pipelin...