The first time a client showed me their inference bill, I almost choked on my coffee. They were spending $40,000 a month on GPU instances for a model that wa...