LLMs — Navigating the Paradox of Cost Saving
A reflection on the Cobra Effect and why optimising a visible AI cost can create less visible costs elsewhere.

Avoiding a visible AI subscription cost can create a larger, less visible infrastructure bill. The useful comparison is not “paid model versus free model”; it is the total cost of running the workflow at the scale and reliability you need.
The apparent saving
When I wrote the original version of this essay in February 2024, I wanted the productivity benefits of language models without paying the monthly fee attached to a premium service. The obvious alternative seemed to be an open-source model: avoid the direct subscription and keep the capability.
Running that model efficiently introduced another requirement—a machine with a suitable GPU. Renting GPU capacity in the cloud appeared to solve the hardware problem, but its usage cost could quickly become greater than the subscription I had set out to avoid.
The Cobra Effect in technology choices
This is the pattern behind the Cobra Effect: an intervention meant to solve a problem can worsen it because the incentive focuses attention on the wrong measure. In my case, the visible monthly fee became the target. Eliminating that one line item did not eliminate the cost of computing.
The experience is a reminder to widen the frame before calling something a saving. A model licence or subscription is only one part of a system. Hardware, cloud usage, operations, and the shape of actual demand all affect the result.
A usage-based middle path
My conclusion was to keep the flexibility of open-source models while consuming them on a usage-based basis. That approach avoids owning idle GPU capacity and lets cost scale with actual use. It preserves the ability to choose an open model without pretending that the infrastructure needed to run it is free.
This is not a universal answer for every workload. It is a way to match the economic model to the demand: look at what the full solution costs, then choose the combination of model and infrastructure that avoids creating a larger problem in the name of solving a smaller one.
This essay was first published on LinkedIn on 8 February 2024. The pricing comparison above reflects the context in which it was originally written.