What cloud infrastructure taught me about building billing systems

For years, I thought of billing as a domain I would have to learn from scratch. Ledgers. Taxes. Proration. Regulations. Money. After 11 years building cloud datacenter management systems, I assumed my infrastructure experience would take me only so far.

I was wrong.

A few years ago, when I began working on billing, what surprised me most wasn’t how different the domain was. It was how familiar the hardest engineering problems felt. I found myself thinking about the same things I had spent years thinking about in infrastructure: state machines, lifecycle transitions, retries, partial failures, reconciliation and the gap between what a system believes and what is actually happening.

A provisioning system answers questions like: What exists? What changed? What happens when a transition fails halfway through? A billing system asks the same questions, with money attached. The domain was new. The engineering was not.

The important difference is what happens when you get the state wrong. In infrastructure, a lifecycle error can waste capacity. In billing, the same kind of error can charge a customer. That distinction turns familiar distributed-systems problems into financial ones.

Provisioning is a state machine, not an action

Nobody who has built infrastructure resource provisioning systems believes in just “create.” Create is a journey: validate the request, reserve capacity, allocate resources, configure, activate. Each step can fail, and when one fails halfway through, you are left with something that is neither quite a resource nor quite nothing. Half-created objects are a fact of life in provisioning systems at scale. The code that finds, finishes or sweeps them is where much of the real engineering lives.

Billing has the same shape. A subscription is not an event; it is a state machine: trial, active, past due, paused, canceled. The interesting bugs all live in the transitions: the upgrade that half-applied, the plan change scheduled for period end that fired twice, the trial that converted but kept its trial price. If you have ever chased a VM that exists in the database but not on any host, you already understand the subscription that bills for a plan the customer cannot see.

Consider a customer who upgrades mid-cycle. The plan change commits on the subscription record, but the entitlement service never hears about it. The customer pays for the upgrade but does not receive the corresponding access. Three weeks later, support gets a ticket. Support can see the charge but not the missing entitlement. This is the subscription twin of the VM that exists in the database but not on any host: the create journey got halfway and nobody swept it.

The lesson is broader than either domain: lifecycle transitions are where distributed systems become difficult.

Create and delete must be idempotent

In provisioning, retry safety is survival. A client that times out will retry, and if your create operation is not idempotent, you get two of everything. Delete is worse: a retried delete must be safe against a resource that is already partially gone. I spent years making sure that an operation applied twice, out of order or after a crash, would still converge on the same end state.

Billing depends on the same rule. Charge attempts are retried after a deployment. Invoice generation restarts mid-run. An event arrives twice because a queue redelivered it. The operation does not need to succeed exactly once. It needs to produce the correct result even when the system tries more than once. The systems that survive are the ones where every record has a stable identity and reprocessing is boring.

That last property is underrated. In a distributed system, retries, replay and recovery should not be exceptional paths. They should be ordinary operations that the system can safely perform again.

Stopping is harder than starting

The dirty secret of infrastructure is that deprovisioning can be harder than provisioning. Creates get attention because someone is waiting for them. Deletes fail quietly: a reference someone forgot, a dependency that will not release, a cleanup job that errors at 3 a.m. and pages no one. The result is orphaned state: resources that exist, consume capacity and belong to no one.

Billing has its own version of the orphan: a resource torn down on the 14th but billed for the entire month. A seat that is removed but keeps charging is not an arithmetic bug. It is a lifecycle bug: a stop event was lost or a transition never fired. The math is fine. The state is wrong.

Consider a team lead who removes a seat on the 9th. The remove call times out. The UI shows the seat as gone, and everyone moves on. But the stop event never landed, so the seat keeps billing. Nobody notices until renewal season, when the customer’s finance team reconciles the invoice against headcount and asks why eleven seats cost more than the nine people they employ. The fix is

[…]
Content was trimmed to protect the source. Please visit the original article for the full text.

This article has been indexed from InfoWorld

Read the original article: