What cloud infrastructure taught me about building billing systems
InfoWorld ·

For years, I thought of billing as a domain I would have to learn from scratch. Ledgers. Taxes. Proration. Regulations. Money. After 11 years building cloud datacenter management systems, I assumed my infrastructure experience would take me only so far. I was wrong. A few years ago, when I began working on billing, what surprised me most wasn’t how different the domain was. It was how familiar the hardest engineering problems felt. I found myself thinking about the same things I had spent years thinking about in infrastructure: state machines, lifecycle transitions, retries, partial failures, reconciliation and the gap between what a system believes and what is actually happening. A provisioning system answers questions like: What exists? What changed? What happens when a transition fails halfway through? A billing system asks the same questions, with money attached. The domain was new. The engineering was not. The important difference is what happens when you get the state wrong. In infrastructure, a lifecycle error can waste capacity. In billing, the same kind of error can charge a customer. That distinction turns familiar distributed-systems problems into financial ones. Provisioning is a state machine, not an action Nobody who has built infrastructure resource provisioning systems believes in just “create.” Create is a journey: validate the request, reserve capacity, allocate resources, configure, activate. Each step can fail, and when one fails halfway through, you are left with something that is neither quite a resource nor quite nothing. Half-created objects are a fact of life in provisioning systems at scale. The code that finds, finishes or sweeps them is where much of the real engineering lives. Billing has the same shape. A subscription is not an event; it is a state machine: trial, active, past due, paused, canceled. The interesting bugs all live in the transitions: the upgrade that half-applied, the plan change scheduled for period end that fired twice, the trial that converted but kept its trial price. If you have ever chased a VM that exists in the database but not on any host, you already understand the subscription that bills for a plan the customer cannot see. Consider a customer who upgrades mid-cycle. The plan change commits on the subscription record, but the entitlement service never hears about it. The customer pays for the upgrade but does not receive the corresponding access. Three weeks later, support gets a ticket. Support can see the charge but not the missing entitlement. This is the subscription twin of the VM that exists in the database but not on any host: the create journey got halfway and nobody swept it. The lesson is broader than either domain: lifecycle transitions are where distributed systems become difficult. Create and delete must be idempotent In provisioning, retry safety is survival. A client that times out will retry, and if your create operation is not idempotent , you get two of everything. Delete is worse: a retried delete must be safe against a resource that is already partially gone. I spent years making sure that an operation applied twice, out of order or after a crash, would still converge on the same end state. Billing depends on the same rule. Charge attempts are retried after a deployment. Invoice generation restarts mid-run. An event arrives twice because a queue redelivered it. The operation does not need to succeed exactly once . It needs to produce the correct result even when the system tries more than once. The systems that survive are the ones where every record has a stable identity and reprocessing is boring. That last property is underrated. In a distributed system, retries, replay and recovery should not be exceptional paths. They should be ordinary operations that the system can safely perform again. Stopping is harder than starting The dirty secret of infrastructure is that deprovisioning can be harder than provisioning. Creates get attention because someone is waiting for them. Deletes fail quietly: a reference someone forgot, a dependency that will not release, a cleanup job that errors at 3 a.m. and pages no one. The result is orphaned state: resources that exist, consume capacity and belong to no one. Billing has its own version of the orphan: a resource torn down on the 14th but billed for the entire month. A seat that is removed but keeps charging is not an arithmetic bug. It is a lifecycle bug: a stop event was lost or a transition never fired. The math is fine. The state is wrong. Consider a team lead who removes a seat on the 9th. The remove call times out. The UI shows the seat as gone, and everyone moves on. But the stop event never landed, so the seat keeps billing. Nobody notices until renewal season, when the customer’s finance team reconciles the invoice against headcount and asks why eleven seats cost more than the nine people they employ. The fix is a refund and an apology. The root cause is a delete that failed quietly because nobody was waiting on it. Starting something creates an expectation; stopping it creates an obligation to clean up every piece of state it created. Drift is the default. Reconciliation is the fix Every control plane drifts from reality. The system believes a cluster has forty free slots. The cluster disagrees. So, you build reconciliation loops that continuously compare desired state with actual state and repair the difference. Not once a quarter. Continuously. Billing drifts too, between what the contract believes is active and what the invoice says. The fix is similar: reconciliation as a first-class system, continuously comparing what was provisioned, what was used and what was charged, with explicit thresholds for investigation. A reconciliation report nobody reads is a spreadsheet. A reconciliation loop with teeth is how the number stays true. I saw the infrastructure version repeatedly: a team provisions a test cluster for a two-week experiment, the experiment ends, and the teardown job errors at 3 a.m. and pages no one. The cluster quietly burns budget for months until a capacity review finds it. Billing has the same class of problem. A lifecycle transition can leave an invoice line item without a living entitlement behind it. A reconciliation system can detect that inconsistency before the customer does. Same orphan, different currency. Snapshots tell you what. Events tell you why A state snapshot tells you what exists now. An event log tells you how it got that way. A provisioning system that keeps only snapshots cannot answer, “when did this change?” A system that keeps only events cannot answer, “what do I have?” without replaying history. You need both. Billing needs both answers constantly. Proration is essentially the question “when did this change?” asked with money on the line. When someone asks, “Why was I charged this amount?” the answer depends on what changed, when it changed and which lifecycle transitions occurred. This is why billing systems cannot treat history as an implementation detail. The current state explains what the system believes today. The event history explains how it got there. Both are necessary when the system needs to explain a number of months later. Correctness is not the same everywhere Correctness has different consequences in these domains even when the underlying engineering patterns are similar. In infrastructure, correctness might mean that a desired resource eventually exists in the desired state. A temporary inconsistency may consume capacity. An orphaned resource may waste money. A cleanup process can often reclaim it later. There is room for bounded uncertainty. Billing needs something stronger. The amount charged must be correct, reproducible, explainable and traceable to the underlying lifecycle. The system cannot simply know the right answer. It needs to be able to explain why the answer is right. That distinction changes the engineering discipline. The distributed-systems principles remain familiar: retries happen, events arrive more than once, components fail independently and state drifts. But the required precision changes completely. An infrastructure system might tolerate a cleanup window. An invoice cannot. Money numbers do not have error bars. What transfers The engineering principles transfer surprisingly well. Idempotency as a first principle. Design operations so retries and reprocessing converge safely. Lifecycle as a state machine. Treat transitions as first-class engineering concerns because that is where partial failures and inconsistent state appear. Reconciliation as a continuous loop. Assume systems drift and continuously compare what should be true with what is actually true. Corrections as new records. When something is wrong, record the correction rather than silently rewriting history. The system should preserve enough information to explain how it arrived at the current state. The most valuable skill is not memorizing every domain’s vocabulary. It is recognizing the underlying system. The cost of being wrong changes completely. An orphaned VM can eventually be reclaimed. A wrong invoice has a customer on the other side of it. The hard part is not learning that state matters. It is learning that this time, the state is someone’s money.
For years, I thought of billing as a domain I would have to learn from scratch. Ledgers. Taxes. Proration. Regulations. Money. After 11 years building cloud datacenter management systems, I assumed my infrastructure experience would take me only so far. I was wrong. A few years ago, when I began working on billing, what surprised me most wasn’t how different the domain was. It was how familiar the hardest engineering problems felt. I found myself thinking about the same things I had spent years thinking about in infrastructure: state machines, lifecycle transitions, retries, partial failures, reconciliation and the gap between what a system believes and what is actually happening. A provisioning system answers questions like: What exists? What changed? What happens when a transition fails halfway through? A billing system asks the same questions, with money attached. The domain was new. The engineering was not. The important difference is what happens when you get the state wrong. In infrastructure, a lifecycle error can waste capacity. In billing, the same kind of error can charge a customer. That distinction turns familiar distributed-systems problems into financial ones. Provisioning is a state machine, not an action Nobody who has built infrastructure resource provisioning systems believes in just “create.” Create is a journey: validate the request, reserve capacity, allocate resources, configure, activate. Each step can fail, and when one fails halfway through, you are left with something that is neither quite a resource nor quite nothing. Half-created objects are a fact of life in provisioning systems at scale. The code that finds, finishes or sweeps them is where much of the real engineering lives. Billing has the same shape. A subscription is not an event; it is a state machine: trial, active, past due, paused, canceled. The interesting bugs all live in the transitions: the upgrade that half-applied, the plan change scheduled for period end that fired twice, the trial that converted but kept its trial price. If you have ever chased a VM that exists in the database but not on any host, you already understand the subscription that bills for a plan the customer cannot see. Consider a customer who upgrades mid-cycle. The plan change commits on the subscription record, but the entitlement service never hears about it. The customer pays for the upgrade but does not receive the corresponding access. Three weeks later, support gets a ticket. Support can see the charge but not the missing entitlement. This is the subscription twin of the VM that exists in the database but not on any host: the create journey got halfway and nobody swept it. The lesson is broader than either domain: lifecycle transitions are where distributed systems become difficult. Create and delete must be idempotent In provisioning, retry safety is survival. A client that times out will retry, and if your create operation is not idempotent , you get two of everything. Delete is worse: a retried delete must be safe against a resource that is already partially gone. I spent years making sure that an operation applied twice, out of order or after a crash, would still converge on the same end state. Billing depends on the same rule. Charge attempts are retried after a deployment. Invoice generation restarts mid-run. An event arrives twice because a queue redelivered it. The operation does not need to succeed exactly once . It needs to produce the correct result even when the system tries more than once. The systems that survive are the ones where every record has a stable identity and reprocessing is boring. That last property is underrated. In a distributed system, retries, replay and recovery should not be exceptional paths. They should be ordinary operations that the system can safely perform again. Stopping is harder than starting The dirty secret of infrastructure is that deprovisioning can be harder than provisioning. Creates get attention because someone is waiting for them. Deletes fail quietly: a reference someone forgot, a dependency that will not release, a cleanup job that errors at 3 a.m. and pages no one. The result is orphaned state: resources that exist, consume capacity and belong to no one. Billing has its own version of the orphan: a resource torn down on the 14th but billed for the entire month. A seat that is removed but keeps charging is not an arithmetic bug. It is a lifecycle bug: a stop event was lost or a transition never fired. The math is fine. The state is wrong. Consider a team lead who removes a seat on the 9th. The remove call times out. The UI shows the seat as gone, and everyone moves on. But the stop event never landed, so the seat keeps billing. Nobody notices until renewal season, when the customer’s finance team reconciles the invoice against headcount and asks why eleven seats cost more than the nine people they employ. The fix is a refund and an apology. The root cause is a delete that failed quietly because nobody was waiting on it. Starting something creates an expectation; stopping it creates an obligation to clean up every piece of state it created. Drift is the default. Reconciliation is the fix Every control plane drifts from reality. The system believes a cluster has forty free slots. The cluster disagrees. So, you build reconciliation loops that continuously compare desired state with actual state and repair the difference. Not once a quarter. Continuously. Billing drifts too, between what the contract believes is active and what the invoice says. The fix is similar: reconciliation as a first-class system, continuously comparing what was provisioned, what was used and what was charged, with explicit thresholds for investigation. A reconciliation report nobody reads is a spreadsheet. A reconciliation loop with teeth is how the number stays true. I saw the infrastructure version repeatedly: a team provisions a test cluster for a two-week experiment, the experiment ends, and the teardown job errors at 3 a.m. and pages no one. The cluster quietly burns budget for months until a capacity review finds it. Billing has the same class of problem. A lifecycle transition can leave an invoice line item without a living entitlement behind it. A reconciliation system can detect that inconsistency before the customer does. Same orphan, different currency. Snapshots tell you what. Events tell you why A state snapshot tells you what exists now. An event log tells you how it got that way. A provisioning system that keeps only snapshots cannot answer, “when did this change?” A system that keeps only events cannot answer, “what do I have?” without replaying history. You need both. Billing needs both answers constantly. Proration is essentially the question “when did this change?” asked with money on the line. When someone asks, “Why was I charged this amount?” the answer depends on what changed, when it changed and which lifecycle transitions occurred. This is why billing systems cannot treat history as an implementation detail. The current state explains what the system believes today. The event history explains how it got there. Both are necessary when the system needs to explain a number of months later. Correctness is not the same everywhere Correctness has different consequences in these domains even when the underlying engineering patterns are similar. In infrastructure, correctness might mean that a desired resource eventually exists in the desired state. A temporary inconsistency may consume capacity. An orphaned resource may waste money. A cleanup process can often reclaim it later. There is room for bounded uncertainty. Billing needs something stronger. The amount charged must be correct, reproducible, explainable and traceable to the underlying lifecycle. The system cannot simply know the right answer. It needs to be able to explain why the answer is right. That distinction changes the engineering discipline. The distributed-systems principles remain familiar: retries happen, events arrive more than once, components fail independently and state drifts. But the required precision changes completely. An infrastructure system might tolerate a cleanup window. An invoice cannot. Money numbers do not have error bars. What transfers The engineering principles transfer surprisingly well. Idempotency as a first principle. Design operations so retries and reprocessing converge safely. Lifecycle as a state machine. Treat transitions as first-class engineering concerns because that is where partial failures and inconsistent state appear. Reconciliation as a continuous loop. Assume systems drift and continuously compare what should be true with what is actually true. Corrections as new records. When something is wrong, record the correction rather than silently rewriting history. The system should preserve enough information to explain how it arrived at the current state. The most valuable skill is not memorizing every domain’s vocabulary. It is recognizing the underlying system. The cost of being wrong changes completely. An orphaned VM can eventually be reclaimed. A wrong invoice has a customer on the other side of it. The hard part is not learning that state matters. It is learning that this time, the state is someone’s money.