
Modeling Billing So the Code Matches the Pricing Page, Not the Other Way Around
A base-fee-plus-seat-overage pricing model needs two Stripe subscription items, not one; webhooks confirm state changes but never decide them; and a single four-state subscription machine replaces a separate `isActive` flag that could quietly disagree with it
Part 3 of 7 — InvocaCare Architecture series
Pricing pages look simple. "$79/month, up to 200 calls, includes your first seats, $X per extra seat after that." One sentence, easy to read, easy to sell. Then you sit down to model it in Stripe and discover that "simple pricing" and "simple billing code" are not the same claim at all — and every shortcut you take on the second one eventually shows up as a support ticket, a double charge, or an argument about what a customer actually owes.
Here's how InvocaCare's billing model turned out, and which shortcuts I took on purpose versus which ones I ruled out after actually trying them.
Decision: two subscription items, not one
The obvious way to model "base fee plus per-seat overage" in Stripe is a single subscription item, quantity equal to seat count, price scaled accordingly. I tried it. It can't actually express the pricing: a flat base fee that doesn't move with seat count, plus an overage that only kicks in past an included-seat threshold, isn't a single linear quantity-times-price relationship. You either lose the base fee or lose the free included seats.
So every InvocaCare subscription carries exactly two Stripe subscription items:
| Item | Role | Quantity |
|---|---|---|
| Base | flat monthly tier fee | always 1 |
| Seat | per-additional-seat overage | max(0, seats − includedSeats) |
Each item is tagged with its own metadata (role: 'base' / role: 'seat'), so the code that reconstructs a subscription record from Stripe's API response can identify each item reliably regardless of array order — a small detail, but the kind of thing that turns into a confusing bug six months later if you skip it.
I also looked at using Stripe Checkout Sessions for tier changes — upgrade/downgrade via a hosted checkout flow — and ruled it out for the same reason: it creates a window where there's no active subscription, and it can't update seat quantities on an existing one. Tier and seat changes happen through direct Stripe API calls instead, kept in sync with our own records synchronously.
Decision: webhooks confirm, they don't decide
The tempting pattern is to let Stripe webhooks drive your billing state — a tier changes, Stripe fires an event, your handler reacts. I didn't build it that way. Outbound calls from InvocaCare's billing service update Stripe and DynamoDB synchronously, in the same operation that initiated the change. Webhooks exist to confirm that the action actually completed and to correct any divergence — they never initiate a tier change themselves.
The reasoning: webhook delivery is asynchronous and, occasionally, out of order. If webhooks are your source of truth for decisions, you've built a system where a delayed or reordered event can put a tenant's access into the wrong state for however long that delay lasts. Reconciliation is exactly what webhooks are good at — it's just a different job than decision-making, and conflating the two was the actual mistake I wanted to avoid.
Decision: one state machine, no separate isActive flag
Early on there was a boolean isActive flag living alongside a Stripe-derived subscription status. Two sources of truth for the same underlying question — can this tenant use the product right now — is a bug waiting to happen; they will eventually disagree, and whichever one your access-control check reads becomes the wrong answer for someone.
It's now a single four-state machine:
trial → active → suspended → cancelled
trial and active allow API access; suspended and cancelled block it. suspended fires specifically on a third consecutive payment failure — read directly from Stripe's invoice.attempt_count, not a hand-rolled local counter that has to stay in sync with Stripe's own retry logic. I also pulled past_due out of this state machine entirely; it's a Stripe billing concept, not a fact about whether a tenant should be let in the door, and modeling it as a tenant-lifecycle state was importing Stripe's vocabulary into a place it didn't actually belong.
Decision: what I deliberately didn't build
Two things were on the initial spec and got cut, on purpose, for v1:
A custom billing portal. Sign-up creates a Stripe customer automatically, and a tenant can already manage payment methods and view invoice history through Stripe's own hosted Billing Portal, wired up behind a single createPortalSession call — no custom UI needed for that. What's still deliberately out is a self-serve plan change through that portal: a tier or seat change stays a direct API call into our own state machine, for the same reason webhooks don't get to decide — I want one place, not two, that can put a tenant's plan in the wrong state.
A revenue dashboard. Stripe's dashboard already reports MRR, churn, and tier breakdown — better than a first version of a homegrown one would. The one gap SystemAdmin actually needed was a single narrow read endpoint for one tenant's subscription state, not a general analytics surface. Building the dashboard would have been building a worse copy of a tool I already had.
Both cuts follow the same rule: don't build the general version of something Stripe already gives you for free, until there's a specific gap Stripe doesn't cover.
The throughline
None of these are exotic decisions individually — idempotency keys on every mutating Stripe call so retries can't double-charge, lookup keys instead of hardcoded price IDs so a pricing change is zero-code, item IDs cached in our own records so we're not re-fetching from Stripe before every seat update. What ties them together is a single instinct: model billing state so there's exactly one place that can be wrong, and make that place the one that's easiest to get right and cheapest to verify. Every one of the alternatives I ruled out — the single-item pricing model, webhook-driven decisions, the extra isActive flag, the early dashboard — failed that test in a different way, and the fastest way to find that out was to build the wrong version first and notice specifically why it didn't hold up.