· 3 min read

Charging for AI generations without double-charging

by · Project: DRAPAI

#drapai #postgres #architecture

DRAPAI is an AI virtual try-on app: you upload a photo of a person and a garment, and it generates a preview of that outfit on that person. Each generation costs a token. That sounds simple until you look at the path a request takes:

mobile app → backend → workflow engine → AI model → back to the database

Every arrow can fail, and some fail ambiguously. If a call times out, the generation may still finish a minute later. Charge too early and users pay for failures. Charge too late and a retry can produce a free generation. This post is about the ideas that keep the token balance correct.

A job is a state machine, enforced by the database

Every generation is a job that moves in one direction:

pending → processing → succeeded | failed

The database itself rejects any transition that is not on that path, and a finished job can never move back. The app only treats finished jobs as final. Putting the rule in the database means no service, workflow or future bug can bend it.

Reserve, then settle

Tokens are not deducted when the user taps Generate. They are held:

  1. When a job is created, a token is reserved for it, after checking that the caller owns the job.
  2. The generation is dispatched, with ownership checked again on the server.
  3. If the job succeeds, the hold is committed. If it fails, the hold is released.

Reserving and dispatching are both idempotent per job, so a double tap or a network retry costs nothing extra.

Timeouts are not failures

An early version treated a timed-out dispatch as a failed job. That was wrong: the work can keep running after the HTTP call gives up, so a job could be refunded and still produce an image.

Now a timeout means "unknown", not "failed". A scheduled reconciliation job periodically settles holds for finished jobs and resolves jobs that have been stuck for too long. How long counts as "too long" depends on where the job got stuck: a job that was never picked up is resolved sooner than one that is known to be running, and a job whose dispatch outcome is unknown gets the most patience before it is refunded.

Keeping privileged operations out of reach

The operations that move tokens need more rights than a normal user has. They are kept out of the client-facing API entirely, and clients reach them only through thin entry points that check ownership first.

Getting this right took a few attempts. One change that looked stricter on paper broke every real call, because the entry point no longer had the rights it needed underneath. The lesson: when you tighten permissions, test end to end from a real user session, not only by checking that the migration applied.

Subscriptions reset balances instead of adding to them

Paid plans grant a monthly token allowance through the App Store and Google Play. Both stores can report the same renewal more than once. Adding tokens on every report would duplicate them, so each billing period resets the balance to the plan amount, keyed by the store's own transaction ID. Seeing the same ID again changes nothing.

Free tokens attract abuse

New users get a few free tokens to try the app. Free generations on a paid AI model invite people to create accounts in bulk, so the grant has conditions: it is tied to the user's identity rather than to a fresh account, it requires finishing onboarding, and signups are rate limited. When a check cannot decide, it fails closed.

What I would keep for the next project

  • Put state rules in the database, not in every caller.
  • Hold, then settle. Never charge on the way in.
  • Treat a timeout as "unknown", and have a scheduled job resolve unknowns.
  • Make grants idempotent on an ID the payment provider controls.