TL;DR
- Transaction volume alone doesn't predict when reconciliation breaks; the real driver is the interaction of transaction count, provider count, live contract versions, and how many days the close allows.
- Four walls hit in order as load grows: sampling (which forfeits unsampled recoverable charges), a compressing batch window, a growing exception backlog, and contract sprawl across amendments and effective dates.
- The exception backlog is the wall that actually matters: once median exception age passes the shortest contractual dispute window, unworked overcharges expire and become permanent write-offs rather than delays.
- The diagnostic questions are concrete: what share of transactions were fully repriced (coverage), median exception age versus dispute window, expired value last quarter, and how many live contract versions are in force.
Short answer
A million transactions a month is a plausible-sounding threshold, and it is the wrong number.
Volume on its own predicts almost nothing. What breaks a verification operation is the interaction of four things: how many transactions you process, how many providers bill you, how many contract versions are live, and how many days the close gives you. A smaller business billing across five providers can be under more strain than a much larger one on a single contract.
Four walls arrive as that load grows, and they arrive in a predictable order. Only one of them turns a delay into a permanent loss.
Wall one: sampling
The first thing a growing team gives up is completeness.
It happens quietly. Checking every transaction stops being feasible, so somebody proposes checking a representative slice instead, and that is a sensible response to a real constraint. Statistically the slice is fine. Commercially it is a decision to stop recovering money.
A sample can establish that a systematic error exists. It cannot establish what that error cost you, because a provider refunds what you evidence, at the transaction level, and it has no obligation to extrapolate your sample on your behalf. Every transaction outside the sample is an unevidenced charge.
So the day sampling begins, recoverable value starts leaking by design. The finding gets sharper and the claim gets smaller.
Wall two: the batch window
Provider files arrive on the provider's schedule. Your close happens on the calendar's schedule. Verification has to fit in the gap between them.
That gap is fixed. It does not lengthen because your volume grew. So as processing time rises, the window compresses, and something gets dropped to make the deadline. Usually it is the slowest check, which is usually the most valuable one, because the checks that take longest are the ones that reprice a transaction properly rather than compare two totals.
There is a second-order version of this that catches people out. Providers reissue files. A restatement three days before close either forces a rerun you do not have time for, or gets skipped and quietly invalidates the run you already did.
The tell that you have hit this wall is a conversation about which checks to turn off during close week.
Wall three: the exception backlog
This is the wall that matters, and it is the one most operations misread as a staffing problem.
Exception volume scales with transaction volume. Headcount does not. So a queue that was worked daily becomes a queue that is worked weekly, then a queue with a backlog, then a backlog with a queue attached.
A backlog looks like a delay. It is not, and the reason is contractual. Provider agreements cap how long you have to dispute a charge. That clock is not yours to reset, and it runs while the exception sits unworked. An overcharge that expires in the queue is not deferred revenue. It is gone, with the evidence intact and the entitlement expired.
So the real threshold is not a transaction count. It is the day your median exception age passes your shortest dispute window. Before that day, a backlog costs you working capital timing. After it, the backlog is a write-off machine, and it will keep running quietly because nothing in a normal finance dashboard reports it.
Most teams cannot answer either half of that comparison. They do not know their median exception age, and they have not read the dispute window in every provider contract they hold.
Wall four: contract sprawl
The last wall is combinatorial, and it arrives without any change in transaction volume at all.
Each provider brings a contract. Each contract accumulates amendments, each with its own effective date. Ten providers with four amendments each is forty rate regimes, and any given transaction has to be priced under exactly one of them, chosen by its settlement date.
Past that point the failures change character. They stop being arithmetic mistakes and become knowledge failures: a term nobody remembered, a threshold nobody noticed had moved, a corridor priced under last year's schedule. The person who modelled the original contract has usually moved on.
This is also where an expected cost gets genuinely hard to compute, because contract terms interact and the order you apply them in changes the answer.
What each wall costs, in CFO terms
None of these show up as a line item, which is exactly why they persist.
| Wall | What it costs you |
|---|---|
| Sampling | The difference between the leakage you can prove and the leakage you found. You keep the insight and lose the claim. |
| Batch window | The checks you disable under deadline pressure, which are the ones that find priced-wrong rather than missing. |
| Exception backlog | Recoverable money that expires. This is the only one of the four that is permanent. |
| Contract sprawl | Errors that persist for quarters because nobody is aware the governing term exists. |
The unrecovered figure never appears in a headcount plan, and it is usually larger than the headcount question it competes with. We have set out that comparison in the real cost of a fintech reconciliation team.
What to measure instead of transaction volume
Four numbers. A finance leader can ask for all four this week, and the answers are diagnostic on their own.
Coverage. What share of last month's transactions were actually repriced against the contract, as opposed to sampled or summarised? Anything short of full coverage is a known, quantifiable gap rather than a clean bill of health.
Median exception age, against your shortest dispute window. One number beside another number. If the first is approaching the second, everything else on this list is secondary.
Expired value, last quarter. The total of exceptions that aged out before they were submitted. If nobody tracks this, that itself is the finding, because it means the loss has never been visible.
Live contract versions. How many distinct rate regimes are currently in force across your providers. This is your true complexity number, and it is the one that predicts knowledge failures.
The honest limit of tooling
Scale problems are not solved by deciding to try harder, and they are not solved by tooling alone either.
Software can restore coverage, hold the batch window, and keep an exception queue against a contractual clock. It cannot read an ambiguous contract clause for you, and it cannot have the commercial conversation. The first contract model is judgment work, and it stays judgment work at any volume.
What changes at scale is that the judgment work becomes worth doing properly, because it now runs against every transaction instead of a slice of them.
The bottom line
Ask what breaks past a million transactions a month and you get an answer about infrastructure. Ask when your exceptions start expiring and you get an answer about money.
The second question is the one a CFO should be asking, because it has a date attached and the date is set by a contract you already signed.
Bluefyn reconstructs contract pricing and checks fees transaction by transaction, so expected against actual becomes provable rather than sampled. Bluefyn never moves, holds or custodies funds. It analyses transaction and provider data.
A charge can reconcile perfectly and still be wrong. At scale, the charges you never checked are the ones that stay that way.
Frequently asked questions
At what transaction volume does reconciliation break?
There is no single volume. The load is the product of transaction count, provider count, live contract versions and the length of your close. A smaller business billing across five providers can be under more strain than a much larger one on a single contract, so volume alone is a poor predictor.
What is the difference between reconciliation software and this problem?
Reconciliation software in the general market is built for the financial close: matching the ledger to the bank so the books balance. That answers whether two records agree. It does not answer whether the charge in both records was priced correctly under the contract, which is a separate question with a separate reference document.
Why is sampling a problem if it is statistically valid?
Because a provider refunds evidenced charges, one transaction at a time. A sample proves an error pattern exists but not what it cost, and no provider is obliged to extrapolate your sample on your behalf. Sampling preserves the insight and forfeits most of the claim.
What is the single most useful number to ask for?
Median exception age set beside your shortest contractual dispute window. That comparison tells you whether your backlog is a timing issue or a permanent loss, and most operations have never put the two numbers next to each other.
Does adding headcount fix an exception backlog?
It can, temporarily, and it does not change the shape of the problem. Exception volume grows with transactions while headcount grows in steps, so the backlog returns. The durable fix is raising coverage and enforcing a queue deadline set by the dispute window rather than by internal preference.
For the cost side of scaling a reconciliation function, see how to cut the cost of your reconciliation team.



