Every engagement in this series starts from a client's own priorities. This client's priority was uptime: a redesigned security model touches every role in the tenant, and this business doesn't get a maintenance window long enough to cut all of it over at once. So they made a completely reasonable call, deliver the redesign in batches, one role family at a time, verify each one against the live environment, move to the next. Nobody on their side did anything wrong. What follows is what that decision actually costs on our side, batch after batch, and it's worth writing down honestly because it isn't a story about anyone's mistake. It's a story about what a phased delivery model structurally requires that a single cutover doesn't.

What a single cutover buys you

When a redesign ships as one delivery, the comparison is simple: here is the target state, here is the current state, here is the diff. You verify it once, you ship it once, and the "current state" you verified against is still the current state by the time it goes live, because nothing else touched it in between. That's not a small convenience. It's the thing that makes a diff trustworthy in the first place, the target you compared against is still the target.

What changes the moment batch two exists

Batching breaks that assumption on purpose, and for a good reason: the business keeps running between batches. People still need role changes, support tickets still get raised, and a live production environment doesn't wait for the next delivery. So the "current state" that batch two gets diffed against isn't the state batch one modeled. It's batch one as actually delivered, plus whatever else has happened since, including changes the client's own team made directly in production to keep things working. One of those, a navigation privilege patched straight into a live role between our deliveries, had to be recognized and deliberately preserved rather than silently reverted the next time we touched that role. That's not a data quality problem. That's the client's team doing exactly what they should, keeping the business running, and it's exactly the kind of thing a single cutover never has to account for because there's no gap in time for it to happen in.

Timeline diagram: Batch 1 ships, correction v2 narrows revocation scope from 210 to 75 rows, correction v3 restores 11 wrongly-held-back additions, the next batch finds 3 leftover legacy roles, and the latest batch hits a NO-GO because the gap was only visible at that intersection.

The corrections cascade: v2, then v3

Batch one delivered three core roles, 156 duties, and 709 privileges. It also came with a revocation list, the old roles being retired in favor of the new ones, scoped down from 210 rows across five legacy roles to 75 rows across the three that had a verified replacement in this specific batch. That narrowing wasn't a mistake, it was the honest state of a picture still being assembled one batch at a time: the other 135 rows genuinely didn't have a confirmed replacement yet, because their replacement roles hadn't shipped in this batch. Holding them back was correct. It also meant the revocation list itself had to be revisited in a later pass, once more of the picture existed.

The correction after that was narrower but telling in a different way. A first delta held back eleven Deny-duty additions pending review, additions that only tighten access, never loosen it, so there was no compliance reason to hold them. They'd been treated cautiously because, in isolation, any addition to a role mid-batch looks worth double-checking. Once it was clear these were pure tightening, they went back in. Neither correction happened because the underlying redesign was wrong. Both happened because a batch-by-batch delivery means re-evaluating, batch by batch, exactly how much of the full picture you actually have in front of you yet.

The leftover-roles discovery

Starting work on a later batch turned up three legacy add-on roles still live in production that didn't match anything, not the original batch-one delivery, not the corrected version of it. They were uncleaned leftovers from a phase before this one. In a single cutover, there's no such thing as "a phase before this one" for something to be left over from, everything that exists gets accounted for in the one diff. In a phased delivery, every earlier batch is a place something can quietly survive past its intended lifespan, invisible until the next batch happens to look in that specific corner.

The NO-GO

The most recent batch was meant to deliver the rest of this client's roles, everything not already covered by the earlier batches. It hit a hard stop instead: a gap in how the picture was being assembled, one that had no way of showing up in any earlier batch, because it only existed at the specific intersection of "everything delivered so far" and "this new slice of scope." Nothing shipped from that batch. Two independent review passes were run before touching production again, specifically because this is exactly the kind of gap that a phased delivery makes possible and a single cutover structurally can't produce, since a single cutover only ever has one intersection to check, not a growing number of them.

Why the review pass grows, not just the delivery

It would be easy to read all of this as a testing problem, as if a more thorough QA pass on batch one would have prevented the corrections that followed. It wouldn't have, because the thing each correction was catching didn't exist yet at the time batch one shipped. The eleven Deny-duty additions weren't a defect sitting in batch one waiting to be found, they became a question worth re-litigating only once there was a second delta to compare against. The three leftover legacy roles weren't visible from inside batch one's own scope, they only became findable once a later batch's review happened to look at that specific corner of production. Each individual review pass was proportionate to what it was reviewing. What grows isn't the rigor of any single review, it's the amount of accumulated state every new review has to hold in its head before it can trust its own diff. The most recent batch needed two independent review passes before anything touched production again, not because the batch itself was unusually large, but because "everything delivered so far" is a bigger and more entangled thing to verify against with every batch that's already shipped.

✓ What this actually costs

Batching by role family isn't a mistake, and we wouldn't tell this client to stop. But it has a real, specific, structural cost: every batch reopens the question "what's actually true right now," and answering that question gets more expensive, not less, the more batches have already shipped. That's not overhead we'd design away with more careful testing. It's the literal shape of incremental delivery against a moving target.

What we'd tell the next client

If the reason for batching is uptime risk, not every way of slicing the work carries the same cost. Batching by role family, redesigning one department's roles at a time, means every batch touches a fresh slice of the whole security model and has to be re-diffed against everything delivered so far plus whatever's changed since. Batching by user population instead, delivering the full redesign to one group of users at a time, keeps the model itself frozen and only staggers who's using which version of it. Each user still gets the whole finished design, verified once, the moment their batch goes live, instead of getting an early version of a design that's still being corrected underneath them. It's a real conversation worth having before committing to the batching shape, not after the third correction pass, and it's one we now raise explicitly at the start of any engagement where uptime is the reason batching is even on the table.

Questions we get asked

Isn't this just normal iteration on a big project?

Some of it is. The difference is that ordinary iteration responds to new information about what the design should be. This is different: the design didn't change between v2 and v3, the ground it was being compared against did, because the business kept running in between deliveries. That's a property of the delivery model, not of how well the design was scoped.

Why not just tell the client to deliver it all at once?

Because their reason for batching is legitimate: this business can't absorb every role changing on the same day without a real operational risk. Our job isn't to argue them out of a real constraint, it's to be honest about what that constraint costs, and to help them choose a batching shape that costs less of it, if one's available.

Does this mean batched delivery is always the wrong call?

No. For a business that genuinely can't pause, it's often the only responsible option. The point isn't that batching is wrong, it's that it isn't free, and knowing exactly what it costs, and why, is what lets a client make that trade-off with open eyes instead of discovering it one correction pass at a time.

Could this have been caught earlier with better upfront planning?

Some of it, yes, the batching-shape conversation now happens before the first batch, not after the third correction. But not all of it. The leftover legacy roles and the NO-GO both depended on production having changed between deliveries, in ways no amount of upfront planning can predict, because the changes hadn't happened yet when the plan was made. Some of this cost is genuinely irreducible for as long as the delivery stays phased.