Two of our earlier pieces covered how a single privilege gets its verdict: stays as-is, downgrades to read, or drops, and why "no evidence" has to be handled with real care. Neither one produces a role. A disposition table is, deliberately, just the decision surface, thousands of rows, one per user per privilege, each with a verdict. Someone still has to turn that into an actual set of roles a security team can assign, audit, and maintain.
That's a genuinely different problem, and it doesn't have an obvious right answer.
The needed set, and the two shapes it can take
Start from what each user actually needs: everything that stays_as_is at its native tier, plus everything downgrade_to_read at its cheaper read tier. Privileges the evidence couldn't confirm are held aside unless keeping them would raise the user's license tier; anything genuinely unused and observable is already gone. That's the raw material, a binary user-by-privilege matrix.
There are two honest ways to turn that matrix into roles, and they sit at opposite ends of a real trade-off.
The first: give every user their own role, built from exactly what they need. Zero over-provisioning, by construction, every single time. It also produces roughly as many roles as there are users, which is not a role model, it's a spreadsheet with extra steps. No security team can administer that.
The second: keep the handful of broad roles the organization already has, and just trim what's inside them. Fast, familiar, and it only works if the roles were reasonably coherent to begin with. When a role's members turn out to do genuinely different jobs, wide as it is, no amount of trimming produces something both cheap and correct at the same time.
Neither extreme is usually the right answer. What sits between them is a genuine optimization problem: minimize the number of roles, subject to every user keeping everything the evidence says they need. That has a name in the literature, the Role Mining Problem, a Boolean matrix decomposition that's formally equivalent to set cover, and known to be NP-complete. Naming it correctly matters in practice, not just academically: ordinary similarity-based clustering has no concept of "must cover", it will happily grant a privilege three people need to everyone in a group, or silently drop one that two people depend on. That's a fine trade-off for a marketing segment. It's not an acceptable one for a person's ability to do their job.
Core roles: the part that's never allowed to inflate anyone
The first half of the answer is a core role per job family: whatever privileges the clear majority of that group's active users genuinely need, at a shared license tier. The threshold for "clear majority" isn't fixed; it's swept across a range and chosen deliberately, because pushing it up or down trades role count against how many people carry something they don't personally use.
One rule doesn't move, though: a privilege never enters a core if it would raise even one member's license tier above what their own evidence supports. Cores are built to never inflate anyone, by construction, not by later cleanup. Anything that would blow that rule for even a single person gets pushed out to the second half of the model instead.
Add-ons: the residual, mined rather than guessed
Whatever's left over, the genuinely individual, occasional, or role-crossing access that didn't fit a shared core, becomes the add-on layer. This is where set-cover mining actually happens: group the residual privileges into license-homogeneous bundles that recur across many users, and select the smallest, cheapest set of bundles that still covers everyone's needs.
Naming an add-on well matters more than it sounds. When a mined bundle happens to line up with a real, named business function, an accounts-payable duty, a reporting duty, that's what the add-on is called and scoped to. When granting that whole named function would hand a few of its members more than they need, the add-on shrinks to just the common subset actually shared. Only when neither approach fits does the system fall back to naming the add-on after the pattern of usage itself. Reusing real business language whenever it's honestly available is what keeps the result something an administrator can reason about, instead of a wall of algorithmically-labeled clusters nobody recognizes.
What actually gets handed to the client
None of this is useful if the output is a table of privilege IDs. The delivered role catalog is deliberately written for a non-technical reviewer: a plain-English description of what the role is actually for, the single license a holder is required to carry (the cheapest license that covers everything the role grants, not the most expensive one theoretically possible), how many additional licenses a holder would need only if they exercised every privilege in the role at once, and, separately, which license families the role touches at all, a diagnostic figure, not a requirement. Keeping those last two numbers distinct matters: conflating "touches" with "requires" is exactly the kind of overstatement that makes a savings number impossible to defend later.
The core-threshold and the mining floor aren't set once and forgotten, they're swept across a range, and the resulting cost-vs-role-count curve has a visible "knee": a point past which accepting a few more roles stops buying meaningful savings, and a point before which chasing a few more dollars costs a disproportionate number of extra roles. That knee, not a single "optimal" run, is the number we bring to a client, because it's the one point on the curve that's honestly defensible as a starting position for a conversation about trade-offs, not a claim that a single design is the only correct one. We go deep on exactly which of these dials actually matters, tested across two very differently shaped clients, in our parameter sensitivity study.
Standard roles are left alone on purpose
Every object this process mints carries a version suffix, distinct from what already exists, and only modified-standard or fully custom roles are ever candidates for restructuring. An unmodified, standard Microsoft role is left exactly as it is. This isn't a shortcut: touching Microsoft's own baseline roles adds real risk (breaking assumptions other configuration or ISV code makes about a standard role's shape) for savings that are rarely there in the first place, since standard roles are usually reasonably scoped by the people who wrote them. The savings live in what's been customized, extended, or copied and modified since, and that's exactly where this process spends its effort.
A disposition table answers "should this specific access stay." A role model answers "what's the smallest, most defensible set of roles that respects every one of those answers at once." They're different problems, solved in sequence, and skipping straight from the first to a redesign without the second is how you end up with either a role per user or a handful of roles nobody can actually reason about.
Questions we get asked
Isn't a core role that grants slightly less than the old role a risk in itself?
Less risk than the alternative. A core is built strictly from evidence, plus anything genuinely unobservable is protected by the same guardrail covered in our disposition articles, never silently dropped. What's removed is access nobody in the group has shown a need for over the full evidence window, not a guess.
How many roles should we expect to end up with?
There's no universal number, it's set by how many genuinely distinct job functions exist in your own population, not a target we pick in advance. What we can tell you in advance is the shape of the trade-off: fewer roles costs more license spend, and we'll show you that curve before you commit to a point on it.
What happens to roles we've already customized ourselves?
They're exactly what this process targets. Unmodified standard Microsoft roles are left alone; it's the modified-standard and fully custom roles, the ones most likely to have drifted from what people actually do, that get re-mined.
Glossary
| Term | Meaning |
|---|---|
| Needed set | Everything a user's own evidence supports keeping, at native or downgraded tier, the raw input to role mining |
| Core role | A shared role for a job function, built to never grant any single member more than their own evidence supports |
| Add-on | A smaller, mined role covering access that doesn't belong in any one core, named after a real business function whenever possible |
| The knee | The point on the cost-vs-role-count curve where adding roles stops buying meaningful savings, the operating point we recommend |
| V2 object | Any role, duty, or privilege minted by this process, version-suffixed and kept distinct from the client's original security objects |