A model that is missing evidence can look cheaper than the truth. On one client, restoring the missing read-tier usage data raised the honest annual cost by at least 378,998 dollars and flipped the saving from positive to negative. The higher number was the correct one.

Missing data made the design look better than it was

On one client, the as-shipped cost rose from 1,833,942 dollars to 2,212,940 dollars per year once the missing read-tier usage data was populated, and the saving went from positive 15.9 percent to negative 1.5 percent. Nothing in the design got worse. The earlier, cheaper number was simply built on incomplete evidence.

This is the counterintuitive part that matters for anyone who trusts a low number: less data produced a lower cost and a better-looking saving. More data produced a higher cost and a worse-looking saving. The higher one was honest.

Why missing evidence biases the cost downward

The mechanism is specific. The optimization engine saves money by proposing read-only substitutions, swapping a cheaper read-tier grant in where a user only reads data. To propose that substitution, the engine needs read-tier usage evidence showing the user reads in that module.

When the read-tier data is missing, the engine cannot see those read-only candidates at all. They never enter the set of things the user is judged to need, so they silently drop out of the cost. The model looks lean because it is blind to a whole category of privilege the user genuinely holds.

Once the read-tier data is present, those privileges are correctly kept in the needed set. And because a read-only substitution still ships the write privilege until a dedicated read-only variant exists, each one is priced at the write tier. The cost rises because the model can finally see privileges it was previously missing, and it prices them the way they will actually ship.

So the increase was not a regression. It was the removal of a downward bias. The missing-data version was understated, and restoring the evidence corrected it upward to the number the client will actually be billed.

A caution on how the figure was first obtained

The 378,998 dollar figure is sound, but the way it was first produced was not, and that is worth stating because it is the exact trap this kind of analysis invites.

The figure was originally calculated by subtracting a table written a week earlier from one written minutes earlier, while the client's new build had not finished running. That is a stale-read mistake: comparing two different vintages of data as though they were the same run. The first version of the number came from that error.

It only became trustworthy after re-derivation. The cost was confirmed against a settled build read of 2,212,940 dollars, and the mechanism was verified in the engine code itself. The number stood up. The original derivation did not. The lesson is to re-derive from a settled, single-vintage read rather than re-quote a figure, however right it later turns out to be.

It is a lower bound, not a tidy attribution

One more honesty note. Three changes landed between the two builds being compared: the read-tier restoration, a fix to the cost cover logic, and a rule excluding license-free users. Attributing the full delta to any single one of them would be wrong.

For this client the license-free rule was exactly zero, because its license-free users carry no priced licenses. The cover fix moves cost down, not up. So the read-tier effect is at least 378,998 dollars, because the one change working in the other direction can only have reduced the total. The direction of the conclusion survives the fact that three variables moved at once. The precise attribution does not, so it is reported as a lower bound.

Bottom line

A low cost built on missing evidence is not a win, it is a measurement error waiting to be found. Restoring the read-tier data took this client from an apparent 15.9 percent saving to a 1.5 percent cost increase, a swing of at least 378,998 dollars, and the worse-looking number is the one the client would actually pay. More evidence produced a higher, truer figure. Quote the number that survives a settled read, and derive it fresh rather than repeating one that happened to be close.

Frequently asked questions

Why would more data ever make a saving worse?

Because the engine can only propose a read-only, cheaper substitution for a privilege it can see the user reading. Missing read-tier data hides those privileges, so they drop out of the cost and the model looks cheaper than reality. Restoring the data puts them back and the cost rises to the truth.

If the privileges are read-only, why are they priced at the write tier?

Because a read-only substitution still ships the underlying write privilege until a dedicated read-only variant is built. The user still holds write access, so the license is billed at the write tier until that variant exists.

Can we recover the saving that disappeared?

Yes, by building the read-only privilege variants so the read-only substitutions actually ship as read-only grants. At that point those privileges price at the read tier and the saving returns as a real, deliverable number rather than a modeling artifact.