Back to Blog

Data Quality: Uncertainty Tags and How to Mark What Your Data Cannot Support

AIBig DataInformation Technology

Last updated: October 4, 2026

Data quality is whether a value can support the decision it feeds. An uncertainty tag marks, field by field, what a value cannot support yet, why, and who will settle it. The tag then travels with the number into every report built on top of it.

Weighted average net leverage, 4.6x.

That was the line on page three of a fund’s quarterly letter to its investors, set in the same type as everything around it. Nothing marked it. I asked where each borrower’s EBITDA had come from. It took the team two days to answer. That was the first finding.

Here is what sat under the number. I’ve rounded the shares and kept the shape.

Where the EBITDA came fromBorrowersShare of exposure
Compliance certificate, tied out to the borrower’s financials2461%
Compliance certificate, add-backs not yet reviewed722%
Last quarter’s figure, carried forward because the certificate was late411%
Sponsor’s model at closing, no certificate delivered yet36%

Thirty-eight borrowers. Sixty-one percent of the exposure rested on figures somebody had checked against a source. The rest was a mix of claims, stale numbers, and a sponsor’s forecast, and every one of them was averaged into a single figure with one decimal place.

Nobody lied. Nothing on that page was wrong in the sense an auditor would mean.

Recomputed with the unreviewed add-backs taken out, the same book came to 5.0x. A reader of the letter had no way to know that, and neither, until that week, did most of the team. The distance between what a number says and what it can support is the subject of this piece, along with the simple mark that closes it.

KORE1, whose site you’re on, recruits the people who end up owning this work, and its recruiters for data engineers and data scientists fill both of the seats I name near the end. The tagging itself starts in a spreadsheet, and I’d start it there.

A man in rimless glasses and a charcoal quarter-zip sweater reading a single sheet of paper at an oak desk, a pad of orange page flags beside him

What Data Quality Means When a Number Has to Hold Up

Data quality is the degree to which data can support a specific use. Most definitions measure it across a whole dataset, along dimensions such as accuracy, completeness, consistency, timeliness, validity, and uniqueness. For a credit fund the unit that matters is smaller. One value, for one borrower, in one period, feeding one decision. That changes the tooling.

I’d still run the dimensions. They answer a fair question for the data team, which is how clean a table looks this month. The person signing a covenant compliance memo is asking something else. Can this number, the one in front of me, carry the weight I’m about to put on it?

That’s the tag’s job. I define a data foundation as three things. One agreed record per borrower, per policy, per claim. Written rules deciding which fields deserve trust. And a tag on any field that hasn’t earned it. The first two get the budget. The third is the part a reader of the output actually sees. Usually it’s missing.

Data quality dimension scoreUncertainty tag
UnitA table or a columnOne value, in one record, for one period
Question it answersHow clean is this dataset?Can this number support this decision today?
Who reads itThe data teamWhoever signs the report
ExampleEBITDA column 97% completeOne borrower’s June quarter EBITDA, add-backs unreviewed, analyst named
When it changesAt the next profiling runWhen someone settles it, and the record says who

One warning about the word. Snowflake has a feature called tags, and its documentation defines one as “a schema-level object that can be assigned to another Snowflake object,” with uses like finding columns that hold sensitive data or tracking spend by cost center. Snowflake object tagging and Databricks Unity Catalog tags both label containers, down to the column. So the tagging feature in the tool you already bought marks the EBITDA column as financial data. It can’t say that one borrower’s June number is a guess lifted from the sponsor’s closing model, that the guess will be replaced when the certificate arrives, or that a named analyst is waiting on it. The words match. The jobs don’t. I’ve watched the confusion eat the first hour of a project kickoff.

What a Tag Has to Carry

A tag that says only “uncertain” gets ignored by the second month. People stop seeing it the way they stop seeing a check-engine light that’s been on since spring. So it needs more.

In a document-to-decision pipeline, tagging is step five, and each field comes out in one of three states, trusted, needing review, or unreadable. That’s the state. A tag needs five more parts. Below is a single tag, written out in full, for one borrower.

Part of the tagOne borrower’s June quarter EBITDA
What it marksBorrower 1047, Consolidated EBITDA, quarter ended June 30
StateNeeds review. Usable in reports only with a mark beside it
ReasonAdd-backs claimed, not yet reviewed against the agreement
EvidenceCompliance certificate, schedule 2, line 14. $3.2 million of add-backs against $11.9 million of reported EBITDA
OwnerThe portfolio analyst who covers the credit
What settles it, and by whenEach add-back matched to the definition of Consolidated EBITDA, or the cap applied. Before the September monitoring call

The reason comes from a fixed list. Free text feels more honest, and within a quarter it turns into forty ways of saying a certificate was late. Keep the list short enough that the person who would do the work can recite it. I keep mine on one page. For covenant figures the list usually starts with six.

  • Late certificate. Last quarter’s figure was carried forward.
  • Add-backs claimed and not reviewed against the agreement’s definition.
  • The number came from the sponsor’s closing model. Nobody has seen an actual yet.
  • Restated. The borrower revised a prior period and the restatement hasn’t been run through.
  • Two sources disagree and nobody has decided which one governs, which at most funds means the loan system and the administrator, often because they don’t agree on which legal entity a name refers to.
  • Read off a PDF by a model, not yet checked against the page.

Borrowing base certificates need their own list, and I wrote a separate set of reason codes for borrowing base fields. Different document, different failure.

Of the six parts, teams skip the last one. Every time. I’ve stopped being surprised. “What settles it” is what turns a tag from a complaint into a task with a finish line, and without it a tag sits on a record until the whole office has learned to look past it.

The evidence line earns its place more quietly. Somebody opening a tag should land on the page and the line, ready to read the record, rather than on a search through a shared drive. In that example, $3.2 million is 27% of what the borrower reported. That’s a lot. Some credit agreements cap certain add-backs, and whether this one did, and how, is exactly the question the analyst has to answer from the definitions section. Over a few seasoned deals, the same records also show how often the add-backs actually came true. A tag that sends people hunting is a tag nobody clears.

A hand pressing a burnt orange page flag onto the edge of a stack of white paper on a walnut table

The Tag Travels With the Number

Measurement science settled this question decades ago. The Guide to the Expression of Uncertainty in Measurement, published by the Joint Committee for Guides in Metrology, opens its introduction with this sentence. “When reporting the result of a measurement of a physical quantity, it is obligatory that some quantitative indication of the quality of the result be given so that those who use it can assess its reliability.”

Obligatory. For a thermometer.

A leverage figure is a measurement too, assembled from other people’s measurements, and in most funds it gets less care than a lab thermometer does. Three rules close most of the gap.

The first is inheritance. A number built from tagged inputs is tagged. A covenant test with one input under review is a test under review, however comfortably it passes, and comfort is worth stating rather than assuming. A test that clears its threshold by a wide margin on unreviewed add-backs may well clear it again once they’re reviewed. The tag lets you say so. In writing.

Second, carry the share along with the flag. “4.6x, with 39% of exposure resting on EBITDA still under review” is a sentence an investment committee can act on. “4.6x” with an asterisk invites a phone call.

Third, where you can bound it, bound it. Recompute with the tagged inputs at their least favorable plausible value, which for unreviewed add-backs means taking them out. For the book at the top of this piece that gave 5.0x, and printing both figures would have cost the letter one line. They didn’t print it.

Rollups are where tags die. I check them first. A weighted average across 38 borrowers dissolves seven borrowers under review, worth 22% of exposure, along with four stale figures and three sponsor forecasts, into one tidy decimal unless the tag is carried into the calculation itself. NAV does the same thing. So do portfolio yield and the watchlist count. At most funds the one check on this is still done by hand, by whoever rereads the letter the night before it goes out, and that person is usually checking the arithmetic rather than asking what sits under each figure.

The Standards Already Ask for It

None of this is new outside private credit. Two examples, from different corners of finance.

Actuaries in the United States work under ASOP No. 23, Data Quality, issued by the Actuarial Standards Board. The current edition applies to work where the data were provided or developed on or after April 30, 2017. It doesn’t require an audit. It requires disclosure. Section 4.1 lists, among other things, “any limitations on the use of the actuarial work product due to uncertainty about the quality of the data” and, in summary form, “unresolved concerns the actuary may have about questionable data values.” Read that second phrase slowly and it’s an uncertainty tag written as a paragraph. I picked the habit up there and now use it on loan books. Nothing requires it there.

The second sits closer to a credit fund. The SEC adopted Rule 2a-5 on fair value determinations in December 2020, effective March 8, 2021, and it covers registered funds and business development companies. Among the valuation risks a board or its valuation designee is expected to assess and manage, the adopting release lists “the extent to which each fair value methodology uses unobservable inputs.” Under ASC Topic 820, unobservable inputs are Level 3. Much of private credit sits there. The same release notes that BDCs, which must invest at least 70% of their assets in certain smaller private or public U.S. companies, usually hold securities valued with Level 2 or Level 3 inputs.

So your valuation team already tags uncertainty. It calls the tag Level 3 and discloses it every quarter. Nobody calls that bureaucracy.

The mark stops at the asset, though. Underneath a Level 3 loan sit the EBITDA, the add-backs, the late certificate, and the sponsor’s projection, and those usually live in a workbook with no marks at all. Field-level tags are the same discipline, applied one layer down, where the numbers are actually typed in.

How Tagging Goes Wrong

I’ll tell this one in order, because at one fund the failures arrived in order.

The first pass marked 14% of the covenant fields on the book. The fund’s CFO looked at the count and asked me, fairly, whether we’d built a smoke detector that goes off when somebody makes toast. I didn’t think so. Those fields had been in that condition for years. Nobody had counted them before.

Week three, we found the blank. A borrower’s EBITDA cell was empty, a formula upstream treated empty as zero, the ratio threw an error, and an IFERROR wrapper somebody had added years earlier quietly removed the borrower from the weighted average. It was one of the weaker credits on the book. Its absence had been flattering the headline figure for at least four quarters, and nobody had touched that workbook with bad intent. I still think about that wrapper.

By the end of the first month the tags existed, but in the wrong place, a ticket queue the operations team liked and the investment committee had never opened. We moved them into the workbook as columns beside each value. Ugly. It worked.

Month two brought the confidence-score argument. The extraction tool scored every field it read, and someone proposed treating anything above 0.9 as settled. A score like that is the model grading its own reading of the page. It can’t tell you whether an add-back is permitted under the agreement, which is most of what we were tagging. I’ve gone into that score at more length in why OCR projects stall.

The worst moment came at the first quarter-end. Eleven tags were cleared on the last afternoon, before the letter went to print, by an analyst who was trying to help. None of the eleven had met its settle condition. We reopened them the next morning and added a rule that whoever clears a tag is named in the record beside the date. Two quarters after that first pass, open tags were under 3% of covenant fields, and certificates were arriving earlier, because borrowers’ deal teams had learned that a particular person at the fund was waiting for each one.

Who Owns the Tags

Two seats, in most firms I work with. At a small fund it’s often one person wearing both.

Somebody has to keep the reason list short, write the settle conditions, and sweep the open tags every week. At a larger manager that’s a data governance analyst. At a 30-person fund it tends to be the controller, or the senior portfolio analyst who already gets asked everything. Titles vary. I care about one qualification above the others, and it’s plain enough. Hand them a credit agreement. Can they find the definition of Consolidated EBITDA and tell you which add-backs it permits, and with what cap?

The plumbing is an engineer’s problem. Every join and every rollup between the raw file and the investor letter is a chance for a tag to fall off, so the tag needs its own columns on the fact tables and each aggregate needs to report what share of its inputs is still under review. In a layered build that work lands in the middle and top layers, in the order I laid out in the build order for a credit fund’s data layers. The person who designs those tables is a data warehouse engineer by trade, whatever their offer letter calls them.

Both roles work fine as a contract hire for the first quarters, while the opening backlog of tags is worked down. Ninety-two percent of the people KORE1 places are with the same client a year on. For this work I’d weigh that more than usual. Continuity is the point. A reason list kept in one analyst’s head just moves the old problem somewhere new, the problem behind one person is the database.

A woman in a burnt orange blouse explaining a point to a bearded colleague in a light blue shirt beside a tall office window

Questions I Get Before Anyone Tags Anything

Isn’t an Uncertainty Tag Just a Confidence Score With a New Name?

A confidence score is a model’s opinion of its own reading, while an uncertainty tag records why a person cannot rely on a value yet and what would settle it.

They can coexist. A low score is one reason a field might get tagged. Most tags on a loan book have nothing to do with models, though, because a late certificate isn’t a machine-learning problem.

How Many Open Tags Should a Healthy Book Carry?

Low single digits, as a share of covenant fields, once the first two quarters of tags have been worked through.

Zero is a warning sign. Every book has late certificates and unreviewed add-backs somewhere, so a book with no open tags usually means nobody is looking. I’d ask why.

Do Investors Ever See the Tags?

They see the effect, usually as one line beside a headline figure that says how much of it rests on inputs still under review.

Nobody sends an investor the tag table itself. Operational due diligence teams do ask how reported figures get produced and checked, though, and a fund that can show its tag history answers with a record rather than a description of a process. Investors notice.

Which Data Quality Dimensions Should a Credit Fund Measure First?

Timeliness and consistency come first, because late certificates and systems that disagree cause most of the tags I see on a loan book.

Completeness is a close third. Accuracy matters most of all, of course, but it’s also the hardest dimension to score in bulk, and it’s the one tags exist to handle one value at a time.

Can We Start Without Hiring Anyone?

You can start this quarter with four extra columns in the workbook you already use and a named owner for each borrower.

State, reason, owner, settle-by date. I say scorecard now, ML when your data can support it, and tags follow the same logic at the scale of a single number. Hire when the open tags outgrow the person clearing them, and you’ll know when that happens, because the settle-by dates start slipping.

Does Tagging Slow Down Month-End?

About a day, the first time.

After that it tends to give the time back. Fewer questions arrive after the numbers go out, since the questions were answered on the page, and the late certificates that cause most tags start coming in earlier once borrowers notice someone tracks them.

Tag One Number This Quarter

Pick the headline figure in your last investor letter. Leverage, NAV, yield, whichever one people quote back to you.

List every input under it. It takes an afternoon. Beside each, write where it came from and whether anyone checked it against the source document. Add up the share of the figure that rests on inputs nobody checked, then print that share next to the figure in the next letter. Even if it’s small. Especially then. If you’d rather start with the whole picture, a four-week data foundation diagnostic counts the hand work first.

If you’d like a second read on your reason list, send it over after you connect with me on LinkedIn. And if nobody on staff can own the tags, talk to KORE1’s recruiters about who could.