Data Governance: Data Quality (/ Integrity) | The Danger of Data ROT


The ask? Perform an audit of the company’s data estate. Like many organizations these days, everyone wants to quickly automate their legacy processes, build robust business intelligence (BI) reports, and fully embrace artificial intelligence (AI). However, per the audit findings, before any of those initiatives can properly get off of the ground, there is a lot of identified data ROT that should be addressed first.


Although the phrase is cliché, it still rings true: “Garbage in, garbage out.” Regardless of which initiative is planned first, the output will only be as good as the input. Data ROT will only complicate the automation builds, BI data sourcing, and AI prompt responses. For the uninitiated, ROT stands for Redundant, Obsolete, Trivial:


In short, redundant data is duplicate data stored in multiple locations or systems. Copies of the same exact data often needlessly lives in various places, not as backups, but as active duplicated working content, so there is no longer a true “Golden Record.” As a result, this raises several questions:

  • Which copy does AI use as grounding data?
  • Which copy is the latest and greatest for the BI reports?
  • Which copy is the automation meant to run against?
  • And ultimately, who is responsible for making these decisions?

Obsolete data, as the name suggests, is data that has outlived its usefulness. This typically happens when new policies, procedures, codes, and/or regulations are introduced, and already existing records aren’t purged. Storage eventually becomes a problem, but for the more immediate concerns, this is a problem for AI solutions. When grounding their data, people often ground to folders, not individual files. Now, the AI is grounded on a folder with obsolete data and depending on the prompt, could provide inaccurate, conflicting responses. Here at least, if the data cannot be deleted, then it should at least be archived to prevent AI hallucinations.


Lastly, trivial data is data that if it were deleted right now, there would be no measurable impact to the business. This data offers no business value, so the BI reports would still be exact, the automations would still run reliably, and the AI would still be accurate if this data went away. Still, this data is kept around. As with the obsolete data, if this data cannot be deleted, given there would be no impact, it should at least be archived, not cluttering the file directories. This clutter is why many orgs overspend on cloud storage costs.


Finally, with the data classified, action can be taken. Again, ROT should ideally be deleted, purged from all systems, and policies put in place to prevent future ROT. This helps makes AI perform better, BI be more accurate, and automations be more performant. For instances where the data can’t be deleted, archiving is an option. Ultimately, this helps with the initiatives, but also helps manage storage volumes and storage costs, so win-win-win.


Conclusion:
No one is expected to solve their data ROT problem overnight. Captain Grace Hopper talked about this back in the 1980s, so if it takes you a few months or years, that’s perfectly fine. Still, this is something that pays dividends once you get the ball rolling.

“When I liberate myself, I liberate others. If you don’t speak out ain’t nobody going to speak out for you.”

Fannie Lou Hamer

#BlackLivesMatter

Leave a comment