Expert Opinion

How Close Are We Really to Fully Autonomous Data Cleansing?

As enterprise datasets grow to millions of records, how much data cleansing can AI safely handle on its own, and where should automation stop before a plausible-looking error becomes a costly one?

Fully Autonomous Data Cleansing

Follow Us

LinkedIn Facebook Twitter YouTube Instagram

Data cleansing looks almost perfectly suited to automation. Much of the work is repetitive, AI can compare records at scale, and the models are increasingly getting good at spotting anomalies, duplicates, and inconsistencies. Yet fully autonomous data cleansing at enterprise scale remains difficult.

The reason is not that AI cannot clean data without human approval. It already can, particularly when the correct answer can be established through clear rules or reliable reference data.

The harder part is defining what happens when an intelligent data cleansing tool detects a wrong value but cannot establish, with enough evidence, what the correct value should be.

That is where most of the difficult data cleansing work begins, and where AI starts stumbling.

The Ground Reality of Automated Enterprise Data Cleansing

Automated Enterprise Data Cleansing

Across 15+ data cleansing projects we handled over the last six months, from roughly 100,000 to 200,000 records per database, around 35–40% of the data we reviewed was wrong. I don’t mean formatting errors, extra spaces, or obvious duplicate records. Just plain wrong - wrong person, wrong company, wrong email, stale job title.

Much of this data came from automated scraping, and nearly half the datasets had already passed through AI data cleansing tools. Many records didn't look bad, which is likely why the AI tool did not flag them. But when my team manually validated the data, we saw particular patterns everywhere: “neat-looking” data fields that were actually errors with severe consequences.

1. Email addresses constructed from predictable corporate patterns rather than discovered from reliable records

For example, a scraper may infer firstinitial.lastname@company.com from the naming conventions it sees elsewhere and add it to the record as though it were a confirmed address.
The problem is that a valid-looking format does not establish that the mailbox exists. The address may never have been created, may belong to someone else, or may follow a different naming convention than the scraper assumed. Because the result looks structurally correct, it can pass automated cleansing even when the underlying contact information is wrong.

2. Emails validated only by checking whether a domain or mail server will accept a verification request, but not whether a specific person can actually receive the email

Catch-all domains, for example, may accept verification requests for addresses that do not exist, producing a false positive. Corporate security gateways can behave similarly by accepting or masking verification requests rather than revealing whether a specific mailbox exists. Greylisting can create the opposite problem: defense systems temporarily reject initial pings from unrecognized servers and tell them to retry in 15 minutes or so. Verification tools can misinterpret this pause as a dead mailbox and discard the email, even if it belongs to active, high-value prospects.
Such automated validation can introduce errors in both directions. The tool may pass fabricated email addresses while discarding valid ones, affecting teams that depend on that data for outreach.

3. Job titles that may be internally consistent but factually out of date/incorrect

Job titles can appear reliable simply because the same information has been repeated across multiple sources. But consistency across sources does not necessarily mean the information is current. The person may have been promoted, moved into another function, or left the company altogether. If the underlying sources are outdated, an automated cleansing system can reinforce the same stale information instead of flagging it as an error.

4. Entire contact profiles built around an incorrect assumption

Once an automated tool associates a person with the wrong company or individual, subsequent data enrichment or data cleansing tools can continue building the record around that incorrect assumption. The resulting profile may look complete and contain a plausible company name, realistic job title, matching location, valid phone number, and other details that appear to confirm one another. Each field may look reasonable, while the underlying identity is still wrong.
Because automated cleansing systems typically check consistency between fields rather than independently confirming the person's identity, a complete profile can pass multiple validation checks without raising an obvious flag.

Companies pouring billions into AI data management solutions think they're saving a lot, but within months, they could end up paying 7-10x or more of that in hidden remediation costs.

Read more: AI Data Quality Management Problems

Detecting an Error and Knowing the Correct Answer Are Different Problems

AI can increasingly detect suspicious records. That does not mean every suspicious record has an obvious replacement value.

Suppose three sources say a person is Director of Sales, while one newer source says Vice President of Revenue. Is the old title stale? Was the person promoted? Are two people with similar names being combined? A sales rep would investigate further by checking current company pages, professional profiles, internal CRM history, external databases, and other sources. An AI without that business context (what to look for, across which channels, and how to decide whether the evidence is strong enough to justify changing the record) cannot do that.

Simply put, the key requirement for autonomous data cleansing is sufficient context, not just greater model intelligence.

Full Autonomy Is Already Possible where the Definition of a “Correct Field” Is Clear

Some automated data cleansing tasks already run without meaningful human intervention. For example, data cleansing tools can manage date formatting, trim whitespace, standardize country codes, validate values against authoritative lists, and correct deterministic syntax issues on their own. In fact, more contextual tasks can also become autonomous when several independent signals point to the same answer.

The boundary therefore is not:

Rules = automation, judgment = humans.

It is closer to:

Strong evidence + low consequence = greater autonomy.

Conflicting evidence + high consequence = greater oversight.

Consider these two changes with similarly high confidence.

  • Correcting stainles steel to stainless steel carries little downstream risk.
  • Merging two customer accounts can alter revenue attribution, sales ownership, contract history, reporting, or customer communication.

Therefore, an enterprise should not apply the same automation threshold to both.

"Human in the Loop" Data Cleansing will Always be Necessary when the AI Reaches an Evidence Gap

Consider an AI data cleansing system evaluating two supplier records. It finds:

  • 96% name similarity
  • identical address
  • identical website domain
  • different tax IDs
  • different payment accounts

Should they be merged?

The algorithm may be perfectly capable of identifying the contradiction. What it cannot safely do is pretend the contradiction does not matter. A reviewer may discover that the organizations are separate subsidiaries operating from the same location and should be kept as similar records. Automated data-cleansing software may not have the evidence to reach that conclusion.

Also, Some Cleansing Decisions Are Really Governance Decisions

There is another reason complete autonomy remains difficult for data cleansing tools. Some apparent data quality problems do not have a universally correct answer.

Consider a global customer represented by several regional subsidiaries. Should the CRM contain:

  • one global account,
  • separate legal entities,
  • country-level accounts,
  • purchasing entities,
  • billing entities,
  • or some combination?

There is no mathematical answer. The correct data structure depends on how the organization sells, reports revenue, manages contracts, assigns territories, and evaluates customer relationships.

The same problem occurs across enterprise data.

  • Should two product SKUs be consolidated because the products are materially identical?
  • Should an inactive customer be marked churned?
  • Should a physician record be associated with a hospital where they currently practice or the network employing them?
  • Should an old supplier address be replaced or preserved for historical transaction records?

These are not simply data defects. They are interpretations of business reality. AI may eventually execute the resulting policy autonomously. But somebody first has to decide what that policy should be. This is where the human role shifts from record-level correction to rule-level governance.

There Is Also an Accountability Question Here

When an automated system changes a field, the issue is not only whether the change was statistically likely to be correct. An enterprise also needs to know who is responsible for that decision and what happens if the correction turns out to be wrong. Otherwise, you’d be implementing automation without defining accountability and increasing your liability.

That matters more as the consequences of an error increase. A typo corrected automatically is unlikely to require anyone to investigate who approved the change. But if an automated merge changes a customer's legal identity, alters revenue attribution, or removes information needed for a contract, the organization needs a clear record of why the change was made, what evidence supported it, and whether a human or a defined business rule authorized that type of decision.

This is why autonomy cannot be measured by model confidence alone. The more consequential the decision, the more important it is to have not only sufficient evidence, but also a clear chain of responsibility.

A Mature Fully Autonomous Cleansing Model that Makes Sense for Enterprise Datasets

A mature autonomous data cleansing workflow would not treat every record equally. It would separate decisions by evidence and risk, aiming for maximum safe autonomy, and escalate records to a human where evidence conflicts or ambiguity exists. And they’d keep escalating until a definitive rule can be created from the pattern of corrections humans apply to the problematic records.

Deterministic Corrections

Tier 1: Deterministic Corrections

Examples include formatting, character normalization, reference validation, standard code conversion, and exact duplicate removal.

These can usually run automatically.

Contextual Corrections

Tier 2: Contextual Corrections

Examples might include updating an outdated company address confirmed by trusted registries or correcting an organization name supported by strong entity resolution.

These can increasingly run automatically with logging and monitoring.

Ambiguous Decisions

Tier 3: Ambiguous Decisions

For a certain record/field, several plausible values could be true.

The system should collect evidence, recommend an action, and escalate the decision.

Policy or High-Impact Decisions

Tier 4: Policy or High-Impact Decisions

For changes that affect legal entities, financial reporting, customer identity, regulatory obligations, healthcare records, or important business definitions.

Here, organizations may deliberately retain human authorization even when AI confidence is high.

So, How Close Are We?

We are already quite far along the path to autonomous data cleansing. The question is whether we are close to fully autonomous data cleansing. At enterprise scale, we are not there yet! At this point, achieving autonomous data cleansing depends on how precisely one marks the boundary between human and machine decision-making.

In fact, the difficult cases in enterprise data cleansing (where the data looks plausible, multiple sources disagree, the correct value depends on business context, or changing the record has significant consequences) are exactly the cases we continue to encounter in enterprise datasets, including datasets that have already passed through automated cleansing tools. We increasingly deal with exceptions, define business rules, resolve ambiguous cases, and determine what constitutes a correct record. That ground reality tells me that achieving autonomous data cleansing is now about building a system that knows which decisions it can make, which decisions require more evidence, and which decisions belong to the business rather than the machine.

And as that boundary becomes clearer, the amount of data cleansing that can safely happen without human intervention will continue to grow. But for the foreseeable future, the highest-value data cleansing decisions will still depend on something AI cannot manufacture from a confidence score alone: sufficient evidence and the business context to interpret it.