How to Clean and Maintain CRM Data

Clean in a fixed order, because doing it out of order creates work. Deduplicate first, since every later step done on duplicates has to be redone. Then fix formatting on the fields you actually filter by. Then fill the gaps that matter and accept the ones that do not. Then archive genuinely dead records rather than deleting them. Only then set up the weekly routine, which is what makes it stay clean. In HubSpot, contacts auto-deduplicate on email and companies on domain, so most duplicates arrive through imports or manual entry under a different address. Budget four to eight hours for the first pass and fifteen minutes a week after that.

Every CRM decays. People change jobs, companies get acquired, someone imports a list with slightly different formatting, and a salesperson creates a new contact because search did not find the existing one.

The decay rate is commonly put at roughly 2% a month, which compounds to a quarter of your database being wrong within a year. Nobody notices until a campaign goes out addressed to the wrong person at a company that no longer exists.

The sequence

Clean in this order

The order is not arbitrary. Each step invalidates work done before it if you get it wrong:

  1. Deduplicate. Any formatting or enrichment done on a record you later merge away is wasted.
  2. Fix formatting on the fields you filter and segment by.
  3. Fill gaps that change a decision, and deliberately accept the rest.
  4. Archive the dead. Doing this earlier means enriching records you were about to remove.
  5. Set the routine. Without it, you repeat all four next year.

Most CRM cleanups fail not because a step was done badly but because they stopped at step four.

Before starting, decide what the CRM is actually for. That sounds like a detour and it determines everything downstream. A CRM used to remember who to call next needs accurate contact details and activity dates and almost nothing else. One used to report on channel performance needs lead source on every record. One used to forecast needs honest deal stages and close dates above all. Cleaning fields that serve none of your actual purposes is the most common way a cleanup consumes a week and changes nothing.

Step one

Deduplicate

HubSpot deduplicates automatically on some fields — contacts on email address, companies on domain. That handles the obvious cases and misses the common ones: the same person entered twice with a work and a personal address, or a company entered as both "Acme Inc" and "Acme Incorporated" on different domains.

The tool for the rest is the data quality command center, under Data Management → Data Quality, which surfaces likely duplicates for contacts and companies and lets you review them in a table before acting.

Merging cannot be undone

Once two records are merged in HubSpot, they cannot be unmerged. The surviving record inherits activities, associations and most property values from both. That is the right behavior, and it means reviewing before merging rather than bulk-accepting suggestions — a wrong merge is permanent and takes two customer histories with it.

Which record survives

Pick the survivor deliberately, using this priority:

  • The one with the most activity history — emails, calls, meetings. History is the hardest thing to reconstruct.
  • The one associated with open deals.
  • The one with the more current contact details.

Where these conflict, activity history usually wins. A stale email address on a record with three years of context is a five-second fix; the context is not recoverable.

Step two

Fix formatting on the fields that matter

Not every field. Only the ones you filter, segment or report on, because those are the only ones where inconsistency changes an outcome.

  • Job title — "VP Sales", "V.P. Sales" and "Vice President, Sales" are three values to any filter. If you target by seniority, this matters; if you do not, leave it.
  • Company name — legal suffixes, ampersands and abbreviations.
  • State and country — two-letter codes or full names, chosen once.
  • Phone — consistent enough to dial and to match on.
  • Industry and lifecycle stage — these should be dropdowns rather than free text, and if they are free text, that is the actual fix.
Fix the input, not just the data

If a field is free text and it matters, convert it to a dropdown before cleaning it. Otherwise you are cleaning a field that will be re-dirtied by the next person who types into it, which is how a business ends up doing this annually forever.

Step three

Fill the gaps that matter

The instinct is to complete every field. Resist it. A missing value only matters if it stops you doing something.

Ask of each empty field: what decision does this change?

  • Missing email on an active contact — blocks everything. Fill it.
  • Missing industry when you segment campaigns by industry — fill it.
  • Missing company size on a record you will never target by size — leave it.
  • Missing lead source — fill it if you make budget decisions on channel performance, which is the main reason to have the field at all.

A CRM with four accurate fields per record beats one with twelve fields of which five are guesses, because the guesses are indistinguishable from the facts once entered.

Step four

Archive the dead

Records that are genuinely finished:

  • Contacts whose email has hard-bounced repeatedly.
  • People who have left, where you know the company but not the successor.
  • Deals lost more than a year ago with no subsequent activity.
  • Anyone who asked not to be contacted — these must be suppressed rather than deleted, so the suppression itself is recorded.
Archive, do not delete

Deleting destroys history you may need — why a deal was lost, what was discussed, whether someone opted out. Archiving removes them from working views while keeping the record. Storage is cheap; a deleted opt-out request is a compliance problem.

Step five

The weekly routine that makes it stick

This is the step that separates a CRM that stays usable from one that gets cleaned every eighteen months in a burst of frustration. Fifteen minutes, weekly:

  • Review new duplicates. A handful weekly is trivial; a year's worth is a project.
  • Log the week's activity — calls, emails and meetings against the right records.
  • Update stage on anything that moved.
  • Flag deals gone quiet — no activity in 30 days on an open deal is either a follow-up you forgot or a deal that is dead and inflating your pipeline.

HubSpot can send a weekly data quality digest summarizing changes in issue volume, which turns the review into reading a notification rather than remembering to go and look.

Measurement

What good actually looks like

"Clean" is not a state you reach, so measuring it as one leads nowhere. Four numbers describe CRM health well enough to manage, and all four take a minute to pull:

Duplicate rate

Duplicates as a share of total records. Under about 2% is normal decay; above 5% means records are being created without searching, which is a process problem no amount of merging fixes.

Activity coverage

The share of open deals with activity logged in the last 30 days. This is the single most revealing number in a small business CRM. If it is under half, the CRM is not describing your pipeline — it is describing the part of your pipeline someone remembered to type in.

Required-field completeness

Completeness on the four or five fields you actually filter by, not on every field. A record with those five right is more useful than one with twenty fields where six are guesses.

Pipeline age

The share of open deals with no movement in 60 days. This is the number owners least want to look at, because it usually reveals that a comfortable-looking pipeline is substantially historical. A pipeline that is 40% stale is not a forecast, it is a filing cabinet.

Trend beats absolute

None of these has a correct value, and chasing one is a waste of time. What matters is direction. A duplicate rate rising month over month tells you something specific is broken in how records get created; a flat 3% tells you the routine is working.

Troubleshooting

Common problems

The team creates duplicates faster than you merge them

A process problem, not a data problem. Usually search is not being used before creating, often because it is genuinely bad at partial matches. Train the search step and make required fields minimal, since long creation forms are what drive people to skip the search.

Nobody logs activity

The most common CRM failure of all. If logging is manual it will not happen consistently. Connect email and calendar so activity is captured automatically, and reserve manual entry for things that genuinely need judgment.

The pipeline is full of dead deals

Because closing a deal as lost feels like admitting failure, so it sits in "proposal sent" forever. Set an automatic rule: no activity in 60 days moves it to a review stage. An accurate small pipeline is worth more than an impressive fictional one.

Two sources of truth

Someone keeps a personal spreadsheet because the CRM is missing something they need. That is feedback about the CRM. Find out what the spreadsheet has and add it, or the spreadsheet will quietly become the real system.

Expectations

How long this takes

4–8 hrs
First cleanup, small CRM
15 min
Weekly maintenance
2–5 days
A neglected CRM of 10k+ records

The initial cleanup scales with neglect rather than size. Ten thousand records maintained weekly is faster to clean than two thousand that have never been touched.

Fifteen minutes a week is about thirteen hours a year, which is far less than the cleanup it prevents — and unlike the cleanup, it is never urgent, which is exactly why it gets skipped.

The bridge

When it stops being worth doing by hand

The judgment in this work is small and specific: which record survives a merge, whether a quiet deal is dead. Everything else — surfacing duplicates, logging activity against the right record, attaching meetings to deals, flagging what went quiet — is assembly.

And it is weekly, forever, with no deadline forcing it. That combination is why almost every small business CRM is out of date.

Have the CRM stay current without anyone typing

Weekly CRM Maintenance cleans the records, logs the week's calls and emails against the right contacts, attaches meetings to open deals, and flags what went quiet — every Friday by 5 PM. $397 a month on its own, or part of Books + Pipeline at $1,697. Start with one closed month for $497 and see the underlying data first.

Get Your First Close — $497 See the plans
Questions

Common questions

How often should I clean my CRM?

Fifteen minutes weekly, not a big annual project. The annual version is how you end up with a five-day cleanup that undoes itself within a year.

Should I delete duplicate records?

Merge rather than delete, so activity and associations survive. In HubSpot merges are permanent and cannot be undone, so review before accepting rather than bulk-merging.

Which record should survive a merge?

Usually the one with the most activity history. A stale email on a rich record is a five-second fix; three years of context is not recoverable.

Do I need to fill in every field?

No. Fill a field only if a missing value stops you doing something. Guessed values are worse than blanks because they look like facts.

What causes duplicates in the first place?

Imports with inconsistent formatting, and people creating records without searching first — usually because search is poor at partial matches or the creation form is long.

Should I delete people who opted out?

No. Suppress them. Deleting removes the record of the opt-out itself, which is the thing you need to be able to prove.

Keep reading

Related

WG
William A. Green Jr.

Principal of William Delaney Consulting, in Wetumpka, Alabama. Twenty-seven years implementing enterprise systems where data quality was the difference between a working implementation and an expensive one. More about William →