CRM & Data

Fixing duplicate records in HubSpot and Pipedrive

Fixing duplicates means matching rules, deliberate merge priority, and a check at entry that stops the next one being created.

Fixing duplicate records properly takes three things: real matching rules rather than a manual scroll, deliberate logic for which record survives a merge, and a check at the point of entry that stops the next duplicate forming. Skip the third and you will be back here in six months.

Duplicates are the most visible symptom of messy CRM data, and the most commonly mishandled, because most attempts focus on merging what already exists rather than stopping new ones being created.

Start with matching rules, not a manual scroll

Spotting duplicates by eye, scrolling a contact list for names that look similar, does not scale past a few hundred records. It also misses the less obvious cases: a company entered as "Acme Ltd" in one record and "Acme Limited" in another, or a contact with two slightly different email addresses across two records.

A proper approach uses matching rules. Exact matches on email address as a starting point, then fuzzy matching on company name and domain, to surface the genuine duplicates a manual scroll would miss entirely.

Merge logic has to decide which record wins

Once duplicates are identified, the harder question is which version of the data survives.

If one record has a complete set of fields and the other has more recent activity logged against it, the merge process needs a clear rule for which takes priority, not an arbitrary "keep whichever was created first". Getting this wrong quietly loses good data while fixing the duplicate problem, which is worse than leaving the duplicate in place.

The real fix happens at the entry point

Merging existing duplicates fixes what is already there. It does nothing to stop the next one forming the same way, usually because there is no check when a new contact or company is created that flags "this might already exist".

The effective fix adds that check at entry, so duplicate creation is caught before it happens rather than accumulating until the next clean-up project. This is the same reason data cleaning projects fail to stick: the clean-up treats the symptom and leaves the cause running.

HubSpot and Pipedrive both have native tools

Both platforms offer built-in duplicate management. Used properly, with sensible matching rules configured rather than left on defaults, they catch a meaningful share automatically.

Used blindly, accepting every suggested merge without checking which record should win, they can just as easily merge two genuinely different companies that happen to share a similar name. The tool is not the fix on its own. The judgment behind how it is configured is.

What a practical approach looks like

1. Real matching rules, not just exact-match defaults. 2. Clear priority logic for what survives a merge. 3. A check at the point of entry that stops new duplicates forming.

Skipping any one of these means duplicates either keep reappearing or the clean-up itself quietly damages good data. Where the backlog is already large, that is data cleaning work rather than a configuration change.

The pattern, in short

Fixing duplicates is not really about the merge. That part is mechanical once you know which records match. It is about deciding deliberately which data survives, and stopping the next duplicate before it is created.

The limit: some duplication is unavoidable, and chasing zero costs more than it saves. The target is a rate low enough not to distort your reporting.

Duplicates piling up in HubSpot or Pipedrive again?

Get in touch. We will fix what is there and stop it happening again. No obligation, just a straight answer.