# Matching rules Source: https://neondeerdata.com/docs/platform/duplicate-detection/matching-rules/ Matching rules decide which People and Companies Duplicate Detection puts side by side as potential duplicates: records that share an email address, a domain, a social profile or a name, or a combination you define yourself. ## Finding a match is not deciding it A rule only finds records worth comparing. When a rule matches, the records are grouped into a set for review; nothing is merged and nothing changes in Attio. Whether the records really are the same person or company is a decision you make when you [review the set](https://neondeerdata.com/docs/platform/duplicate-detection/reviewing-duplicates/) and, if you are sure, [approve a merge](https://neondeerdata.com/docs/platform/duplicate-detection/merging-records/). Rules are therefore tuned to find candidates, not to prove identity. Two companies with the same name are a reasonable thing to show you, even though many will turn out to be different companies (see the [Northstar Labs example](https://neondeerdata.com/docs/platform/duplicate-detection/matching-rules/#example)). ## Built-in rules Every built-in rule starts switched on. A workspace admin can switch each rule on or off per object on the `Duplicate detection rules` page, and `Reset to defaults` returns them to the defaults. | Object | Rule | Matches records that have | | --------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | People | Email addresses | The same email address, compared the way the mail provider treats it (see [Email addresses](https://neondeerdata.com/docs/platform/duplicate-detection/matching-rules/#email)) | | People | Name | The same name, or a similar one (see [Exact and similar](https://neondeerdata.com/docs/platform/duplicate-detection/matching-rules/#exact-and-similar)) | | People | LinkedIn | The same LinkedIn profile | | People | Twitter | The same Twitter profile | | People | Facebook | The same Facebook profile | | People | Phone number + name | The same phone number and the same name | | People | Name + company | The same name and the same company | | People | Email domain + name | The same email domain and the same name | | People | Redirected email domain + name | Email domains that redirect to the same website, and the same name | | People | Probable same email (provider alias) | The same mailbox at two of one provider's own domains, such as gmail.com and googlemail.com | | People | Avatar URL | The same avatar URL | | Companies | Domains | A domain in common (see [Domains and redirects](https://neondeerdata.com/docs/platform/duplicate-detection/matching-rules/#domains)) | | Companies | Name | The same name, or a similar one | | Companies | LinkedIn | The same LinkedIn page | | Companies | Twitter | The same Twitter profile | | Companies | Facebook | The same Facebook page | | Companies | Redirected domains | Domains that redirect to the same website | Companies have no email rule because Attio Companies have no email attribute. Social profiles match however the URL was written: the scheme, `www` and tracking parameters are ignored. ## Exact and similar Every rule is exact except the similar half of Name. Exact means the values are the same after the cleanup described for each kind of value on this page, such as ignoring letter case in email addresses. A similar-name comparison first sets aside differences that rarely mean a different person or company: | Set aside | Example (illustrative) | | ------------------------------------------ | ------------------------------------- | | Word order | Lopez Ana and Ana Lopez | | Accents | José Núñez and Jose Nunez | | Punctuation | O'Neill and ONeill | | Common abbreviations | Intl and International | | A trailing legal suffix, on Companies only | Northstar Labs Inc and Northstar Labs | The remaining names are then scored against each other. How close they must be is the workspace's `Scan sensitivity` setting, set separately as `Company name similarity` and `Person name similarity`. The levels run from `Broad` through `Flexible`, `Balanced` and `Precise` to `Near identical`, which is the default. A broader level finds more potential duplicates, and more of them will be different people or companies. ## Signals that need a name A phone number, a company and an email domain are each shared by many unrelated people: a switchboard number, a large employer, a company's email domain. So for People these three only count together with the same name, in the rules Phone number + name, Name + company and Email domain + name. No built-in rule matches on a shared phone number alone. Both values must be present. A record that has the phone number but no name, or the name but no phone number, contributes nothing to that rule. ## Email addresses Addresses are compared the way the mail provider treats them, with letter case ignored. Where a provider documents that two spellings reach the same mailbox, they count as the same address. At any other domain, addresses must match exactly apart from case, because a wrong merge cannot be undone. | Address on one record | Address on the other | Same address? | | --------------------- | ---------------------- | ------------------------------------------------------------------------- | | ana.lopez@gmail.com | analopez+crm@gmail.com | Yes. Gmail ignores dots and plus tags. | | ana+crm@outlook.com | ana@outlook.com | Yes. Outlook.com drops plus tags. | | ana.lopez@outlook.com | analopez@outlook.com | No. Outlook.com keeps dots. | | Ana@Example.com | ana@example.com | Yes. Letter case is ignored. | | jason+1@example.com | jason@example.com | No. At other domains, plus tags and dots are kept. | | ana@googlemail.com | ana@gmail.com | Not for Email addresses. Probable same email (provider alias) matches it. | Probable same email (provider alias) covers providers whose domains reach one account, for example gmail.com and googlemail.com, or icloud.com, me.com and mac.com. It is a separate rule because the addresses really are different and one account only probably holds both. ## Domains and redirects Company domains are compared by the registered domain. The protocol, `www`, paths, query strings and subdomains are ignored, so `https://www.acme.co.uk/about` and `shop.acme.co.uk` are both `acme.co.uk`. Free email domains such as gmail.com are never used as evidence that two companies are the same. Redirected domains matches two different domains whose redirects land on the same website, for example an old domain that now forwards to a company's current one. Redirected email domain + name does the same for People's email domains. Some points about redirect checks: - The lookups are done by [checkredirects.io](https://checkredirects.io/), a redirect-checking service run by Neon Deer Data Labs. While checks are on, the Neon Deer platform sends it company website domains; [Sub-processors](https://neondeerdata.com/sub-processors/) lists what it receives. - A workspace admin can turn them off with `Check domain redirects`. While it is off, no domains are sent and Redirected domains finds no new matches. - Whether they are included depends on your plan. On some plans you use your own checkredirects.io API key: [sign up at checkredirects.io](https://checkredirects.io/), then paste the key into `Redirect lookup API key` under Settings » Domain redirects in the Neon Deer platform and choose `Save key`. Until a key is set, the rule shows `Off until set up`. - They run during scans, not during automatic detection in Attio. See [Scans and schedules](https://neondeerdata.com/docs/platform/duplicate-detection/scans-and-schedules/) for how a first scan can leave them for the next one. ## Protection against common values Some values sit on many records without meaning those records are one company or one person: placeholders such as "Test Co" or "N/A", or one office phone number shared by a whole team. Each rule has `Ignore very common values` (for rules with more than one condition, `Ignore very common combinations`). A value, or combination, that appears on more records than the rule allows is not used as evidence, so it cannot pull many unrelated records into one set. The records can still match each other through other values and rules. ## Custom rules Add your own rules when the built-in ones miss how duplicates appear in your data. A custom rule has one to four conditions, and all of them must agree for records to match. A condition can compare a field on the record itself, or follow one relationship, for example the name of a person's company. | Condition | Matches when | | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | Same value | Both records have the same value, ignoring case and surrounding spaces. | | Same value, or both empty | As above, and two records both missing the value also agree. Needs at least one other condition. | | Similar name | The values are close but not identical, at the sensitivity you choose. Compared as a name. | | Same domain | The registered domains are the same. | | Same email | The addresses are the same, as described in [Email addresses](https://neondeerdata.com/docs/platform/duplicate-detection/matching-rules/#email). | | Same email domain | The part after the @ is the same. This matches colleagues, so pair it with another condition. | | Same phone number | The phone numbers are the same. | | Same LinkedIn profile, Same Twitter profile, Same Facebook page | The profile is the same, however the URL was written. | | Same linked record | Both records link to the same record, not merely to records with the same name. | Location and owner fields cannot be used in a condition. Before you can save a new rule, or a change to what a rule matches, run `Preview matches` to see what it would find in your records. If you change the rule afterwards, the preview reads `Preview out of date` until you preview again. A rule that would match a large share of your records shows a warning, and a rule that is too broad cannot be saved. ## How matches become sets - An exact rule puts every record that shares the value into one set. Three records with the same email address form one set of three. - A similar-name set holds only records whose names all matched each other. If Ana Lopez is similar to Ana Lopes, and Ana Lopes to Anna Lopes, but Ana Lopez is not similar to Anna Lopes, the three do not form one set. A similar-name set can still hold more than two records when every pair matched. - Matches from different rules do not chain together. Each record belongs to one set at a time. If a record also matches records in another set through a different rule, that match appears as evidence on the set for you to review instead of joining the two sets. ### Why a set can hold records that look different Every record in a set matched the others through the rule that formed it. A set can still look like it holds strangers, for two reasons. A rule needs only one value to agree, so the records can differ on every other field; the review comparison shows those fields as `Different`. And a custom rule that follows a relationship matches through a linked record: the two records agree because they link to the same record, and the review screen marks this as an `Indirect match`. ## Example: two companies called Northstar Labs Your workspace holds these two Company records. The data is invented. | Field | Record 1 | Record 2 | | -------- | ----------------------------------- | -------------------------------------- | | Name | Northstar Labs | Northstar Labs | | Domains | northstarlabs.io | northstar-labs.co.uk | | LinkedIn | linkedin.com/company/northstar-labs | linkedin.com/company/northstar-labs-uk | ### 1\. The Name rule finds them A scan puts the two records in a set because their names are the same. Domains, LinkedIn and Redirected domains do not match: the domains are different and, in this example, neither redirects to the other. ### 2\. Review shows they differ In review, Name is `Same`, while Domains and LinkedIn are `Different`. Two separate websites and two separate LinkedIn pages are strong evidence of two companies that happen to share a name. You do not merge them. ### 3\. Decide what later scans should do | You choose | What changes | | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Mark not duplicate | These two records stop being flagged as duplicates of each other, in later scans too. The decision covers this pair only: if a third Northstar Labs record appears, it can still be matched with either of them. See [Reviewing duplicates](https://neondeerdata.com/docs/platform/duplicate-detection/reviewing-duplicates/). | | Change the rules | If same-name companies are mostly noise in your workspace, an admin can switch off the built-in Name rule for Companies and add a custom rule such as Similar name AND Same value on a field your team keeps reliably, for example a Headquarters country select. Later scans then match companies by name only when that field agrees as well. This affects every company, so companies that share only a name are no longer matched by any name rule. | Choosing `Remove from list` instead only hides the set until a later scan finds the same records together again, which here it would.