Data cleaning by people who know the real estate data model.
We clean real estate agency data every day. Doing it properly means knowing what a tenancy schedule, a comparable and a lease expiry actually are — which is exactly where generic data vendors stall.
What we clean
- Duplicate companies, contacts and properties
- Entity resolution across every record type
- Address standardisation and property matching
- Tenancy, lease and occupancy records
- Transaction and deal data across sales, leases and campaigns
- Financial records — commissions, rent rolls and trust accounting
- Lease and sales documentation and file references
- Contact validation, enrichment and decay removal
- Comparable sales and leasing evidence
- Centralisation into a single source of truth
- Migration into a new or existing CRM
This is the work we do most, and the work that decides whether everything else succeeds. Agency data is never clean — twenty years of duplicate companies, four spellings of the same tenant, phone numbers in five formats, addresses that match nothing official, and contacts nobody has touched since 2014.
Cleaning it is not a generic data exercise. It requires knowing the real estate data model well enough to tell which records are genuinely duplicates and which only look alike — that two entries at the same address are a landlord and a tenant, not one company entered twice; that a company name change after a sale is the same entity; that Suite 4 and Unit 4 are the same tenancy but Level 4 is not.
Get a general data vendor to do this and they optimise for row counts. They merge records that shouldn't be merged, flatten relationships that carry the commercial meaning, and hand back something tidy that has quietly lost the information your business runs on.
Why we can do it and others can't
We built and delivered ElevateIQ and rePrecinct. That means we designed the real estate data schema ourselves — entities, properties, tenancies, leases, campaigns and evidence, and the relationships between them. We clean against a model we understand rather than a spreadsheet we were handed.
We are cleaning agency data daily. This isn't a service line we added — it's the work underneath every product we've shipped.
What we actually do
Deduplication and entity resolution
Matching records that refer to the same real-world thing across inconsistent spellings, trading names, legal entity changes, mergers and manual entry error. Companies, contacts, properties and tenancies each resolve differently, and treating them the same is how data gets damaged.
Address standardisation and property matching
Normalising addresses to a consistent, verifiable format and matching them to the actual property — including the unit, suite and level distinctions that generic address tools flatten. This is what lets you group everything you hold about a single building.
Contact validation and enrichment
Email and phone validation, role and title normalisation, removal of decayed contacts, and enrichment where a record is worth keeping but incomplete. A database of 40,000 contacts where 12,000 bounce is worse than one of 28,000 that works.
Relationship repair
Reconnecting contacts to companies, companies to tenancies, tenancies to properties and properties to evidence. This is the layer that generic cleaning destroys and the layer that makes your data actually useful.
Migration
Field mapping and structured migration into whatever CRM you're moving to, with reconciliation so you can prove nothing was lost.
Specialist data sets
Transaction data cleaning
Real estate transactions leave traces across multiple systems — CRM deals, trust accounting ledgers, campaign registers and market evidence. We clean sales, leasing, property management and campaign records so buyer, seller, landlord, tenant, property and deal data resolve to the same entities, dates and amounts. The result is reliable deal history, pipeline and comparable evidence you can query without losing transactions in duplicates.
Financial records data cleaning
Commissions, rent rolls, invoices, disbursements and trust accounting entries often have mismatched property references, inconsistent entity names and legacy codes that no longer match your chart of accounts. We reconcile financial records back to the properties, tenancies and people they belong to, standardise descriptions and amounts, and flag where the source cannot be reconciled so your finance team can make decisions.
Lease and sales documentation data cleaning
Lease schedules, sale contracts, vendor statements, tenancy schedules and property files are usually stored as PDFs, spreadsheets or filenames that contain the only structured data. We extract and clean key fields — lease start and expiry, rent review dates, options, outgoings, purchaser and vendor details — and link them back to the right properties, tenancies and contacts so documents become searchable, not orphaned.
Centralisation
Cleaning is only useful if the data ends up somewhere consistent. We centralise cleaned records into a single source of truth — whether that is a CRM, a data warehouse, an internal database or a structured feed for AI platforms — with defined schemas, unique identifiers, audit trails and rules that stop the same mess from re-forming.
Why this is the prerequisite for AI
Language models don't fail loudly on bad data — they produce fluent, confident, wrong answers. If your CRM holds four versions of the same tenant, an agent asked to find your industrial occupiers in a corridor will return three of them and miss the rest, and nobody will notice until it costs something.
Clean, resolved, well-structured data is what makes retrieval reliable, makes agent output traceable back to a source record, and makes AI worth deploying at all. It's also the part every vendor skips, because it's slower and less impressive than a demo.
How an engagement runs
We start with a paid audit of your actual data — volumes, duplication rate, decay, structural problems, and what it can and can't currently support. You get that assessment in writing whether or not you continue.
From there we clean in stages, starting with the records that matter most to your current work, so value arrives early rather than at the end. You get the cleaned data and the pipeline, so it stays clean rather than decaying back to where it started within a year.
Three layers, cleaned differently.
Entities
Companies, contacts, landlords, tenants, purchasers and their relationships. Resolved by legal entity and trading history, not by name similarity alone.
Property
Buildings, tenancies, suites and levels. Standardised and matched so everything you hold about a single asset groups correctly.
Evidence
Sales and leasing comparables, campaign history and transaction records — the layer that makes analysis and AI retrieval reliable.
Straight answers.
Why does real estate data need specialist cleaning?
Because the meaning sits in the relationships. Two records at the same address might be a landlord and a tenant, not a duplicate. A company that changed names after a sale is still the same entity. Suite 4 and Unit 4 are the same tenancy; Level 4 is not. Generic tools optimise for tidy rows and flatten exactly the distinctions your business runs on.
How messy is too messy?
We have not yet seen a dataset we could not improve. Legacy exports, spreadsheets, shared drives, three CRMs that were never reconciled, decades of inconsistent manual entry — that is the normal starting point, not an unusual one.
Will we lose data in the process?
No. Nothing is destroyed. Merges are reversible and reconciled, and you receive a full record of what was matched, merged and changed so it can be audited or unwound.
Do you clean once, or keep it clean?
Both, and the second matters more. A one-off clean decays measurably within a year. We deliver the cleaned data along with the pipeline and rules that maintain it, so new records are resolved on entry rather than accumulating into the next cleanup.
Can you migrate the cleaned data into a new CRM?
Yes. Field mapping, structured migration and reconciliation into whatever system you are moving to. Because we have built CRM systems ourselves, we have worked on both sides of that migration.
How long does it take?
The audit is two weeks. Cleaning depends on volume and condition, but we work in stages and start with the records that matter to your current work, so you see usable results early rather than waiting for the whole database.
Do we need this before building AI agents?
Almost always, at least for the data the agent will touch. We assess what the specific workflow requires and clean only that, rather than insisting on a full data programme before anything useful ships.
Custom platforms, internal tools and reporting engines built end to end.
AI agentsAI agentsDomain-specific agents for lease abstraction, matching, reporting and triage.
Websites & SEOAgency websitesreal estate agency websites with CRM integration and an SEO strategy behind them.
AI advisoryAI advisoryStack selection, agentic workflow design and adoption across the business.
Send us a sample export.
We'll tell you what's actually wrong with it, what it would take to fix, and whether your data can support what you're planning to build.