Every property team we talk to has the same spreadsheet. It lives somewhere on a shared drive, it has been touched by six different people over three years, and no two columns use the same naming convention. The rent roll for one asset calls it “Tenant Name.” The one from the asset across the road calls it “Occupier.” The one inherited from the previous manager just says “T.”
This is not an unusual situation. It is the normal situation. And it is exactly where an AI engagement either proves its worth or falls apart entirely.
The clean-data fantasy
There is a version of the AI pitch that goes: “Once your data is clean, the AI will unlock serious value.” We hear this framing a lot. It sounds reasonable until you ask the obvious follow-up: when, exactly, does the data get clean? Who is cleaning it? At what cost, and to what end?
In practice, data cleaning is either already happening as a continuous operational task, or it is not happening at all. The firms that have been running disciplined data hygiene for years tend not to need much convincing about AI. They already have a fairly good picture of their portfolio. The firms that most need an AI system to help them see across their assets are, almost by definition, the ones whose data is messy.
The interesting question is not “how do we clean the data first?” It is “what can we extract from the data as it actually exists?”
“The data is never ready. The question is whether your AI system is.”
Teddy James, Tercero Analytics
What the data actually looks like
We did a data review for a firm managing eleven assets across two geographies. Between them, they had rent rolls in four different formats. Some were Excel files with colour-coded cells carrying meaning (which no extraction tool could read). Some were exports from a property management system that had been discontinued. One was a PDF that had been scanned rather than exported, so it was essentially a photograph of a table.
The occupancy figure looked different in every source. One system reported it by unit count. Another by square footage. A third used a blend that nobody could quite explain but that appeared to be weighted by estimated market rent. The same building, queried across three sources, gave three different occupancy numbers.
None of this was anyone’s fault. It was the accumulated result of sensible local decisions made by different people over time. The previous asset manager used the system they were trained on. The data analyst built the Excel model that answered the questions the investment committee was asking at the time. The fund administrator exported what the reporting template required.
The data was not corrupt. It was rich. It just required a system capable of reasoning about inconsistency rather than one that assumed it away.