Master data management on Databricks, and what AI does with bad data

26
Oct 2026
Webinar
26 October, 2026
15:00 to 16:00

How many customers do you actually have 

Data lands in Databricks from your ERP, your CRM, the web shop and whatever else the business has bought over the years. The pipelines run, Unity Catalog handles access, the first AI use cases are lined up and waiting. And still, nobody can agree on how many customers the company has, whether two supplier records are the same company, or which product hierarchy the reports should be using. 

That gap doesn't close by adding another pipeline. It closes by agreeing once on what a customer or a product actually is, then maintaining those definitions somewhere: matching records across systems, resolving the duplicates, and giving the business a practical way to fix things when they go wrong. That's what master data management does, and it's usually the part of the platform that gets postponed. 

For a long time, you could live with that. A report with duplicated customers is annoying, but an analyst notices and fixes the report or adds a caveat before it goes anywhere. An AI agent doesn't. It answers the question, sounds certain, and nobody goes back to check. The tolerance for messy master data has dropped considerably in the last two years, without most data teams having changed anything about how they work. 

In this session, element61 and Profisee go through how master data management fits into a Databricks platform. We look at where mastering belongs relative to bronze, silver and gold, how mastered data gets written back into Delta tables and governed through Unity Catalog, and how MDM sits alongside the data quality and cataloguing tooling you already have. We also spend time on scope, because most MDM programmes fail on ambition rather than technology. 

Profisee will demo their platform running against Databricks: modelling a customer domain, matching and survivorship, a data steward working through exceptions, and mastered data made available to AI tooling. 

What you'll take away 

  • Where master data problems actually show up, and how to recognise them in your own organisation 
  • A reference architecture for MDM alongside Databricks: medallion layers, Delta, Unity Catalog 
  • An honest comparison between building mastering logic yourself and buying a platform for it 
  • How to scope a first domain so you have something running in a quarter 
  • A working definition of “AI-ready data” that survives contact with a real project 

Who should attend 

  • CDOs, heads of data and analytics, data platform owners and data architects
  • Anyone who ends up accountable when the numbers in two reports don't match

Agenda

  • Welcome and introductions.
    Why element61 and Profisee are running this session together, and what you'll get out of it. 
  • Where AI projects run into master data.
    What changes when an agent rather than an analyst consumes your data, what duplicate customers and competing product hierarchies actually cost, and the symptoms we see most often at clients.
  • MDM in a Databricks architecture.
    Where mastering belongs in the medallion architecture, writing mastered data back to Delta tables, governance through Unity Catalog, and how MDM relates to the data quality and cataloguing tools you already run.
  • Live demo: Profisee on Databricks.
    Modelling a customer domain, matching and survivorship, a steward resolving exceptions, mastered data landing in Delta, and governed master data made available to AI tooling.
  • Where to start.
    Choosing a first domain, what a 90-day path looks like, what to measure, and the mistakes that make MDM programmes stall.
  • Q&A

Practicalities

Profisee