Why a generation of progress still hasn’t produced a shared meaning — and why that matters more than ever.
Note: The content of this article is a synthesis of a methodical, theme-based literature review that informed the research. Most topics were not directly examined in the quantitative analysis.
Ask ten data professionals what “data governance” means, and you’ll get twelve answers, each delivered with confidence, sometimes enthusiasm, and occasionally a hint of existential dread. This isn’t because the field is confused or unserious. It’s because data governance grew up backwards — emerging from decades of practical necessity, organizational dysfunction, regulatory panic, and technology shifts long before anyone paused to ask the fundamental question: What is this thing we keep building?
Everyone claims to “have governance,” yet very few organizations can describe it clearly, execute it consistently, or sustain it when things change.
This isn’t a surprise.
Data governance evolved quickly, reactively, and under pressure.
It wasn’t designed — it was assembled.
And as digital transformation and AI raise the stakes, the gaps in definition, practice, and resilience are becoming too costly to ignore.
To understand why data governance remains so amorphous, we need to understand how governance got here in the first place.
Before Data Governance Had a Name: Borrowing From IT Governance
In the 1960s and 70s, governments and large organizations were wrestling with a new problem: electronic data. With the rise of mainframes and early databases, managing data became both essential and chaotic. Agencies in the US and UK collaborated with universities, corporations, and technology manufacturers to create some of the first data governance-like structures — data dictionaries, access controls, quality checks, and reporting standards.
None of this was called “data governance.”
It was simply the work required to keep early systems from collapsing.
By the 1990s, data warehousing and business intelligence created new pressures. Organizations wanted integrated data for decision-making, but inconsistent definitions, quality issues, and siloed ownership created new problems. Regulations such as HIPAA (1996) and the EU Data Protection Directive (1995) required organizations to improve control over personal data. Still, governance remained informal, implicit, and inconsistent across industries.
The need for governance existed.
The language to describe it did not.
The Mid-2000s: Dragging Data Governance Into the Light
The mid-2000s didn’t define what data governance is, but they did mark the era when the term “data governance” entered our professional lexicon, at least amongst data nerds.
Authors like Gwen Thomas, John Ladley, and Sunil Soares were among the first to publish books, articles, and guidance that used the phrase “data governance” explicitly. Their influence wasn’t in defining the discipline but in giving organizations a vocabulary for problems they already recognized:
unclear accountability,
inconsistent practices,
siloed authority,
decision-making driven by personalities rather than process.
Their work named the dysfunction and helped the term “data governance” take root.
Author’s note: If you’re looking for books on Data Governance, Ladley’s book has been my ‘Bible’ for the majority of my career, and in my opinion, it is the most comprehensive. Thomas’s “Alpha Males and Data Disasters” is my favorite — a 222-page mic-drop moment for practitioners. The latter is sadly out of print and very difficult to get. If you find or have one, I will pay a hefty ransom for a copy. I had considered stealing it from an Oklahoma University Library, where I borrowed it, but ethics prevailed.
Around this time, the industry began codifying data governance ideas into actionable frameworks, giving them legitimacy. The first edition of the DAMA DMBoK positioned data governance as a central component of data management, formally embedding the term into a widely adopted industry reference model. It didn’t define governance universally, but it anchored the term inside the profession. The Data Governance Institute (DGI) offered an early definition focused on decision rights and accountabilities, along with templates and guidance. It helped practitioners articulate governance, even if the definitions reflected a specific organizational lens. IBM’s introduction of a Data Governance Maturity Model in 2006 demonstrated that data governance was moving beyond practitioner intuition and that organizations were now seeking formal, repeatable methods.
As data problems became more visible — quality issues, uncontrolled access, inconsistent reporting, regulatory pressure — vendors, consultants, and technology providers began producing services, glossaries, and frameworks using the same terminology. Data Governance could be profitable.
The mid-2000s brought linguistic convergence, not conceptual clarity. But that convergence mattered: the field now had a shared label, even if meaning lagged.
2008 Changed Everything: Governance Becomes Compliance
The 2008 financial crisis revealed that major institutions couldn’t reconcile their own data — a catastrophic failure when billions hinged on reporting accuracy. Organizations’ inability to manage data (and the harm it could cause) was thrust into the spotlight, just as cloud computing, “Big Data”, and an era of accelerated digital transformation were starting to materialize. Regulators responded with regulation (as they do).
In a matter of a few years, amid the hangover of the “Great Recession” and increased concerns about privacy and data subject rights, organizations faced an onslaught of regulations, including Dodd-Frank, Basel III, EMIR, MiFID II, and GDPR. These frameworks required the core tenets of data governance, such as transparency and accountability, and a situational focus on foundational data management capabilities, including metadata, lineage, and data quality management, but only for some processes. Data governance became synonymous with compliance.
This fundamentally changed data governance:
From: a coordination mechanism focused on clarity, quality, and decision rights.
To: a compliance mechanism focused on control, defensibility, and risk mitigation.
The term solidified.
The definition and intent skewed.
2010s–2020s: The Era of Infinite Variations
The explosion of cloud computing, SaaS ecosystems, hybrid architectures, big data platforms, and machine learning, and a lot of marketing added layer upon layer to the narrative around data governance, conflated it with data management, and introduced variations (e.g., Big Data Governance, Cloud Data Governance, AI Governance).
Data governance now could mean:
lifecycle management
privacy and security
data quality
metadata and lineage
stewardship and ownership
data integration and interoperability
compliance and risk management
data strategy
data fluency and literacy
This drives me crazy — yes, I’m ranting a bit, and no, they wouldn’t let me put this in the dissertation. I believe it is a symptom of data management’s industry-driven nature and reactive ethos, which, combined with a lack of understanding of what data governance actually is, leads to unnecessary change in response to macroenvironment shifts, undermining process resilience and sustainability. Does cloud modernization introduce new methods for the data management function? Sure. Does the impact on the management function, in turn, create new decisions that data governance must support? Absolutely. Does it wholly invalidate the principles, policies, authority, and decision-rights that were established for on-prem systems and require a “cloud” qualifier? Absolutely not.
I will spare you the rant on AI Governance for now. We’ll cover that in another article, but I’ll share my perspective — AI is just another use of data. I suspect the reason we might have a different concept of AI governance is that data governance has thus far not been resilient enough to handle the scale of this macroenvironmental shift (and to sell stuff, of course). And every underlying data management vulnerability — poor lineage, unclear ownership, weak controls, inconsistent definitions — became an AI risk.
AI didn’t change data governance.
It exposed it. Then rebranded.
Why We Still Don’t Have a Shared Definition
The definitional ambiguity isn’t a failure. It’s structural.
Different communities built governance for their own needs: security wanted control, regulators wanted compliance, architects wanted metadata, business users wanted quality, data scientists wanted access, and vendors wanted sales.
Current perspectives struggle with the interdisciplinary nature of data governance and the need for systematic collaboration: Effective data governance requires data literacy, legal expertise, risk management skills, organizational change management, business knowledge and acumen, and technical skills, among others. In other words, lots of people need to get along and work for the greater good, beyond their individual performance objectives — and not just for a project, in perpetuity. Most organizations do not have the systems in place to support these operations or incentivize work for the greater good. So, who you hire to lead your data governance program and who the most vocal leaders are probably have a lot to do with how it “looks” and “feels” in your organization.
Governance Still Lacks a Strong Empirical Foundation: Academic work tends to be theoretical, while practitioners rely on experience, improvisation, and vendor guidance that often conflates governance with tooling and generalized frameworks. The result — organizations often implement governance by trial and error. (I’ve been here, this is me!)
The Human aspects of Data Management are often overlooked. Data Governance failures rarely stem from missing tools; instead, the culprits are: unclear roles, lack of communication, insufficient knowledge, cultural resistance, and shifting priorities, to name a few.
Why It’s So Difficult
My close friend and mentor, Daragh O Brien, wrote an excellent piece describing data management as a “Wicked Problem” — it’s a perfect characterization. Read it.
Many data governance problems behave like classic wicked problems:
Recursive, cross-functional, stubborn, and resistant to permanent resolution.
Data is socially constructed.
People interpret it differently.
Systems encode imperfect representations of the world.
AI systems amplify these imperfections at scale.
Governance requires legal literacy, technical fluency, risk awareness, organizational design, and information architecture. No single team holds all of this, and no single leader owns all the incentives.
That’s why governance programs collapse during reorganizations, leadership turnover, or shifting priorities. They often depend on a small group of subject-matter experts whose responsibilities sit outside governance itself. When those individuals are stretched thin or reassigned, governance resilience disappears.
And with it, performance deteriorates, data debt grows, and the cycle of reinvention begins again.
Four Forms of Governance (Why They All Exist — and None Wins)
Across industries, data governance takes four recognizable forms. These forms aren’t theoretical categories — they reflect the very real motivations that lead organizations to build governance in the first place. And most organizations adopt more than one form at a time, often without realizing it. That’s part of why governance feels inconsistent: different people are practicing different versions of it under the same name.
Here’s how these forms play out.
Nearly every organization treats data as an asset (or so they say)— but how they manage that asset varies dramatically. This is why data governance shows up in four distinct forms. These forms coexist because different parts of the business are optimizing for different outcomes: quality, risk, value, or operational performance.
Form 1: Governance for Quality (Data as an Asset Whose Value Depends on Its Condition)
In this form, governance focuses on creating high-quality, consistent, interoperable data that the organization can trust, use, and build on.
Teams working in this mode think of data like a precision instrument: the better the calibration, the more valuable the output.
Their priorities include:
delivering reliable and accurate information products
ensuring interoperability across systems and business lines
strengthening stewardship roles tied to quality outcomes
enabling better customer experiences and analytics
productizing or monetizing data responsibly
This is data as an asset whose value rises or falls with its quality.
Form 2: Governance for Risk and Compliance (Data as an Asset That Carries Exposure)
Here, the emphasis shifts from value to vulnerability.
Risk managers view data as an asset that must be protected because:
inaccurate data can create regulatory violations,
misuse can trigger legal exposure,
poor controls can cause reputational damage.
Stewardship in this form is oriented around:
meeting regulatory requirements
preventing breaches, misuse, and reporting errors
managing operational and regulatory risk
ensuring defensibility
Quality still matters — but mainly because poor quality creates risk.
This is data as an asset whose value is threatened by risk events, requiring control.
Form 3: Governance for Strategic Value (Data as an Asset That Produces Return)
In this form, governance becomes an instrument of strategy and value creation.
Leadership teams see data not just as something to protect or improve, but as something that:
drives revenue,
powers new capabilities,
strengthens competitive advantage,
and justifies investment.
Priorities include:
Executing the data strategy
allocating resources toward high-value domains
enabling digital transformation
evaluating the ROI of data initiatives
building scalable, future-ready capabilities
This is data as an asset with direct financial and strategic value — something that belongs on the balance sheet metaphorically, if not literally.
Form 4: Governance Embedded in Operations (Data as an Asset Managed Through Design and Performance)
This form treats governance as a built-in property of the data ecosystem rather than a separate function. However, it often incorrectly conflates the governance capability with the management function.
Examples include Domain Ownership in Data Mesh, where teams are accountable for:
producing high-quality, well-documented data products
maintaining operational KPIs (accuracy, latency, uptime)
embedding controls and standards directly into pipelines
assuring interoperability through platform services
In this form, the value of the data asset is protected and grown through operational excellence rather than oversight structures.
This is data as an asset, with performance measured through technical and operational KPIs.
The Real Issue: Competing Motivations Under One Label
Every form is grounded in the idea that data is an asset — but each focuses on a different dimension of what it means to manage that asset:
Form 1: quality → maintain the asset
Form 2: risk → protect the asset
Form 3: strategy → invest in the asset
Form 4: operations → operate the asset effectively
In reality, organizations rarely choose one form.
They practice all four at the same time, often without realizing it or the skills or resources to do so effectively.
The differences may seem nuanced, but when quality teams prioritize trust and consistency, risk teams prioritize control and defensibility, strategy teams prioritize enablement and ROI, and engineering teams prioritize automation and performance — can you really make everyone happy with a handful of analysts and part-time stewards?
And because the motivations are rarely made explicit, governance becomes confusing, fragile, or internally contradictory, with each group pulling the data office in a different direction.
This is how organizations end up with governance that feels chaotic, fragile, or perpetually “in redesign.”
What Organizations Actually Do (Patterns From the Field)
Across industries and roles, a few patterns consistently appear in my quantitative research:
Most organizations technically “have data governance” — programs, policies, stewards, and frameworks exist.
But structure does not guarantee behavior.
Where a formal policy exists, enforcement becomes more than three times as likely.
The most common challenges are not technical.
They are human:
insufficient resources and expertise (64.1%)
lack of awareness (63.8%)
limited knowledge (51.1%)
cultural resistance (47.1%)
And even if you believe technology can solve all of our problems, we identified a significant lag — by more than 30% on average — between awareness of industry innovations and their implementation.
We found that organizations that perceive governance as valuable tend to have:
stronger data cultures,
clearer decision-making,
better metadata practices,
and tighter alignment with strategy.
Organizations that struggle often rely on frameworks and tools without addressing the underlying systemtic and social factors that complicate data governance.
What Leaders Should Do Now
While a common definition would demonstrate to the world we know what they heck we’re doing — it’s really not all that important. In practice, we don’t need a definition (and certainly not another one). What we do need is a way out of the cycle of reinvention — something sturdier than a new framework, cleaner than a new slide, and more honest than pretending data governance is a tooling problem, or that data management is a data problem.
Three actions matter more than anything else:
First, decide what you actually want data governance to accomplish.
Not in a broad, inspirational sense. In a specific, operational sense. Pick the dominant motivation for your program: improving quality, reducing risk, creating value, or strengthening operational performance. You’ll still practice all four, but making the primary purpose explicit changes everything — funding, staffing, decision rights, and expectations. It also stops the quiet turf wars that derail progress.
Second, resource data governance like you mean it.
The research is blunt: programs collapse when they rely on part-time stewards, heroics, or borrowed staff. Leaders consistently underestimate the coordination burden, the communication load, and the ongoing change management required. If data is an asset, treat the people who manage it like asset managers, not volunteers. Build persistent roles. Train them. Give them authority. Make governance as routine as finance, audit, or risk — not an extracurricular activity.
Third, fix the social system, not just the technical one.
Every structural weakness uncovered in the research — unclear roles, weak accountability, cultural resistance, knowledge gaps — sits upstream of every technical failure. Metadata tools don’t eliminate political tension. Lineage graphs don’t fix conflicting incentives. Automated controls don’t replace shared understanding. Leaders who focus exclusively on technology will spend the next decade wondering why their expensive platforms didn’t close the trust gap. Governance works when people work together. That requires incentives, education, norms, and leadership support — not just software.
A final point: governance is not something you “implement.” It’s something you run. It evolves as strategy evolves, and it must withstand reorganizations, turnover, and the churn of new technologies. Treat governance as an operating capability, not a project with a finish line.
Data governance didn’t define itself because we never really paused long enough to think about it. Data Governance leaders have that opportunity now. Start by choosing the purpose, building real capacity, and treating the social architecture with the same seriousness as the technical one. That’s how governance becomes resilient — without needing a rebrand every time the world changes.
How Data Governance Evolved Without Ever Defining Itself was originally published in My Column Has NULLs on Medium, where people are continuing the conversation by highlighting and responding to this story.






