Fictional Brenda Null leads data at a mid-sized logistics company. The CEO returns from a conference after watching a competitor’s chatbot answer a shipping question in four seconds flat, and by the next morning there’s a mandate: make the company “AI-driven.”
Brenda is asked to “do a full cleanup so we’re set up to move fast.” The standard corporate jargon for “nobody has decided how they solve the problem, but everyone would like to have already solved it.”
Six months later, Brenda has produced an impressive landscape: eleven years of support tickets, data from four abandoned CRM migrations, and several other sources that even I, the author of this story, can no longer remember. Nobody has used a single byte of it. Why? Nobody decided what it was for before the cleaning began.
This is where most dark data initiatives go wrong before a single dataset gets touched. The first rule of getting value out of dark data is that the use case comes first. Not afterward, clipboard in hand, checking whether the shelves happen to have what it needs.
The numbers, in case you doubted me doing my homework:
£12.7 billion: the final cost of the NHS's scrapped national patient-records programme, more than five times its original £2.3 billion budget, for a flagship system that reached under 10% of the trusts it was contracted to serve12.
80%: the share of corporate data-governance and cleanup efforts Gartner expects will fail by 2027. Why? Nobody tied the work to an actual business problem before starting it3
0.5%: the data share actually being analyzed. The other 99.5%, sitting unexamined, is what “dark data” means4.
Thoroughness is the trap. Most organizations built on a decade of unstructured files, shadow spreadsheets, and orphaned databases reach for the same logical instinct: map it all, clean it all, standardize it all, then figure out what it’s for. Responsible, yes, but it behaves like a black hole.
The Floodlight Instinct
I am calling it the floodlight approach because that’s the mental image my mind captures as I write: Switch on every light in the building at once, so nothing stays in shadow. It sounds responsible. It feels like the logical choice, the opposite of cutting corners. Clean everything first. Then, once it’s all spotless, start painting.
Before we go any further, it is worth understanding that floodlighting has no natural stopping point. There is always one more folder, one more spreadsheet to untangle, one more shared drive that has not been opened since 2016 but is apparently too important to delete.
And because no specific use case is driving the work, there is no real business case attached to it either. No operational problem being solved. No deadline that anyone outside the data team genuinely cares about.
So the project drifts.
It quietly consumes budget for years just because “we’re still cleaning the data” sounds perfect in a stand-up meeting. Until the day it doesn’t. The money runs out, a new VP arrives and asks what all this cleaning has actually produced, and the project simply stops. Quietly. Usually without anyone officially declaring it dead.
The data does not become any less dark.
What a Headlamp Actually Does
A caver entering an unmapped tunnel does not insist on illuminating the entire cave before taking the first step. They wear a headlamp. It lights exactly as much ground as the next step requires. Nobody (except me) calls this reckless. Cavers would probably call it Monday work, if cavers had performance reviews.
If cavers can do that, why can't we? Stop trying to clean everything in the abstract. Choose one use case and let it determine exactly which data needs attention. Build that, and only that. I’ll say this only in a whisper: everything else can remain dark, deliberately, until someone has a genuine reason to use it.
The Headlamp Playbook
The playbook has five steps. They happen in order. None is optional.
Discover cheap, everywhere. Before choosing a use case, take one quick pass across the organization. Find out what exists, roughly where it lives, and approximately how sensitive it is. Think metal detector, not archaeological excavation. This should take days, not years. Its only purpose is to ensure that valuable or dangerous data is not ignored simply because it is dark.
Prioritize on two axes, not one. Score every potential use case according to value and exposure. Value is what the organization gains if the use case succeeds. Exposure is what it could lose if something goes wrong first: a breach, a subpoena, or a terrifying yet polite letter from a regulator.
Build narrow, on purpose. Take the winning use case and build only the data pipeline it requires. Clean only the related data it needs. Define only the master records necessary to make it work. Test, iterate, and the cycle continues till it works as intended.
Promote only on repetition. Wait until a second use case asks for the same dataset: the same customer record, the same vendor definition, the same product hierarchy. That is the signal to make it shared. Data should earn shared status the old-fashioned way: by being requested twice.
Govern lightly, and automatically. Once something becomes shared, give it a small number of non-negotiable rules and let the system enforce them. Ownership, access, quality and retention should be built into the process wherever possible.
Run the loop again for the next use case. The playbook should not be a project plan with an end date and a wrap party. It should be a standing discipline.
Where This Gets Urgent: AI
For years, bad data governance was a quiet problem. A messy shared drive could sit undisturbed for decades. Nobody had to confront it unless they went looking for trouble.
AI changes that.
Point a model at the same shared drive, and it will not be quiet. It will search everything, find something, and answer with confidence, whether the document it found was approved yesterday or quietly replaced three years ago.
This is where the mess becomes expensive.
Picture Brenda’s data landscape again, except now it can speak directly to customers, quoting a returns policy the company retired years ago as though it were still gospel.
Gartner’s explanation for that 80% failure figure is the revealing part: governance programmes stall when they are not connected to a genuine business problem. AI adoption has just handed every department that problem at once.
The headlamp approach is no longer simply the sensible option. It may be the only one fast enough to keep up.
Leave the Rest Dark, but Build the Future Bright
Fifty-six kilometres east of Málaga, outside the town of Nerja, lies a cave system stretching almost five kilometres into the mountain. Only about a third is illuminated and open to visitors: a paved, hour-long route that passes a 32-metre-high column before reaching an underground chamber used for summer concerts.
Apparently, there is nothing better than five million years of acoustics.
The rest of the cave remains dark. That includes galleries containing some of Europe’s oldest cave art. They are not dark because anyone forgot about them or abandoned the work, but due to the simple fact that there is no reason to illuminate passages that nobody should be walking through yet.
Nobody leaves the illuminated section complaining that the tour was incomplete. The same principle applies to data, only at a larger and considerably more expensive scale.
Darkness may look like past failure. From another angle, it is simply everything you do not need yet.
UK Parliament, Hansard, "Information Technology (NHS)" debate, 14 June 2011 (account of the 2002 Downing Street meeting); UK Parliament, Committee of Public Accounts, "Dismantled National Programme for IT in NHS"
UK Parliament, Committee of Public Accounts, "The National Programme for IT in the NHS" (contracted target of 220 trusts in the North, Midlands and East); Digital Health, "Lorenzo: the end of the beginning," July 2016 (actual deployment figures)
Gartner, "Gartner Predicts 80% of D&A Governance Initiatives Will Fail by 2027, Due to a Lack of a Real or Manufactured Crisis," February 2024
IDC/EMC, "Big Data, Bigger Digital Shadows, and Biggest Growth in the Far East," Digital Universe Study, December 2012 (EMC-sponsored; figure is over a decade old and cited here as a directional estimate, not a current measurement)






