What is a Data Lake, and Do You Need One?
A plain explanation of data lakes, how they differ from a warehouse, and an honest answer on whether your organisation needs one.
Volume and variety have since pushed them to their limits, which is why organisations have moved towards a different model: the data lake.
Where all the data actually goes
Digital organisations generate and receive an enormous variety of data — internal and external, structured and unstructured. It matters for trend analysis, record keeping and understanding what users actually do. The practical question is where it lands.
Traditional data warehouses were built for a world of predictable, structured inputs. Volume and variety have since pushed them to their limits, which is why organisations have moved towards a different model: the data lake.
The symptoms are recognisable: nobody can find what exists, nobody trusts what they find, and two analyses of the same question return different answers.
What a data lake is
A data lake is a consolidated, centralised repository holding many forms of data in their native format, drawn from disparate applications across an organisation.
That phrase is the whole distinction. A warehouse requires data to be structured on the way in, which means deciding in advance what questions you will ask of it. A lake accepts data as it is and defers that decision — so analysts and data scientists can locate and interrogate large volumes quickly, including data nobody had a use for when it arrived.
Why organisations move to one
- Volume. Data is outgrowing systems designed for less of it.
- Variety. Documents, logs, images and telemetry do not fit neatly into relational tables.
- Cost. Storing raw data is markedly cheaper than modelling everything before you know whether you need it.
- Flexibility. New questions do not require a new pipeline.
- Analytics and machine learning. Both need volumes of raw historic data that a warehouse tends to have already discarded or aggregated away.
The failure mode: the data swamp
The risk is well documented and worth taking seriously. A repository that accepts anything, with no catalogue, no ownership and no quality standards, becomes a place data goes to be forgotten — the "data swamp".
The symptoms are recognisable: nobody can find what exists, nobody trusts what they find, and two analyses of the same question return different answers.
Best practice that avoids it
The practices that keep a lake usable are governance disciplines rather than technical ones:
- Catalogue everything, so data can be found and its origin understood
- Assign ownership for each domain — unowned data becomes unmaintained data
- Define zones (raw, curated, consumable) rather than one undifferentiated pool
- Apply security and access control at the point of ingestion, not later
- Track lineage, so any figure can be traced to its source
Does your organisation need one?
Not every organisation does. If your data is modest in volume, largely structured, and answering a stable set of questions, a warehouse is simpler and probably sufficient.
The case for a lake gets strong when you have significant unstructured data, when you cannot predict what you will need to ask, or when you are building machine learning that requires history nobody thought to keep.
Get the full guide
Tell us who you are and the guide opens straight away, all 26 pages, to read online or download as a PDF.
Download the guide
What housing providers say
Named people at named organisations, in their own published words.
“We realised that we had a gap around the golden thread of data, in terms of the availability and accessibility of the data we held in seventeen different systems… That meant colleagues could immediately see all non-compliant properties.”
Jake Le Page Head of Building Safety Regulations Notting Hill Genesis
“Neo brought strong Dynamics 365 expertise, worked collaboratively with our internal teams and applied Microsoft best practice within a live operational environment… We would be pleased to recommend them as a Microsoft Dynamics 365 partner within the housing sector.”
Wayne Human Head of IT Change Sage Homes
“The successful deployment of this solution has significantly increased transparency and improved the operational efficiency of our contact centre, allowing us to deliver greater value to our customers.”
Philip Wragg Infrastructure Programme Manager
“The Neo Technology model allows us to scale our development capacity, accelerating our transformation programmes while future-proofing our business, while achieving substantial industry cost savings.”
Group CIO Notting Hill Genesis
Related proof
See how this works in practice
Thirty minutes, walked through by the people who deliver it. Bring your questions.