When data is scattered across dozens of systems, sooner or later the question of centralisation arises. But where to? A warehouse, a data lake, or both? The choice affects cost, flexibility, and who can use the data.

In this article we compare these approaches, explain when each fits, and show why a lakehouse is a natural default for many. We also give concrete questions to help you decide.

When data is scattered, the question is: centralise it into a warehouse, a data lake, or both? The answer depends on what you do with the data.

Warehouse: structure and queries

A warehouse suits structured data and reporting well. Data is modelled and query-ready, but flexibility for new kinds of data is limited.

Data lake: flexibility and ML

A data lake stores raw data in any format, which suits machine learning and experimentation. The cost is that governance and quality take more work.

Lakehouse: both

A lakehouse combines warehouse governance with lake flexibility. For most organisations this is now a natural default.

Do not choose by fashion but by use cases. Start from what reporting and ML actually need.

Cost and governance

A data lake is cheap to store but expensive to govern badly. Without a catalogue and rules it becomes a "data swamp" no one trusts. A warehouse forces structure earlier, which is both a constraint and a benefit.

Phase the migration

Do not migrate everything at once. Start from one valuable dataset, prove the benefit, and expand. A big-bang move fails more often than a series of small, measured steps.

Decision questions

Ask: Is the data structured or mixed? Is it needed for reporting, machine learning, or both? How many users and at what skill level? The answers guide the choice more than any trend.

Performance and users

Different solutions serve different users. Business analysts who query with SQL benefit from a warehouse's speed and structure. Data scientists who experiment and train models need a lake's flexibility and raw data. Consider who actually uses the data and with what skills — this guides the architecture as much as the technical requirements.

Avoid over-engineering

The temptation to build a fancy, all-encompassing data platform is strong — and an expensive mistake if your needs are modest. For many small and mid-sized organisations a well-run warehouse goes a long way. Start from what you truly need now and expand as the need grows. The architecture is allowed to grow with the business — it does not have to be complete on day one.

Common pitfalls

Most failures come not from technology but from design. Typical mistakes are: starting with too large a scope, lacking clear goals, ignoring people and processes, and forgetting maintenance right after launch. Choosing the right architecture succeeds when you keep the solution simple, measure the result, and correct course quickly. Complexity that is not needed is always a risk.

How to measure success

Success cannot be judged without a metric defined in advance. Set a baseline before you start, choose a couple of clear figures tied to the business, and track them regularly. Avoid metrics that look good but do not change decisions. A good metric answers the question: did this work deliver real value, and how much? When the answer is a number, the conversation turns from opinions into facts.

Summary and next steps

The key message is simple: start from a clear need, keep the solution manageable, and measure the result. Do not chase perfection but a direction that delivers value and improves over time. If you would like to discuss how this applies to your own situation, we are happy to help with an assessment and planning the first steps.