Data Contracts: The Data Lake Is Not Broken, the Agreements Are

Why data contracts, ownership, and semantic rules must be defined before information enters an enterprise data lake.

This article is also available in Spanish.
Data Contracts: The Data Lake Is Not Broken, the Agreements Are

🏗️ The Data Lake Is Not Broken; The Agreements Are

The Problem Starts Before Data Reaches the Lake

When a data project goes wrong, the first thing a company points to is the platform. Maybe it was a bad choice, maybe the technical team didn't know how to implement it, maybe the vendor didn't deliver on their promises. That cycle of blame is very common in Colombia and almost always points in the wrong direction.

The chaos in a data lake starts long before the first piece of data arrives at the system.

A data lake is a central repository where a company stores large volumes of data in their original state, before processing them. Think of it like the storage area of a construction site where bricks, pipes, cables, and wood arrive without being labeled or categorized. The warehouse itself is not to blame if the project turns out poorly. The disorder starts in the order.

When the sales department sends a file where the net sales column includes returns and the finance department sends another where it doesn't, both enter the system with the same field name. The analyst spends days discovering the difference and the report comes back. It's like when bricks arrive without size specifications—they enter just like the rest and no one notices until they don't fit where they should.

This reprocessing happens in many Colombian companies with a regularity that nobody questions anymore. Reconciliation meetings become routine, manual controls don't balance out, and someone always has their own version of the report. Excel tables are built to compensate for what the system should show on its own. The disorder didn't start in the warehouse; it started in the order.

Data governance, which are the agreements and rules that define what information enters the system, who generates it, and by what quality criteria, is almost always postponed. It's assumed to be a technical issue that the data team will resolve on its own. But when that team arrives, the disorder is already inside, and separating it costs much more than controlling it at the entry point.

Putting things in order after everything entered without control is like trying to classify materials when four floors have already been built. That time is subtracted from projects, from decisions, and from confidence in reports. And that was exactly the problem you wanted to solve from the beginning.

First agree with your departments what each key indicator means before the data leaves the source. Then assign a responsible party for each domain, someone who answers for its quality and not just for generating it. Then define minimum rules for system entry, such as mandatory fields and standard formats. Finally, review those agreements each time processes change, because the construction site always receives new materials.