A data lake stores large amounts of raw structured and unstructured data for later processing and analysis.

Raw data at broad scale

A data lake can hold logs, documents, sensor readings, images, and tables before every use is known. Cheap storage makes it practical to retain source data for future processing.

Governance prevents a swamp

Catalogs, ownership, access controls, quality checks, and lifecycle rules make a lake usable. Without them, teams may have plenty of files but no reliable way to discover or interpret them.