Apache Hudi is an open data lakehouse platform that adds database functionality to data lakes while keeping data in open file formats. It supports tables, transactions, upserts and deletes, indexes, clustering and compaction, as well as incremental processing for streaming and batch workloads. It is a software library rather than a hosted service: you run it inside your own data platform.
Hudi was developed at Uber, open-sourced in 2017 and has been a top-level project of the Apache Software Foundation, a US non-profit organisation, since 2020. It is licensed under the Apache License 2.0. Because Hudi has no hosting of its own, you decide where data and metadata live: in your own data centre or with a European storage and cloud provider. Open source code and an open table format limit lock-in, and with it dependence on a single supplier or jurisdiction.