Apache Spark is a multi-language engine for data engineering, data science and machine learning. It runs on a single machine or on clusters, including Kubernetes, and processes data in batches and in real time through the same API. You use Spark with Python, SQL, Scala, Java or R. It includes libraries for SQL and DataFrames, Spark Connect, streaming, pandas on Spark and machine learning (MLlib). Spark is software you install yourself, for example through pip or as a Docker container; it is not a hosted service.
Spark is a project of the Apache Software Foundation, a US non-profit organisation, and is released under the Apache License 2.0. Because you manage Spark yourself, you decide where compute and data reside: in your own data centre or with a European cloud provider. Commercial third-party Spark services fall under the jurisdiction of those providers, but the open source code keeps a move to self-managed operation possible.