apache/glutenApache Gluten Docker镜像是预配置的容器化环境,用于快速部署和运行Apache Gluten加速的Apache Spark集群。Gluten作为Spark的查询加速器,通过集成高效的执行引擎(如Meta Velox、ClickHouse)和优化技术(向量化执行、列式存储),可将Spark SQL查询性能提升2-10倍,同时保持与原生Spark的兼容性,降低迁移成本。
bashdocker run -d \ --name gluten-spark \ -e SPARK_MASTER="local[*]" \ -e GLUTEN_BACKEND="velox" \ -e SPARK_DRIVER_MEMORY="4g" \ -e SPARK_EXECUTOR_MEMORY="8g" \ -p 4040:4040 \ apache/gluten:latest
| 环境变量 | 描述 | 默认值 |
|---|---|---|
SPARK_MASTER | Spark集群Master地址(如spark://host:7077或local[*]) | local[*] |
GLUTEN_BACKEND | 选择Gluten后端引擎(支持velox、clickhouse、arrow) | velox |
SPARK_DRIVER_MEMORY | Spark Driver内存大小 | 2g |
SPARK_EXECUTOR_MEMORY | Spark Executor内存大小 | 4g |
GLUTEN_LOG_LEVEL | Gluten日志级别(DEBUG/INFO/WARN/ERROR) | INFO |
SPARK_SQL_EXTENSIONS | Spark SQL扩展类(启用Gluten需设置为io.glutenproject.sql.GlutenSparkSessionExtension) | 自动配置 |
通过spark-defaults.conf自定义Spark和Gluten属性,可通过挂载配置文件实现:
bashdocker run -d \ --name gluten-spark \ -v ./spark-defaults.conf:/opt/spark/conf/spark-defaults.conf \ apache/gluten:latest
示例spark-defaults.conf配置:
ini# 启用Gluten加速 spark.sql.extensions io.glutenproject.sql.GlutenSparkSessionExtension # 配置Velox后端内存限制 spark.gluten.velox.memory_pool.size 16g # 启用向量化执行 spark.gluten.sql.columnar.backend.velox.vectorized true # 优化Shuffle性能 spark.shuffle.manager org.apache.spark.shuffle.sort.ColumnarShuffleManager
yamlversion: '3' services: gluten-spark: image: apache/gluten:latest container_name: gluten-spark environment: - SPARK_MASTER=local[4] - GLUTEN_BACKEND=velox - SPARK_DRIVER_MEMORY=8g - SPARK_EXECUTOR_MEMORY=16g ports: - "4040:4040" # Spark UI端口 - "***:***" # Spark History Server端口 volumes: - ./data:/opt/spark/data # 挂载数据目录 - ./spark-defaults.conf:/opt/spark/conf/spark-defaults.conf restart: unless-stopped
http://localhost:4040bashdocker exec -it gluten-spark /opt/spark/bin/spark-sql \ -e "SELECT count(*) FROM parquet.`/opt/spark/data/sample.parquet`"
bashdocker logs gluten-spark | grep "Gluten backend initialized with"
clickhouse后端需提前部署ClickHouse集群)spark.gluten.velox.memory_pool.size)
manifest unknown 错误
TLS 证书验证失败
DNS 解析超时
410 错误:版本过低
402 错误:流量耗尽
身份认证失败错误
429 限流错误
凭证保存错误
来自真实用户的反馈,见证轩辕镜像的优质服务