Overview
The SkaleData Superset image: preinstalled database drivers, how to extend it, and how to deploy a custom image.
SkaleData publishes a maintained Superset image at
ghcr.io/skaledata/superset.
It's the official apache/superset image with drivers for common warehouses
and databases, so you can connect to the databases in
Preinstalled databases without building an image of
your own.
New SkaleData Superset instances run this image by default. An existing instance moves to it the next time it's applied or restarted: a Save & Apply, a Restart, or a cluster apply.
This page describes ghcr.io/skaledata/superset:6.1.0-dc3cdfa, the default
image.
What's in the image
apache/superset6.1.0, unchanged. Adding the drivers doesn't change the version of any package that the upstream image already has.- The database drivers in Preinstalled databases.
- The shared libraries those drivers load: the MariaDB client for MySQL, Cyrus SASL for Hive and Spark SQL, and the Firebird client library.
- The packages that Superset's built-in Model Context Protocol (MCP) server
needs, so that the
superset mcp runcommand works:fastmcp3.2.4,mcp1.29.1, and their dependencies, such asstarlette1.3.1 anduvicorn. They aren'tapache-supersetextras, so the tables on this page don't list them. The image doesn't start the MCP server.
The image is published for linux/amd64 only. Cluster nodes are amd64. On a
computer with Apple silicon, Docker runs the image under emulation.
Tags
| Tag | Moves? | Use it for |
|---|---|---|
6.1.0 | Yes. It moves to each new release of the image for Superset 6.1.0. | Custom images that should pick up driver fixes when you rebuild. |
6.1.0-<sha7> | No. <sha7> is the commit the image was built from. | Reproducible builds. |
latest | Yes. It follows the newest supported Superset version. | Trying the image out. Don't build production images on it. |
The default image is 6.1.0-dc3cdfa, pinned by digest:
ghcr.io/skaledata/superset:6.1.0-dc3cdfa@sha256:b2328a3d6c564773414993df8bcf27cb696d22fa03559cd12c9c2ea2b216db34The package is public, so pulling it doesn't need credentials:
docker pull --platform linux/amd64 ghcr.io/skaledata/superset:6.1.0-dc3cdfaPreinstalled databases
Superset lists these databases in its Connect a database dialog. The
extra column names the apache-superset extra that the image installs for
each one. The driver packages column names the Python packages that provide
the connection.
| Database | Extra | Driver packages |
|---|---|---|
| Amazon Athena | athena | pyathena |
| Amazon DynamoDB | dynamodb | pydynamodb |
| Amazon Redshift | redshift | sqlalchemy-redshift, psycopg2-binary |
| Apache Doris | doris | pydoris, mysqlclient |
| Apache Drill | drill | sqlalchemy-drill |
| Apache Druid | druid | pydruid |
| Apache Hive | hive | pyhive, thrift, thrift-sasl |
| Apache Impala | impala | impyla |
| Apache Kylin | kylin | kylinpy |
| Apache Pinot | pinot | pinotdb |
| Apache Spark SQL | spark | pyhive, thrift |
| Aurora MySQL | mysql | mysqlclient |
| Aurora MySQL (Data API) | aurora-data-api | preset-sqlalchemy-aurora-data-api |
| Aurora PostgreSQL | postgres | psycopg2-binary |
| Aurora PostgreSQL (Data API) | aurora-data-api | preset-sqlalchemy-aurora-data-api |
| ClickHouse | clickhouse | clickhouse-connect |
| CockroachDB | cockroachdb | cockroachdb, psycopg2-binary |
| CrateDB | crate | sqlalchemy-cratedb |
| Databricks | databricks | databricks-sqlalchemy, databricks-sql-connector |
| Databricks (legacy) | databricks | databricks-sqlalchemy, databricks-sql-connector |
| Databricks Interactive Cluster | databricks | databricks-sqlalchemy, databricks-sql-connector |
| Databricks SQL Endpoint | databricks | databricks-sqlalchemy, databricks-sql-connector |
| Dremio | dremio | sqlalchemy-dremio |
| DuckDB | motherduck | duckdb-engine, duckdb |
| Elasticsearch | elasticsearch | elasticsearch-dbapi |
| Firebird | firebird | sqlalchemy-firebird, fdb |
| Firebolt | firebolt | firebolt-sqlalchemy |
| Google BigQuery | bigquery | sqlalchemy-bigquery, google-cloud-bigquery, pandas-gbq |
| Google Sheets | gsheets | shillelagh |
| MotherDuck | motherduck | duckdb-engine, duckdb |
| MySQL | mysql | mysqlclient |
| OpenSearch (OpenDistro) | elasticsearch | elasticsearch-dbapi |
| PostgreSQL | postgres | psycopg2-binary |
| Presto | presto | pyhive |
| SQLite | Included with Superset | Python's sqlite3 module |
| Shillelagh | Included with Superset | shillelagh |
| SingleStore | singlestore | sqlalchemy-singlestoredb, singlestoredb |
| Snowflake | snowflake | snowflake-sqlalchemy |
| StarRocks | starrocks | starrocks, pymysql |
| Trino | trino | trino |
| Vertica | vertica | sqlalchemy-vertica-python, vertica-python |
The image also installs this driver, which Superset doesn't list in the dialog:
| Database | Extra | Driver packages |
|---|---|---|
| Microsoft SQL Server | mssql | pymssql |
Some notes on the list:
- SQLite and Shillelagh appear in the list, but Superset blocks
connections to them by default. Its
PREVENT_UNSAFE_DB_CONNECTIONSsetting is on, because both can read files on the Superset server. - DuckDB comes with the
motherduckextra, which installs Superset'sduckdbextra. - Databricks: Superset lists four Databricks entries because they share one
SQLAlchemy backend. Use Databricks, which connects with a
databricks://URI. The other three default todatabricks+connector://,databricks+pyhive://, anddatabricks+pyodbc://URIs, and the image doesn't include those drivers.
Connect to Microsoft SQL Server
The image installs pymssql, but Superset lists Microsoft SQL Server in
the Connect a database dialog only when pyodbc is installed. pyodbc
needs the Microsoft ODBC driver, which the image doesn't include. To connect
through pymssql, enter a SQLAlchemy URI:
-
In Superset, click + > Data > Connect database. You can also go to Settings > Database Connections and click + Database.
-
Under Or choose from a list of other databases we support, open the Supported databases list and select Other.
-
In Display Name, enter a name for the connection.
-
In SQLAlchemy URI, enter a
mssql+pymssql://URI:mssql+pymssql://<username>:<password>@<host>:<port>/<database> -
Click Test connection, and then click Connect.
For the URI's options, see the SQLAlchemy SQL Server dialect documentation.
Databases that aren't included
The image doesn't include drivers for these databases:
| Database | Superset extra | Why it's left out |
|---|---|---|
| IBM Db2 | db2 | The ibm-db driver bundles IBM's CLI driver. Its license, the IBM International Program License Agreement, restricts redistribution and hosting. |
| SAP HANA | hana | The SAP Developer License Agreement bars making the hdbcli driver available to third parties. You can add it yourself. |
| Teradata | teradata | The Teradata License Agreement bars distributing or embedding the teradatasql driver without Teradata's written consent. You can add it yourself. |
| Oracle | oracle | cx-Oracle can't connect without Oracle Instant Client, a licensed system library. |
| Azure Synapse | None | It needs Microsoft's ODBC driver, a licensed system library. |
| Databricks over ODBC | None | databricks+pyodbc:// needs the Databricks (Simba Spark) ODBC driver, a licensed system library. Use the Databricks entry instead. |
| Exasol | exasol | Superset's exasol extra requires sqlalchemy-exasol 2.x, and no 2.x release supports SQLAlchemy 1.4, which Superset 6.1 runs on. |
| Airtable, MongoDB | None | Superset has no extra for them. |
| Rockset | None | Rockset shut down its service in 2024. |
Add the SAP HANA or Teradata driver
You can install either driver from PyPI through your own requirements.txt, as
described in Extend the image. When you do, you download
the driver yourself, and you accept the vendor's license directly: the SAP
Developer License Agreement for hdbcli, or the Teradata License Agreement
for teradatasql and teradatasqlalchemy. Read the license, and decide
whether it allows your use, before you build.
For SAP HANA, add the hdbcli driver and the sqlalchemy-hana dialect:
# requirements.txt
hdbcli
sqlalchemy-hana==0.4.0For Teradata, add the teradatasql driver and the teradatasqlalchemy
dialect:
# requirements.txt
teradatasql
teradatasqlalchemy==20.0.0.2Superset 6.1 runs on SQLAlchemy 1.4, which is why both dialects are pinned:
sqlalchemy-hana0.4.0 is the version that Superset's ownhanaextra pins. Without the pin, the install picks a newer release that upgrades SQLAlchemy to 2.0, and Superset stops working.teradatasqlalchemyreleases after 20.0.0.2 need SQLAlchemy 2.0. On SQLAlchemy 1.4, their dialect fails to load, and Superset leaves Teradata out of its database list without showing an error.
After the build, SAP HANA or Teradata appears in the Connect a database dialog.
Extend the image
To add Python packages or Debian packages, build your own image on top of this
one. Put a Dockerfile in a directory, and add a requirements.txt file, a
packages.txt file, or both, next to it:
.
├── Dockerfile
├── packages.txt
└── requirements.txtThe Dockerfile needs one line:
FROM ghcr.io/skaledata/superset:6.1.0packages.txt lists Debian packages, one per line:
# Debian packages, one per line. apt-get installs them as root.
postgresql-client # psql, for testing connections from a podrequirements.txt lists Python packages, in the usual
requirements file format:
# Python packages, one requirement per line.
sqlalchemy-kusto==3.1.1 # Azure Data Explorer (Kusto)Build it for linux/amd64:
docker build --platform linux/amd64 -t my-superset:6.1.0-build1 .The base image declares ONBUILD triggers. They run at the start of your
build, right after your FROM line and before your own instructions:
- If
packages.txtexists,apt-getinstalls each package in it, asroot. - If
requirements.txtexists,uvinstalls it into Superset's virtual environment,/app/.venv, asroot. - The build switches back to the
supersetuser. The image runs assuperset.
Both files are optional. Blank lines and # comments, whether they fill a
whole line or follow a package name, are ignored. A file that contains only
comments installs nothing.
The image has no compiler. If a package you add builds from source, add
build-essential, and any -dev headers the package needs, to
packages.txt.
requirements.txt installs without constraints
The ONBUILD step installs requirements.txt without version constraints. A
package you add can upgrade or downgrade a package that Superset depends on,
such as SQLAlchemy, Flask-AppBuilder, or even apache-superset itself. After
a change like that, Superset can fail to start. Check the build output for
changed or removed packages before you deploy.
Keep Superset's dependencies at their versions
To make the install fail instead of changing a package that the image already
has, install your packages with the image's package list as a constraint.
/app/skaledata/image-freeze.txt lists every package in the image and its
version, except apache-superset and superset-core. Those two install from
/app, not from PyPI, so the constraint doesn't pin Superset itself. Check the
build output for a change to apache-superset.
FROM ghcr.io/skaledata/superset:6.1.0
USER root
ENV HOME=/root
COPY my-requirements.txt /tmp/my-requirements.txt
RUN uv pip install --python /app/.venv/bin/python --no-cache \
--constraint /app/skaledata/image-freeze.txt \
--requirement /tmp/my-requirements.txt
ENV HOME=/app/superset_home
USER superset- Name the file anything but
requirements.txt, such asmy-requirements.txt. TheONBUILDstep installs a file namedrequirements.txtwithout constraints before yourRUNstep runs, so a constraint in yourRUNstep comes too late for it. - The install runs as
root, because Superset's virtual environment belongs toroot.USER supersetswitches back afterward. - The
HOMElines keep files thatrootcreates during the build out of thesupersetuser's home directory.
If a package needs a different version of a package in the image, the build fails with an error like this one:
× No solution found when resolving dependencies:
╰─▶ Because you require humanize==4.13.0 and humanize==4.12.3, we can
conclude that your requirements are unsatisfiable.Deploy a custom image
The deploy script, deploy.sh, builds your image, pushes it to your
cluster's container registry, and deploys it to one Superset app. You don't
need cloud credentials; the script gets short-lived registry credentials from
the SkaleData API.
Before you begin
- Docker,
jq, andcurl7.76 or later. - A SkaleData API key with the
apps:deployscope, or thefullscope. See API key scopes. If you call these endpoints with a console session instead of a key, you need the Operator role or higher; see Roles and permissions. - The cluster's ID. It's the last part of the cluster's page URL in the
console,
/dashboard/clusters/<cluster-id>. - The Superset app's name. You can leave it out if the cluster has only one Superset app.
Get the script and run it
-
Download the script for your Superset app:
export SKALEDATA_API_KEY=sdk_... curl -sS --fail-with-body \ -H "Authorization: Bearer $SKALEDATA_API_KEY" \ -o deploy.sh \ "https://api.skaledata.com/clusters/<cluster-id>/deploy-script?app_type=superset&app_name=<app-name>" chmod +x deploy.sh -
Put
deploy.shin the directory with yourDockerfile. The script's header has an exampleDockerfilethat startsFROMthe exact image tag your cluster runs. -
Run the script from that directory:
./deploy.sh
The script does the following:
- Gets registry credentials from
POST /clusters/<cluster-id>/registry-tokenand signs in to the registry withdocker login. - Builds the image for
linux/amd64, even on a computer with Apple silicon. - Pushes it as
<registry>/<app-name>:<tag>. By default, the tag is the Superset version and a UTC timestamp, such as6.1.0-20260928171233, so each run pushes a new tag. To choose the tag, setIMAGE_TAG, for exampleIMAGE_TAG=6.1.0-build42 ./deploy.sh. - Reads the digest the registry reports for the push. If it can't read a
sha256:digest, it stops before it deploys anything. - Calls
POST /clusters/<cluster-id>/deploy-imagewithapp_type,app_name,image_tag, andimage_digest.
The API responds like this:
{
"status": "deploying",
"image": "<registry>/<app-name>:6.1.0-20260928171233@sha256:<digest>",
"app": "<app-name>"
}The deploy pins the pushed digest. If you push the same tag again with new content and deploy it, the pods still move to the new content.
Tag rules
The API checks the tag before it deploys, and returns 400 Bad Request when a
rule fails:
-
No
latestorcurrent. Each deploy's tag names the build it runs, so you can trace what an app runs back to a build. Use a unique tag for each deploy, such as a git SHA or a timestamp. -
The Superset version must match. The API reads a tag that starts with two dot-separated numbers, with or without a leading
v, as a Superset version. For example,6.1.0-build42,v6.1.0, and6.1all read as Superset 6.1. That version must match your cluster's Superset major and minor version,6.1, because an image on another Superset version migrates the metadata database. A date such as2026.09.28also reads as a version, 2026.9, so it returns400. For example,6.2.0-build1fails with:The tag '6.2.0-build1' looks like Superset 6.2, but this cluster runs Superset 6.1. A different version migrates the metadata database, so build FROM the base image for Superset 6.1. -
A tag without a version is allowed, with a warning. A tag such as a git SHA or
20260928-build1deploys, and the response includes awarningfield that reminds you to buildFROMthe base image for your cluster's Superset version. -
The tag must be a valid Docker tag: up to 128 letters, digits, underscores, periods, and hyphens, not starting with a period or hyphen.
-
image_digestis required for Superset, assha256:followed by 64 lowercase hexadecimal digits.deploy.shsends it for you.
What happens after a deploy
Every deploy runs a Helm upgrade of the Superset app. It moves the web server, the Celery worker, Celery beat, and the database-migration job to the new image. The upgrade takes several minutes. In a test on 2026-09-28, the Helm job started about 4 to 6 minutes after the request, and then the pods rolled in about 1.5 minutes.
deploy-image returns 409 Conflict, and changes nothing, in these cases:
- The cluster isn't in the
readystate. - The cluster's container registry isn't provisioned.
- Superset isn't enabled on the cluster.
- The cluster has more than one Superset app, and the request has no
app_name. - The app is provisioning, restarting, or being deleted. A deploy sets the app
to provisioning until its Helm job finishes, so running
deploy.shagain before the previous deploy finishes returns409. The script pushes the image before it calls the API, so that push lands in the registry but isn't deployed.
Go back to the default image
To drop your custom image and run the SkaleData default image again, send
reset_to_default. The curl command is also in the header of deploy.sh:
curl -sS --fail-with-body -X POST \
"https://api.skaledata.com/clusters/<cluster-id>/deploy-image" \
-H "Authorization: Bearer $SKALEDATA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"app_type": "superset", "app_name": "<app-name>", "reset_to_default": true}'Don't send image_tag or image_digest with reset_to_default. The reset
also runs a Helm upgrade, and the response's image is the default image.
The console doesn't show a Superset app's custom image. To check which image
an app runs, call GET /applications/<app-id> with a key that has the
apps:read or full scope. <app-id> is the last part of the app's page URL in the
console, /dashboard/applications/<app-id>. In the response:
config.last_deployed_imageis the image that the last deploy or reset recorded, pinned by digest.config.superset_custom_imageholds the custom image. After a reset, that field is gone, and the app runs the default image.
Related pages
- API key scopes
- Roles and permissions
- Airflow image, SkaleData's Airflow base image, which you extend the same way
- Superset database drivers, in the Apache Superset documentation