Skip to main content

Install Datagrok with Helm

The Datagrok Helm chart is published as an OCI artifact on Docker Hub. It supports in-cluster PostgreSQL with local PVCs (laptop / single-node clusters), as well as managed cloud databases (RDS / Cloud SQL) with object storage (S3 / GCS) on EKS or GKE.

On AWS?

For turnkey EKS deployments, use the AWS CloudFormation (EKS) template — it provisions the cluster, RDS, S3, and IAM, and installs this chart automatically. This page covers manual Helm installs for any Kubernetes cluster (on-prem, GKE, AKS, kind, k3s, MicroK8s).

Prerequisites​

  • Kubernetes 1.27+ with a default StorageClass (or an existing PVC for stateful data)
  • Helm 3.8+ (for OCI registry support)
  • For cloud installs: a managed Postgres instance, an object storage bucket, and an IAM role / service account that the cluster can use to reach them

Chart version​

The chart is published as an OCI artifact in the shared datagrok/datagrok Docker Hub repo under tags suffixed with -helm to keep the chart and image tag namespaces disjoint. --version 1.27.3-helm pulls a chart that deploys Datagrok 1.27.3, and every sub-service (grok-pipe, grok-spawner, grok-connect) defaults to the same app tag. Override individual service image tags via --set <service>.image.tag=... — see Service versions below.

Chart tags follow the same scheme as the Datagrok image tags with a -helm suffix; see Images and versions for the full convention.

Quick install (in-cluster Postgres + local storage)​

helm install datagrok oci://registry-1.docker.io/datagrok/datagrok \
--version 1.27.3-helm \
--namespace datagrok --create-namespace \
--set datagrok.adminPassword=<password> \
--set postgres.adminPassword=$(openssl rand -base64 24) \
--set ingress.host=datagrok.example.com

This installs PostgreSQL, RabbitMQ, the datagrok app, grok-pipe, grok-spawner, grok-connect, and grok-connect-adbc with default resources. Suitable for evaluation, dev, or single-tenant production. The chart deploys no script execution gateways. The Scripting package starts a scripting-<lang> container on the first run of a script in that language and stops it when idle.

The chart ships no default passwords. Rendering fails without datagrok.adminPassword (the password of the Datagrok admin user) and, for the in-cluster PostgreSQL, postgres.adminPassword. To keep passwords off the command line, store them in a Kubernetes Secret and reference it with datagrok.existingSecret and postgres.existingSecret, or map single fields with datagrok.secretRefs.* and postgres.secretRefs.*. With ingress.host set, RabbitMQ accepts Datagrok-issued tokens and needs no password. Without an ingress host, also set rabbitmq.password.

To track the latest unstable build (rebuilt after every merge to master), use --version bleeding-edge-helm instead of a release version.

Service versions​

Every service image tag defaults to the chart version. The default set is documented on Images and versions and matches the AWS CloudFormation templates. Override individual tags when you need to run a newer grok_connect against an older datagrok core, or to pin a specific service during a rollout:

helm install datagrok oci://registry-1.docker.io/datagrok/datagrok \
--version 1.27.3-helm \
--set datagrok.image.tag=1.27.3 \
--set grokPipe.image.tag=1.19.0 \
--set grokConnect.image.tag=2.6.2 \
--set spawner.image.tag=2.16.0 \
--set registry.proxy.image.tag=1.27.1 \
-n datagrok

RabbitMQ follows its own upstream release cadence and is not pinned to the Datagrok version.

Database connectors​

The chart runs up to three grok_connect endpoints:

ValueDefaultService
grokConnect.enabledtrueJava JDBC connectors, port 1234.
grokConnectAdbc.enabledtrueRust ADBC connector, port 1235. ClickHouse and BigQuery can use it instead of JDBC. Disable it to keep these types on JDBC only.
grokConnectExtended.enabledfalseAmazon Neptune and Cloudera Impala. Their drivers carry known CVEs with no upstream fix, so they ship in a separate, optional image.

To use Neptune or Impala, accept that CVE exposure and enable the extended connectors with --set grokConnectExtended.enabled=true. While it's off, these data sources don't appear in Datagrok, and existing connections of these types can't run.

Common options​

ValueDescription
datagrok.openId.*OpenID Connect single sign-on: enabled, configEndpoint, clientId, secretType (Client Secret or Signed JWT), autoLogin.
serverKeys.backendWhere server keys are stored: local (default), aws_secrets_manager, or gcp_secret_manager.
serverKeys.signingKeyMaxAgeDaysRotates the primary signing key when it is older than this many days. 0 (default) turns rotation off.
serverKeys.deleteRecoveryDaysAWS Secrets Manager recovery window, in days, for a deleted key (7–30, default 7). 0 deletes it immediately.
datagrok.hsts.enabledAdds the HTTP Strict Transport Security header. Enable it when TLS terminates in front of the pod.
datagrok.updateStrategyRolling update maxUnavailable and maxSurge (25% each). On a small node set, use maxSurge: 0 and maxUnavailable: 1.
datagrok.dnsConfigPod DNS resolver options (default ndots: 2, timeout: 2, attempts: 3). Set to {} to use the cluster defaults.
credentials.externalSecrets.*With credentials.source: externalSecrets, reads all passwords from an External Secrets store under remoteKey.
<service>.revisionHistoryLimitNumber of old ReplicaSets kept for rollback (default 10).
postgres.extraConfigExtra PostgreSQL settings as a key-value map, passed as -c key=value. Overrides the named postgres.* settings.
scripting.*Script execution settings, for example scripting.useKernelPool and scripting.kernelPoolSize. These replace the former jkg.* values.

Production install on AWS EKS​

For new AWS stands use the AWS CloudFormation (EKS) template — it provisions EKS, RDS, S3, IAM with IRSA, and installs this chart for you. The steps below are for installing the chart directly into an EKS cluster you already manage.

  1. Configure kubectl:

    aws eks update-kubeconfig --name <cluster-name>
  2. Install the AWS Load Balancer Controller in the cluster — the chart's ingress.className: alb annotations require it.

  3. Provision RDS, S3, and the IRSA role for the Datagrok ServiceAccount yourself (rds-db:connect, s3:Get/Put/List on the bucket, optionally Secrets Manager read).

  4. Save the EKS overlay below as values-prod.yaml, filling in the SET fields:

    postgres:
    internal: false
    external:
    host: your-db.xxxxx.us-east-1.rds.amazonaws.com
    port: 5432
    ssl: true

    storage:
    type: s3
    s3:
    bucket: your-datagrok-bucket
    region: us-east-1

    ingress:
    enabled: true
    className: alb
    host: datagrok.example.com
    annotations:
    alb.ingress.kubernetes.io/scheme: internet-facing
    alb.ingress.kubernetes.io/target-type: ip
    alb.ingress.kubernetes.io/listen-ports: '[{"HTTPS":443}]'
    alb.ingress.kubernetes.io/certificate-arn: arn:aws:acm:...:certificate/...
    tls:
    enabled: false # ALB terminates TLS

    serviceAccount:
    create: true
    annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::ACCOUNT:role/datagrok-role

    registry:
    type: proxy
    proxy:
    backendUrl: https://ACCOUNT.dkr.ecr.REGION.amazonaws.com

    credentials:
    source: externalSecrets
    externalSecrets:
    enabled: true
    remoteKey: datagrok/prod # AWS Secrets Manager key
  5. Install:

    helm install datagrok oci://registry-1.docker.io/datagrok/datagrok \
    --version 1.27.3-helm \
    -f values-prod.yaml \
    -n datagrok --create-namespace

The chart repo also ships a ready-made values-eks.yaml skeleton you can copy.

Production install on GCP GKE​

Same flow as EKS, but use the values-gke.yaml overlay (Cloud SQL + GCS + GKE Workload Identity).

Upgrades​

helm upgrade datagrok oci://registry-1.docker.io/datagrok/datagrok \
--version 1.27.4-helm \
-f values-prod.yaml \
-n datagrok

Always upgrade through consecutive minor versions for production stands; database schema migrations run automatically on the first start of each new app version.

The chart's PostgreSQL StatefulSet, datagrok-data, and datagrok-cfg PVCs are preserved across upgrades. Database schema migrations run automatically on the first start of a new app version.

Reusing existing PVCs​

If you're migrating an existing Datagrok install to the chart and want to keep your data, point the chart at the existing claims:

postgres:
internal: true
existingClaim: my-existing-postgres-pvc

storage:
type: local
local:
existingDataClaim: my-existing-data-pvc
existingCfgClaim: my-existing-cfg-pvc

When postgres.existingClaim is set, the chart skips its own volumeClaimTemplates and mounts the data volume at the PV root (no pgdata subPath), matching the on-disk layout used by Datagrok installs prior to the chart.

Uninstall​

helm uninstall datagrok -n datagrok

PVCs are NOT deleted by helm uninstall. To remove all data permanently:

kubectl delete pvc -n datagrok -l app.kubernetes.io/instance=datagrok
kubectl delete namespace datagrok