prole/etc/init_kong.sh
chrisfu 4a8d9cc90d feat: full GKE/prod deployment pipeline from UI to Artifact Registry
## GCP / Cluster Environment Screen
- Auto-populate Cloud tab from conf/prod/gcp.cfg on screen open (org_id,
  billing_account, billing_project, project_id)
- gcloud auth validity checked on screen startup; friendly modal dialog
  streams gcloud auth login output live so user never leaves the app
- Live GKE cluster browser: fetches clusters via gcloud container clusters
  list, displays with checkmark selector, auto-selects saved cluster
- Selecting a cluster runs get-credentials, sets KUBECONFIG/KUBECONTEXT,
  and syncs the region dropdown to the selected cluster's location
- Region dropdown populated live from gcloud compute regions list with
  checkmark on currently selected region; graceful fallback when offline
- New 'GCP Storage' tab with workload->StorageClass mapping (CNPG->premium-rwo,
  Redis/Monitoring->standard-rwo, Garage->garage-hdd) and Fetch from Cluster
- Provider readonly field styled correctly (no solid-black on macOS)
- Stale prole.cfg/conf/prole.cfg symlinks removed; all config I/O now
  resolves env-specific paths via prole_conf.entrypoint_path()

## GKE Autopilot Compatibility (Common Services)
- Synology iSCSI StorageClass and static PVs guarded behind PROLE_MODE!=k8s
  in init_openbao.sh (GKE Autopilot forbids hostPath/iSCSI volumes)
- In-cluster Docker registry (hostPath) skipped in k8s mode; GCP Artifact
  Registry used instead
- Kong renamed knoe-svc-kong in k8s mode; all health-check kubectl calls in
  init_common_services.sh and status_common_services.sh updated accordingly
- DNS endpoints switched from *.prole.org to *.knoe.dev in k8s mode
  (api.knoe.dev, git.knoe.dev, svc.knoe.dev); ingress uses gce class
- New GKE-clean Kong manifests under deploy/opentofu/k8s/manifests/prole/:
  no k3s node affinity, explicit Autopilot resource requests/limits

## Garage S3 Store (GKE)
- New garage-statefulset-gcp.yaml targeting garage-hdd StorageClass
  (pd-standard, avoids SSD_TOTAL_GB quota exhaustion in us-west3)
- New storageclass-gcp-hdd.yaml (pd-standard, Retain, WaitForFirstConsumer)
- GCP StorageClass manifests skipped on re-runs (Autopilot built-ins are
  immutable; skip-if-exists guard added)
- PVC deletion guard extended to cover any storageClass (not just synology)
  so stale claims are cleaned before StatefulSet recreation

## Topology (GKE Autopilot)
- DaemonSet collector skipped in prod mode (forbidden in kube-system by
  GKE Warden); Kubernetes-only node facts path used instead
- All ready GKE nodes assumed cnpg-eligible and monitoring-eligible without
  taint/synology-mount checks (skip_collector + assume_nodes_eligible flags)

## KUBECONFIG / kubectl (k8s mode)
- actions.py: new elif mode==k8s branch sets KUBECONFIG=~/.kube/config
  and injects KUBECONTEXT from prole_cfg_data into script env
- _build_kubectl_cmd falls back to Global.KUBECONTEXT when
  init_cluster.selected_kubectx is empty
- _activate_selected_gke_cluster persists KUBECONFIG/KUBECONTEXT to
  prole_cfg_data and saves prole.cfg immediately after get-credentials

## Database Build Screen (GKE)
- Registry display shows correct Artifact Registry URL
  (<region>-docker.pkg.dev/<project>/<namespace>/knoe-db) in green
- Build+push: gcloud auth configure-docker, auto-creates AR repository
  named after SERVICE_NAMESPACE (e.g. knoe-system) if missing, then
  docker tag + push; falls back to gcr.io if region unavailable
- GCP config loaded from conf/prod/gcp.cfg on every screen entry;
  keys normalised to lowercase so project_id lookup is always consistent

## Config / Namespace persistence
- prole_conf.py activate_environment: symlink creation removed; sets
  CLUSTER_ENV env-var so all subsequent calls resolve correct env directory
- knoe/ui/screens/__init__.py: startup config load uses entrypoint_path()
  instead of hardcoded conf/prole.cfg; seeds SERVICE_NAMESPACE=knoe-system
  for managed envs so Common Services never defaults to 'default'
- cfg.py _save_prole_cfg: saves to env-specific path via entrypoint_path()
- etc/prole_cfg.sh: removed all ln -snf symlink creation

Co-authored-by: Junie <junie@jetbrains.com>
2026-04-04 12:38:16 -07:00

461 lines
14 KiB
Bash
Executable File

#!/usr/bin/env bash
set -euo pipefail
# init_kong.sh
# Purpose:
# - Deploy Kong API Gateway (DB-less) into the service namespace
# - Replaces the prole nginx deployment as the API endpoint
# - Routes /backup/* to knoe-db-manager
# - Creates the kong declarative config as a ConfigMap
# - Applies the kong deployment and service manifests
# - Provides start/stop/status/restart actions
SCRIPT_DIR=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
# Shared option parsing for common core scripts
# shellcheck disable=SC1090
source "$SCRIPT_DIR/common_core_lib.sh"
# Inject default config if not provided
_has_config=0
for _arg in "$@"; do
[[ "$_arg" == "-c" || "$_arg" == "--config" || "$_arg" == -c=* || "$_arg" == --config=* ]] && _has_config=1
done
if [[ $_has_config -eq 0 && -f "$SCRIPT_DIR/../conf/prole.cfg" ]]; then
set -- "-c" "$SCRIPT_DIR/../conf/prole.cfg" "$@"
fi
unset _has_config _arg
common_core_preparse_config "$@"
# shellcheck disable=SC1090
source "$SCRIPT_DIR/prole_cfg.sh"
set -- "${COMMON_CORE_ARGS[@]}"
common_core_parse_args "$@"
if [[ -z "${PROLE_MODE:-}" ]]; then
export PROLE_MODE="k3s"
fi
if [[ "${COMMON_CORE_HELP:-0}" == 1 ]]; then
common_core_usage "$0"
exit 0
fi
if [[ -n "${COMMON_CORE_PARSE_ERROR:-}" ]]; then
echo "ERROR: ${COMMON_CORE_PARSE_ERROR}" >&2
common_core_usage "$0"
exit 2
fi
ACTION="$COMMON_CORE_ACTION"
NAMESPACE="$(common_core_resolve_namespace "default")"
common_core_apply_namespace "$NAMESPACE"
PROLE_HOME=${PROLE_HOME:-$(cd "$SCRIPT_DIR/.." && pwd)}
KONG_IMAGE="${KONG_IMAGE:-kong:3.9}"
# k8s/GKE prod mode uses knoe.dev domain and knoe-svc-kong; all other modes use prole.org
if [[ "${PROLE_MODE:-}" == "k8s" ]]; then
KONG_NAME="${KONG_NAME:-knoe-svc-kong}"
KONG_CONFIG_NAME="${KONG_CONFIG_NAME:-knoe-svc-kong-config}"
SERVICE_HOSTNAME="${SERVICE_HOSTNAME:-svc.knoe.dev}"
AUTH_HOSTNAME="${AUTH_HOSTNAME:-api.knoe.dev}"
GITEA_HOSTNAME="${GITEA_HOSTNAME:-${GITEA_DOMAIN:-git.knoe.dev}}"
else
KONG_NAME="${KONG_NAME:-prole-svc-kong}"
KONG_CONFIG_NAME="${KONG_CONFIG_NAME:-prole-svc-kong-config}"
SERVICE_HOSTNAME="${SERVICE_HOSTNAME:-svc.prole.org}"
AUTH_HOSTNAME="${AUTH_HOSTNAME:-api.prole.org}"
GITEA_HOSTNAME="${GITEA_HOSTNAME:-${GITEA_DOMAIN:-git.prole.org}}"
fi
KONG_PROXY_PORT="${KONG_PROXY_PORT:-8000}"
KONG_ADMIN_PORT="${KONG_ADMIN_PORT:-8001}"
KONG_GITEA_SSH_PORT="${KONG_GITEA_SSH_PORT:-3022}"
SERVICE_TLS_SECRET_NAME="${SERVICE_TLS_SECRET_NAME:-${SERVICE_HOSTNAME//./-}-tls}"
SERVICE_TLS_CLUSTER_ISSUER="${SERVICE_TLS_CLUSTER_ISSUER:-letsencrypt-prod}"
# Legacy: svc-check used to own svc.prole.org. We now route the service hostname
# to Grafana, so remove any leftover svc-check resources to avoid conflicts.
SVC_CHECK_NAMESPACE="${SVC_CHECK_NAMESPACE:-svc-check}"
# kubectl robustness knobs (timeouts/retries for transient apiserver slowness)
KUBECTL_REQUEST_TIMEOUT="${KUBECTL_REQUEST_TIMEOUT:-30s}"
KUBECTL_APPLY_RETRIES="${KUBECTL_APPLY_RETRIES:-5}"
KUBECTL_APPLY_RETRY_DELAY="${KUBECTL_APPLY_RETRY_DELAY:-2}"
# Upstream service defaults
DB_MANAGER_SERVICE="${DB_MANAGER_SERVICE:-knoe-db-manager}"
DB_MANAGER_PORT="${DB_MANAGER_PORT:-80}"
DB_MANAGER_NAMESPACE="${DB_MANAGER_NAMESPACE:-${DATABASE_NAMESPACE:-knoe-db}}"
PROLE_SERVICE_UPSTREAM_URL="${PROLE_SERVICE_UPSTREAM_URL:-http://prole-svc.${NAMESPACE}.svc.cluster.local:8080}"
GRAFANA_UPSTREAM_URL="${GRAFANA_UPSTREAM_URL:-http://kps-grafana.monitoring.svc.cluster.local:80}"
GITEA_HTTP_UPSTREAM_URL="${GITEA_HTTP_UPSTREAM_URL:-http://gitea-http.gitea.svc.cluster.local:3000}"
GITEA_SSH_UPSTREAM_HOST="${GITEA_SSH_UPSTREAM_HOST:-gitea-ssh.gitea.svc.cluster.local}"
GITEA_SSH_UPSTREAM_PORT="${GITEA_SSH_UPSTREAM_PORT:-22}"
# SSO wiring knobs
PROLE_GRAFANA_SSO_ENABLED="${PROLE_GRAFANA_SSO_ENABLED:-0}"
GRAFANA_PROXY_UPSTREAM_URL="${GRAFANA_PROXY_UPSTREAM_URL:-http://prole-grafana-proxy.${NAMESPACE}.svc.cluster.local:80}"
KNOE_AUTH_UPSTREAM_URL="${KNOE_AUTH_UPSTREAM_URL:-http://knoe-auth.${SERVICE_NAMESPACE:-${NAMESPACE}}.svc.cluster.local:8080}"
usage() {
cat <<USAGE
Usage: $0 [--mode MODE] [-n NAMESPACE] [${COMMON_CORE_ACTIONS//|/|}] [-c conf/prole.cfg]
Actions:
start Create ConfigMap and deploy Kong
stop Remove Kong deployment, service, and ConfigMap
status Show Kong pod/service status
restart Restart Kong pods
update Create or update Kong resources
USAGE
exit 1
}
ensure_tools() {
for t in kubectl; do
command -v "$t" >/dev/null || { echo "Missing required tool: $t" >&2; exit 1; }
done
}
ensure_namespace() {
if ! kubectl get namespace "$NAMESPACE" >/dev/null 2>&1; then
echo "Creating namespace '$NAMESPACE' ..."
kubectl create namespace "$NAMESPACE" >/dev/null 2>&1 || true
fi
}
kubectl_rt() {
kubectl --request-timeout="$KUBECTL_REQUEST_TIMEOUT" "$@"
}
kubectl_apply_retry() {
local attempt=1
local delay="$KUBECTL_APPLY_RETRY_DELAY"
while true; do
if kubectl_rt apply "$@"; then
return 0
fi
local rc=$?
if [[ "$attempt" -ge "$KUBECTL_APPLY_RETRIES" ]]; then
return "$rc"
fi
echo "WARN: kubectl apply failed (attempt ${attempt}/${KUBECTL_APPLY_RETRIES}); retrying in ${delay}s ..." >&2
sleep "$delay"
attempt=$((attempt + 1))
delay=$((delay * 2))
done
}
# Sets KONG_CONFIG_CHANGED=1 in the caller's scope when the ConfigMap was
# created or updated; leaves it 0 when kubectl reported "unchanged".
# A temp file is used to communicate the result out of the subshell.
KONG_CONFIG_CHANGED=0
_KONG_CONFIG_CHANGED_FILE=""
create_kong_config() {
echo "Creating/updating Kong declarative config '$KONG_CONFIG_NAME' in namespace '$NAMESPACE' ..."
local grafana_url
grafana_url="$GRAFANA_UPSTREAM_URL"
case "${PROLE_GRAFANA_SSO_ENABLED:-0}" in
1|true|TRUE|True|yes|YES|on|ON)
grafana_url="$GRAFANA_PROXY_UPSTREAM_URL"
;;
esac
local kong_yml
kong_yml=$(cat <<KONGEOF
_format_version: "3.0"
_transform: true
services:
- name: prole-service
url: ${PROLE_SERVICE_UPSTREAM_URL}
routes:
- name: prole-k3s-kubeconfig
hosts:
- ${SERVICE_HOSTNAME}
paths:
- /k3s/kube_config.sh
strip_path: false
- name: db-manager
url: http://${DB_MANAGER_SERVICE}.${DB_MANAGER_NAMESPACE}.svc.cluster.local:${DB_MANAGER_PORT}
routes:
- name: backup-route
hosts:
- ${SERVICE_HOSTNAME}
paths:
- /backup
strip_path: false
- name: grafana
url: ${grafana_url}
routes:
- name: grafana-root
hosts:
- ${SERVICE_HOSTNAME}
paths:
- /
strip_path: false
- name: knoe-auth
url: ${KNOE_AUTH_UPSTREAM_URL}
routes:
- name: knoe-auth-root
hosts:
- ${AUTH_HOSTNAME}
paths:
- /
strip_path: false
- name: gitea-http
url: ${GITEA_HTTP_UPSTREAM_URL}
routes:
- name: gitea-root
hosts:
- ${GITEA_HOSTNAME}
paths:
- /
strip_path: false
- name: gitea-ssh
host: ${GITEA_SSH_UPSTREAM_HOST}
port: ${GITEA_SSH_UPSTREAM_PORT}
protocol: tcp
routes:
- name: gitea-ssh-tcp
protocols:
- tcp
destinations:
- port: ${KONG_GITEA_SSH_PORT}
KONGEOF
)
# Use a subshell to ensure the cleanup trap doesn't leak globally (and trip `set -u` later).
# A sentinel file is used to communicate a config change back to the parent shell.
_KONG_CONFIG_CHANGED_FILE="$(mktemp)"
(
tmp="$(mktemp)"
trap 'rm -f "${tmp:-}"' EXIT
kubectl create configmap "$KONG_CONFIG_NAME" \
--namespace="$NAMESPACE" \
--from-literal=kong.yml="$kong_yml" \
--dry-run=client -o yaml >"$tmp"
local apply_out
apply_out=$(kubectl_apply_retry -f "$tmp" 2>&1)
echo "$apply_out"
if ! echo "$apply_out" | grep -q 'unchanged'; then
echo "1" >"$_KONG_CONFIG_CHANGED_FILE"
fi
)
if [[ -f "$_KONG_CONFIG_CHANGED_FILE" ]] && [[ "$(cat "$_KONG_CONFIG_CHANGED_FILE")" == "1" ]]; then
KONG_CONFIG_CHANGED=1
fi
rm -f "$_KONG_CONFIG_CHANGED_FILE"
_KONG_CONFIG_CHANGED_FILE=""
echo "ConfigMap '$KONG_CONFIG_NAME' ready."
}
cleanup_legacy_svc_check() {
# Best-effort cleanup: older installs applied a static check page (svc-check)
# that owned the service hostname via its own Ingress and injected routes into
# the shared Kong declarative config ConfigMap.
kubectl -n "$NAMESPACE" delete ingress svc-check-ingress --ignore-not-found >/dev/null 2>&1 || true
kubectl delete namespace "$SVC_CHECK_NAMESPACE" --ignore-not-found >/dev/null 2>&1 || true
}
apply_service_ingress() {
local host="${SERVICE_HOSTNAME:-}"
if [[ -z "$host" ]]; then
echo "WARN: SERVICE_HOSTNAME is empty; skipping service Ingress." >&2
return 0
fi
local auth_host="${AUTH_HOSTNAME:-}"
local gitea_host="${GITEA_HOSTNAME:-}"
local tls_hosts_extra=""
local rules_extra=""
if [[ -n "$auth_host" && "$auth_host" != "$host" ]]; then
tls_hosts_extra=$'\n - '"${auth_host}"
rules_extra=$(cat <<EOF
- host: ${auth_host}
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: ${KONG_NAME}
port:
number: ${KONG_PROXY_PORT}
EOF
)
fi
if [[ -n "$gitea_host" && "$gitea_host" != "$host" && "$gitea_host" != "$auth_host" ]]; then
tls_hosts_extra+=$'\n - '"${gitea_host}"
rules_extra+=$'\n'$(cat <<EOF
- host: ${gitea_host}
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: ${KONG_NAME}
port:
number: ${KONG_PROXY_PORT}
EOF
)
fi
local ingress_name="svc-knoe-ingress"
local extra_annotations=""
if [[ "${PROLE_MODE:-}" != "k8s" ]]; then
extra_annotations=$' traefik.ingress.kubernetes.io/router.priority: "10"\n'
fi
local ingress_class
if [[ "${PROLE_MODE:-}" == "k8s" ]]; then
ingress_class="gce"
else
ingress_class="traefik"
fi
echo "Applying service Ingress for host '${host}' -> ${KONG_NAME}:${KONG_PROXY_PORT} (namespace=${NAMESPACE}) ..."
(
tmp="$(mktemp)"
trap 'rm -f "${tmp:-}"' EXIT
cat >"$tmp" <<EOF
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: ${ingress_name}
namespace: ${NAMESPACE}
annotations:
kubernetes.io/ingress.class: ${ingress_class}
${extra_annotations} cert-manager.io/cluster-issuer: ${SERVICE_TLS_CLUSTER_ISSUER}
spec:
tls:
- hosts:
- ${host}
${tls_hosts_extra}
secretName: ${SERVICE_TLS_SECRET_NAME}
rules:
- host: ${host}
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: ${KONG_NAME}
port:
number: ${KONG_PROXY_PORT}
${rules_extra}
EOF
kubectl_apply_retry -f "$tmp"
)
}
deploy() {
echo "Deploying $KONG_NAME to namespace '$NAMESPACE' ..."
local manifests_dir
if [[ "${PROLE_MODE:-}" == "k8s" ]]; then
manifests_dir="$PROLE_HOME/deploy/opentofu/k8s/manifests/prole"
else
manifests_dir="$PROLE_HOME/deploy/opentofu/k3s/manifests/prole"
fi
# Determine whether the Deployment already exists before applying manifests.
local deployment_existed=0
kubectl_rt get deployment "$KONG_NAME" -n "$NAMESPACE" >/dev/null 2>&1 && deployment_existed=1
local deploy_out
deploy_out=$(kubectl_apply_retry -f "$manifests_dir/kong-deployment.yaml" -n "$NAMESPACE" 2>&1)
echo "$deploy_out"
local svc_out
svc_out=$(kubectl_apply_retry -f "$manifests_dir/kong-service.yaml" -n "$NAMESPACE" 2>&1)
echo "$svc_out"
echo "Waiting for $KONG_NAME rollout ..."
kubectl rollout status deployment/"$KONG_NAME" -n "$NAMESPACE" --timeout=120s
# ConfigMaps do not trigger a Deployment rollout by default. Only restart
# Kong when something actually changed: either this is a fresh deployment or
# the declarative config ConfigMap was modified. Skipping the restart when
# nothing changed prevents a new ReplicaSet from being created every run.
local need_restart=0
[[ "$deployment_existed" -eq 0 ]] && need_restart=1
[[ "${KONG_CONFIG_CHANGED:-0}" -eq 1 ]] && need_restart=1
if [[ "$need_restart" -eq 1 ]]; then
echo "Restarting $KONG_NAME to reload declarative config ..."
kubectl rollout restart deployment/"$KONG_NAME" -n "$NAMESPACE" >/dev/null 2>&1 || true
kubectl rollout status deployment/"$KONG_NAME" -n "$NAMESPACE" --timeout=120s >/dev/null 2>&1 || true
else
echo "$KONG_NAME config unchanged; skipping rollout restart."
fi
echo "$KONG_NAME deployed successfully."
}
stop() {
echo "Removing $KONG_NAME from namespace '$NAMESPACE' ..."
kubectl delete deployment "$KONG_NAME" -n "$NAMESPACE" --ignore-not-found=true
kubectl delete service "$KONG_NAME" -n "$NAMESPACE" --ignore-not-found=true
kubectl delete configmap "$KONG_CONFIG_NAME" -n "$NAMESPACE" --ignore-not-found=true
echo "$KONG_NAME removed."
}
status() {
echo "=== $KONG_NAME pods ==="
kubectl get pods -n "$NAMESPACE" -l app="$KONG_NAME" 2>/dev/null || echo "No pods found"
echo ""
echo "=== $KONG_NAME service ==="
kubectl get svc "$KONG_NAME" -n "$NAMESPACE" 2>/dev/null || echo "No service found"
}
restart() {
echo "Restarting $KONG_NAME ..."
kubectl rollout restart deployment/"$KONG_NAME" -n "$NAMESPACE"
kubectl rollout status deployment/"$KONG_NAME" -n "$NAMESPACE" --timeout=120s
echo "$KONG_NAME restarted."
}
action_update() {
ensure_tools
ensure_namespace
cleanup_legacy_svc_check
create_kong_config
apply_service_ingress
deploy
}
case "$ACTION" in
start|initialize|update|reload)
action_update
;;
stop)
ensure_tools
stop
;;
status)
ensure_tools
status
;;
restart)
ensure_tools
restart
;;
*)
usage
;;
esac