# CLAUDE.md — knoe project context > Loaded automatically by Claude Code (CLI and IntelliJ plugin) as project context. --- ## Repo role: knoe-db platform repo This is `knoe-db` (remote: `git@git-ssh.knoe.dev:knoe.dev/knoe-db.git`, configured as both `origin` and `knoe`). It's the platform's source of truth — `authority/`, `knoe/`, `etc/init_*.sh`, `deploy/gcp/gke/*`, the test pipeline, all live here. Customer deploys are intended to live as **branches** in this repo (e.g. a future `customer/prole.org`), not as separate forks. As of this writing, no customer branch is active — the prole→knoe rebrand has merged into `main` and there is no separate `~/dev/prole` working tree under development. See [`docs/plans/customer-deploy-resync.md`](docs/plans/customer-deploy-resync.md) for the original plan, currently dormant. ## Master TODO Single source of truth for unfinished work, including the reality-vs-intent gaps flagged below: **[`docs/TODO.md`](docs/TODO.md)**. ## Active plans - [`docs/plans/README.md`](docs/plans/README.md) — Index and conventions for this directory. - [`docs/plans/knoe-auth-round-1.md`](docs/plans/knoe-auth-round-1.md) — Identity backbone (Kerberos KDC + invite-OTP + TOTP). **Shipped.** - [`docs/plans/deployment-modes.md`](docs/plans/deployment-modes.md) — Four installer modes (`min` / `k3d` / `k3s` / `gke`). **Shipped (Phase 0).** - [`docs/pipeline-phases.md`](docs/pipeline-phases.md) — Autobuild & test pipeline phase reference. Phase 0 ✅, Phase 1 next (first task: f-string fix at `knoe/core/ops/cloudnative_pg.py:1372`). - [`docs/plans/customer-deploy-resync.md`](docs/plans/customer-deploy-resync.md) — The original plan to converge `~/dev/prole` onto `knoe-db/main` as a customer-deploy branch. **Dormant** (rebrand is now in main, no separate prole tree active). --- ## Dual-cluster GKE architecture This project uses **two separate GKE Standard clusters** in `us-west3`, both currently provisioned with `e2-standard-2` × 3 nodes (2 vCPU / 8 GB each, ~7.1 GB allocatable). Verify with `gcloud container clusters list`. | Cluster | Context | Role | |---|---|---| | `knoe-dev-0` | `gke_plenary-truck-485623-p7_us-west3_knoe-dev-0` | App cluster — Garage, registry, OpenBao, Kong, GitLab, monitoring | | `knoe-dev-cnpg-0` | `gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0` | DB cluster — CNPG/PostgreSQL only | **Storage quota:** the project has `SSD_TOTAL_GB = 300 GB` in `us-west3`, **fully consumed by CNPG**. All non-CNPG PVCs must use `standard` (pd-standard / HDD) — not `standard-rwo` / `premium-rwo`, which are SSD-backed and will fail to provision with a quota error. `GITLAB_GITALY_STORAGE_CLASS = standard` is set in `conf/gke.cfg` accordingly. ### Resource allocation | Resource | Cluster | Namespace | |---|---|---| | CNPG operator (v1.29.0) | `knoe-dev-cnpg-0` | `cnpg-system` | | PostgreSQL cluster (`knoe-db`) | `knoe-dev-cnpg-0` | `knoe-db-0` | | Barman Cloud plugin (v0.12.0) | `knoe-dev-cnpg-0` | `cnpg-system` | | cert-manager | `knoe-dev-cnpg-0` | `cert-manager` | | Garage (S3 object store) | `knoe-dev-0` | `knoe-system` | | Registry | `knoe-dev-0` | `knoe-system` | | OpenBao | `knoe-dev-0` | `knoe-system` | | Kong API gateway | `knoe-dev-0` | `knoe-system` | | Monitoring | `knoe-dev-0` | `monitoring` | **Garage runs ONLY in `knoe-dev-0`.** The DB cluster (`knoe-dev-cnpg-0`) has none — was removed 2026-04-29. Do NOT redeploy Garage to the DB cluster (use GCS for backups there). ### CNPG backups → GCS Backups use **GCS with Workload Identity**: - Data + WAL: `gs://knoe-0-backups/` (single bucket; `knoe-db/base/` and `knoe-db/wals/` prefixes). `gs://knoe-0-wal/` exists but is unused. - GCP SA: `cnpg-backup@plenary-truck-485623-p7.iam.gserviceaccount.com` (roles on bucket: `storage.objectAdmin`, `storage.legacyBucketReader`) - K8s SA: cluster pods run as **`cnpg-backup-sa`** in `knoe-db-0` (set via `cluster.spec.serviceAccountName`, requires CNPG ≥ v1.29.0). The SA is annotated with `iam.gke.io/gcp-service-account=cnpg-backup@…`. Two `RoleBinding` subjects (`knoe-db` and `knoe-db-barman-cloud`) include `cnpg-backup-sa` so the pod has the same RBAC the auto-generated SA would have had. - ObjectStore manifest: [`k8s/knoe/knoe-db-barman-objectstore-gcs.yaml`](k8s/knoe/knoe-db-barman-objectstore-gcs.yaml) — includes `googleCredentials.gkeEnvironment: true` (required by plugin-barman-cloud v0.12.0). Setup script: [`etc/init_cnpg_gke.sh`](etc/init_cnpg_gke.sh) (creates buckets, GCP SA, WI binding, applies CNPG cluster). --- ## Service mesh (Cloud Service Mesh / Istio) Both clusters are registered in the **knoe-0** GCP fleet with automatic Cloud Service Mesh management. This is automated in `scripts/reset_clusters.sh` (Phase 7) — no longer requires GCP web console. ```bash # Check mesh provisioning status (~10 min after cluster creation): gcloud container fleet mesh describe --project=plenary-truck-485623-p7 # Manual re-registration if needed: gcloud container fleet memberships register knoe-dev-0 \ --gke-cluster=us-west3/knoe-dev-0 \ --enable-workload-identity \ --project=plenary-truck-485623-p7 gcloud container fleet mesh update \ --management=automatic \ --memberships=knoe-dev-0 \ --project=plenary-truck-485623-p7 ``` --- ## install.sh / deploy.sh pre-flight checklist ### `conf/gke.cfg` (interactive installer — `./install.sh`) Before running `./install.sh` (especially "Initialization Scripts"), confirm these are correct: ```ini [Inputs] init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 init_cluster.db_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 env_setup.APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 env_setup.DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 [Global] KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 CNPG_ELIGIBLE_NODES = ``` ### `conf/service/prod.cfg` (unattended deploy — `./deploy.sh`) Same cluster context entries are required here too: ```ini [Inputs] init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 init_cluster.db_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 env_setup.APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 env_setup.DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 [Global] KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 ARTIFACT_REGISTRY = us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system SERVICE_NAMESPACE = knoe-system REGISTRY_NAMESPACE = knoe-system ``` **Why these matter:** `Milestone._get_script_env()` (in `knoe/milestone.py`) reads these to set `KUBECONTEXT=app_ctx` for common services and `DB_CLUSTER_KUBECONTEXT=db_ctx` for CNPG ops. Without them, all kubectl calls use the ambient context, which may be the DB cluster. Missing `init_cluster.app_cluster_kubecontext` → `_cluster_kubecontext("app")` returns `""` → installer falls back to `Global.KUBECONTEXT` for **both** app and db environments → **Garage deploys to knoe-dev-cnpg-0** (wrong). > **Env-contamination guard (live):** `deploy.sh` calls > [`etc/preflight_kubecontext.sh`](etc/preflight_kubecontext.sh) and refuses > to proceed when `kubectl config current-context` doesn't match the > `[Global] APP_CLUSTER_KUBECONTEXT` of the active config. `install.sh` > prints the inherited context up-front (mode-aware strict gate is the > Python TUI's responsibility once the welcome screen records a mode). > Bypass with `KNOE_SKIP_KUBECONTEXT_GUARD=true` for deliberate > cross-cluster maintenance. **History:** the guard was filed in response > to the 2026-04-28 14:00 UTC outage — an `install.sh --mode k3d` run with > the shell pointed at GKE replaced the GCS-backed ObjectStore with a > Garage-backed one, then Garage filled up and backups silently failed for > hours. Closes drift R4 / queue item #1. > **Env-contamination guard (live):** `deploy.sh` calls > [`etc/preflight_kubecontext.sh`](etc/preflight_kubecontext.sh) and refuses > to proceed when `kubectl config current-context` doesn't match the > `[Global] APP_CLUSTER_KUBECONTEXT` of the active config. `install.sh` > prints the inherited context up-front (mode-aware strict gate is the > Python TUI's responsibility once the welcome screen records a mode). > Bypass with `KNOE_SKIP_KUBECONTEXT_GUARD=true` for deliberate > cross-cluster maintenance. **History:** the guard was filed in response > to the 2026-04-28 14:00 UTC outage — an `install.sh --mode k3d` run with > the shell pointed at GKE replaced the GCS-backed ObjectStore with a > Garage-backed one, then Garage filled up and backups silently failed for > hours. Closes drift R4 / queue item #1. ### Get current CNPG node names ```bash kubectl --context=gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 get nodes -o name ``` --- ## Key scripts | Script | Purpose | |---|---| | `scripts/reset_clusters.sh` | Delete + recreate both GKE clusters (`e2-standard-2` × 3 each) | | `scripts/patch_garage_cross_cluster.sh` | Remove garage from DB cluster, configure GCS barman ObjectStore | | `etc/init_cnpg_gke.sh` | Provision CNPG on GKE with GCS backup via Workload Identity | | `etc/init_cnpg_backup.sh` | Configure barman plugin + trigger initial backup | --- ## Cluster code constants (`knoe/core/actions.py`) ```python DEFAULT_APP_CLUSTER_NAME = "knoe-dev-0" DEFAULT_DB_CLUSTER_NAME = "knoe-dev-cnpg-0" DEFAULT_APP_CLUSTER_MACHINE_TYPE = "e2-standard-2" DEFAULT_DB_CLUSTER_MACHINE_TYPE = "e2-standard-2" ``` --- ## GCP project - Project: `plenary-truck-485623-p7` - Region: `us-west3` - Artifact Registry: `us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system` - CPU quota: 16 vCPUs (all regions) — 2 clusters × 3 × `e2-standard-2` = 12 vCPUs used - SSD quota: `SSD_TOTAL_GB = 300 GB` — fully consumed by CNPG; all other PVCs must use `standard` (pd-standard / HDD) --- ## Reality TODOs / Drift Log Quick reference. Each entry links to the master index where context, owner, and rank live. | # | Drift | Where described above | Where tracked | |---|---|---|---| | _(none currently)_ | | | | **Closed in 2026-04-29 stabilization session:** Garage on DB cluster removed; cluster pods migrated to `cnpg-backup-sa` via CNPG v1.29.0 `spec.serviceAccountName`; both operators restarted clean. **Closed 2026-05-01:** R4 — installer env-contamination guard now live in `deploy.sh` (strict) + `install.sh` (informational notice). Helper at [`etc/preflight_kubecontext.sh`](etc/preflight_kubecontext.sh).