- knoe-db.yaml: switch to CNPG-managed TLS cert with serverAltDNSNames (pg.prole.org + knoe-db-rw cluster service) — removes static serverTLSSecret/serverCASecret - dns.yml: add pg.prole.org A record to prole_k3s_dns_records (10.0.0.3, 10.0.0.6) for Ansible-managed split-horizon DNS via Samba AD DC - k3s.cfg: align KNOE_HOME paths to ~/dev/prole, add PROLE_KDC_* vars, remove hardcoded KUBECTL_CONTEXT (kubeconfig current-context is authoritative) - prod.cfg: add PROLE_KDC_STORAGE_CLASS = prole-iscsi - onepassword.py: skip vault check gracefully when no 1Password session active (non-TTY) - CLAUDE.md: document production postgres connection string and DNS/CA cert ops Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
12 KiB
CLAUDE.md — knoe project context
Loaded automatically by Claude Code (CLI and IntelliJ plugin) as project context.
Repo role: knoe-db platform repo
This is knoe-db (remote: git@git-ssh.knoe.dev:knoe.dev/knoe-db.git, configured as both origin and knoe). It's the platform's source of truth — authority/, knoe/, etc/init_*.sh, deploy/gcp/gke/*, the test pipeline, all live here.
Customer deploys are intended to live as branches in this repo (e.g. a future customer/prole.org), not as separate forks. As of this writing, no customer branch is active — the prole→knoe rebrand has merged into main and there is no separate ~/dev/prole working tree under development. See docs/plans/customer-deploy-resync.md for the original plan, currently dormant.
Master TODO
Single source of truth for unfinished work, including the reality-vs-intent gaps flagged below: docs/TODO.md.
Active plans
docs/plans/README.md— Index and conventions for this directory.docs/plans/knoe-auth-round-1.md— Identity backbone (Kerberos KDC + invite-OTP + TOTP). Shipped.docs/plans/deployment-modes.md— Four installer modes (min/k3d/k3s/gke). Shipped (Phase 0).docs/pipeline-phases.md— Autobuild & test pipeline phase reference. Phase 0 ✅, Phase 1 next (first task: f-string fix atknoe/core/ops/cloudnative_pg.py:1372).docs/plans/customer-deploy-resync.md— The original plan to converge~/dev/proleontoknoe-db/mainas a customer-deploy branch. Dormant (rebrand is now in main, no separate prole tree active).
Dual-cluster GKE architecture
This project uses two separate GKE Standard clusters in us-west3, both currently provisioned with e2-standard-2 × 3 nodes (2 vCPU / 8 GB each, ~7.1 GB allocatable). Verify with gcloud container clusters list.
| Cluster | Context | Role |
|---|---|---|
knoe-dev-0 |
gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 |
App cluster — Garage, registry, OpenBao, Kong, GitLab, monitoring |
knoe-dev-cnpg-0 |
gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 |
DB cluster — CNPG/PostgreSQL only |
Storage quota: the project has SSD_TOTAL_GB = 300 GB in us-west3, fully consumed by CNPG. All non-CNPG PVCs must use standard (pd-standard / HDD) — not standard-rwo / premium-rwo, which are SSD-backed and will fail to provision with a quota error. GITLAB_GITALY_STORAGE_CLASS = standard is set in conf/gke.cfg accordingly.
Resource allocation
| Resource | Cluster | Namespace |
|---|---|---|
| CNPG operator (v1.29.0) | knoe-dev-cnpg-0 |
cnpg-system |
PostgreSQL cluster (knoe-db) |
knoe-dev-cnpg-0 |
knoe-db-0 |
| Barman Cloud plugin (v0.12.0) | knoe-dev-cnpg-0 |
cnpg-system |
| cert-manager | knoe-dev-cnpg-0 |
cert-manager |
| Garage (S3 object store) | knoe-dev-0 |
knoe-system |
| Registry | knoe-dev-0 |
knoe-system |
| OpenBao | knoe-dev-0 |
knoe-system |
| Kong API gateway | knoe-dev-0 |
knoe-system |
| Monitoring | knoe-dev-0 |
monitoring |
Garage runs ONLY in knoe-dev-0. The DB cluster (knoe-dev-cnpg-0) has none — was removed 2026-04-29. Do NOT redeploy Garage to the DB cluster (use GCS for backups there).
CNPG backups → GCS
Backups use GCS with Workload Identity:
- Data + WAL:
gs://knoe-0-backups/(single bucket;knoe-db/base/andknoe-db/wals/prefixes).gs://knoe-0-wal/exists but is unused. - GCP SA:
cnpg-backup@plenary-truck-485623-p7.iam.gserviceaccount.com(roles on bucket:storage.objectAdmin,storage.legacyBucketReader) - K8s SA: cluster pods run as
cnpg-backup-sainknoe-db-0(set viacluster.spec.serviceAccountName, requires CNPG ≥ v1.29.0). The SA is annotated withiam.gke.io/gcp-service-account=cnpg-backup@…. TwoRoleBindingsubjects (knoe-dbandknoe-db-barman-cloud) includecnpg-backup-saso the pod has the same RBAC the auto-generated SA would have had. - ObjectStore manifest:
k8s/knoe/knoe-db-barman-objectstore-gcs.yaml— includesgoogleCredentials.gkeEnvironment: true(required by plugin-barman-cloud v0.12.0).
Setup script: etc/init_cnpg_gke.sh (creates buckets, GCP SA, WI binding, applies CNPG cluster).
Service mesh (Cloud Service Mesh / Istio)
Both clusters are registered in the knoe-0 GCP fleet with automatic Cloud Service Mesh management. This is automated in scripts/reset_clusters.sh (Phase 7) — no longer requires GCP web console.
# Check mesh provisioning status (~10 min after cluster creation):
gcloud container fleet mesh describe --project=plenary-truck-485623-p7
# Manual re-registration if needed:
gcloud container fleet memberships register knoe-dev-0 \
--gke-cluster=us-west3/knoe-dev-0 \
--enable-workload-identity \
--project=plenary-truck-485623-p7
gcloud container fleet mesh update \
--management=automatic \
--memberships=knoe-dev-0 \
--project=plenary-truck-485623-p7
install.sh / deploy.sh pre-flight checklist
conf/gke.cfg (interactive installer — ./install.sh)
Before running ./install.sh (especially "Initialization Scripts"), confirm these are correct:
[Inputs]
init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
init_cluster.db_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
env_setup.APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
env_setup.DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
[Global]
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
CNPG_ELIGIBLE_NODES = <comma-separated node names from knoe-dev-cnpg-0>
conf/service/prod.cfg (unattended deploy — ./deploy.sh)
Same cluster context entries are required here too:
[Inputs]
init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
init_cluster.db_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
env_setup.APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
env_setup.DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
[Global]
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
ARTIFACT_REGISTRY = us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system
SERVICE_NAMESPACE = knoe-system
REGISTRY_NAMESPACE = knoe-system
Why these matter: Milestone._get_script_env() (in knoe/milestone.py) reads these to set KUBECONTEXT=app_ctx for common services and DB_CLUSTER_KUBECONTEXT=db_ctx for CNPG ops. Without them, all kubectl calls use the ambient context, which may be the DB cluster.
Missing init_cluster.app_cluster_kubecontext → _cluster_kubecontext("app") returns "" → installer falls back to Global.KUBECONTEXT for both app and db environments → Garage deploys to knoe-dev-cnpg-0 (wrong).
Env-contamination guard (live):
deploy.shcallsetc/preflight_kubecontext.shand refuses to proceed whenkubectl config current-contextdoesn't match the[Global] APP_CLUSTER_KUBECONTEXTof the active config.install.shprints the inherited context up-front (mode-aware strict gate is the Python TUI's responsibility once the welcome screen records a mode). Bypass withKNOE_SKIP_KUBECONTEXT_GUARD=truefor deliberate cross-cluster maintenance. History: the guard was filed in response to the 2026-04-28 14:00 UTC outage — aninstall.sh --mode k3drun with the shell pointed at GKE replaced the GCS-backed ObjectStore with a Garage-backed one, then Garage filled up and backups silently failed for hours. Closes drift R4 / queue item #1.
Env-contamination guard (live):
deploy.shcallsetc/preflight_kubecontext.shand refuses to proceed whenkubectl config current-contextdoesn't match the[Global] APP_CLUSTER_KUBECONTEXTof the active config.install.shprints the inherited context up-front (mode-aware strict gate is the Python TUI's responsibility once the welcome screen records a mode). Bypass withKNOE_SKIP_KUBECONTEXT_GUARD=truefor deliberate cross-cluster maintenance. History: the guard was filed in response to the 2026-04-28 14:00 UTC outage — aninstall.sh --mode k3drun with the shell pointed at GKE replaced the GCS-backed ObjectStore with a Garage-backed one, then Garage filled up and backups silently failed for hours. Closes drift R4 / queue item #1.
Get current CNPG node names
kubectl --context=gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 get nodes -o name
Key scripts
| Script | Purpose |
|---|---|
scripts/reset_clusters.sh |
Delete + recreate both GKE clusters (e2-standard-2 × 3 each) |
scripts/patch_garage_cross_cluster.sh |
Remove garage from DB cluster, configure GCS barman ObjectStore |
etc/init_cnpg_gke.sh |
Provision CNPG on GKE with GCS backup via Workload Identity |
etc/init_cnpg_backup.sh |
Configure barman plugin + trigger initial backup |
Cluster code constants (knoe/core/actions.py)
DEFAULT_APP_CLUSTER_NAME = "knoe-dev-0"
DEFAULT_DB_CLUSTER_NAME = "knoe-dev-cnpg-0"
DEFAULT_APP_CLUSTER_MACHINE_TYPE = "e2-standard-2"
DEFAULT_DB_CLUSTER_MACHINE_TYPE = "e2-standard-2"
GCP project
- Project:
plenary-truck-485623-p7 - Region:
us-west3 - Artifact Registry:
us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system - CPU quota: 16 vCPUs (all regions) — 2 clusters × 3 ×
e2-standard-2= 12 vCPUs used - SSD quota:
SSD_TOTAL_GB = 300 GB— fully consumed by CNPG; all other PVCs must usestandard(pd-standard / HDD)
Reality TODOs / Drift Log
Quick reference. Each entry links to the master index where context, owner, and rank live.
| # | Drift | Where described above | Where tracked |
|---|---|---|---|
| (none currently) |
Closed in 2026-04-29 stabilization session: Garage on DB cluster removed; cluster pods migrated to cnpg-backup-sa via CNPG v1.29.0 spec.serviceAccountName; both operators restarted clean.
Closed 2026-05-01: R4 — installer env-contamination guard now live in deploy.sh (strict) + install.sh (informational notice). Helper at etc/preflight_kubecontext.sh.
k3s CNPG database (production)
The prole.org k3s CNPG cluster is this project's production PostgreSQL database.
psql "host=pg.prole.org port=5432 user=chrisfu dbname=postgres sslmode=verify-full sslrootcert=$HOME/.knoe/knoe-db-ca.crt"
| Detail | Value |
|---|---|
| External hostname | pg.prole.org:5432 |
| Internal service | knoe-db-rw.knoe-db.svc.cluster.local:5432 |
| kubectl context | prole-service-cluster |
| Namespace | knoe-db |
| CA cert | ~/.knoe/knoe-db-ca.crt |
| sslmode | verify-full |
DNS: pg.prole.org resolves internally via split-horizon DNS on myrddin.prole.org (Samba AD DC) to the k3s ServiceLB node IPs (10.0.0.3, 10.0.0.6). External DNS resolves to the public IP — do not access from outside the LAN without a VPN.
CA cert refresh (after CNPG cert rotation):
kubectl --context=prole-service-cluster -n knoe-db \
get secret knoe-db-ca -o jsonpath='{.data.ca\.crt}' | base64 -d > ~/.knoe/knoe-db-ca.crt
Node mobility: to move the postgres LoadBalancer to a different node, update the Samba DNS A records:
ssh myrddin.prole.org "sudo samba-tool dns delete myrddin.prole.org prole.org pg A <OLD_IP> -U Administrator"
ssh myrddin.prole.org "sudo samba-tool dns add myrddin.prole.org prole.org pg A <NEW_IP> -U Administrator"