feat(gke): finalize dual-cluster app/db topology

- separate default app and db cluster contexts (knoe-dev-0 and knoe-cnpg-0) across actions, ops flows, and installer config\n- update CNPG placement and GKE manifests to keep PostgreSQL on the db cluster while app services stay on the app cluster\n- add GCS-backed Barman ObjectStore manifest and cross-cluster Garage patch workflow\n- refresh UI cluster/database/service screens and monitoring wiring for dual-cluster operation\n- add reset/patch scripts, db context selection test coverage, and architecture/network documentation updates

Co-authored-by: Junie <junie@jetbrains.com>
This commit is contained in:
chrisfu 2026-04-08 23:12:51 -07:00
parent 64e1cef16c
commit 5cd43c873c
23 changed files with 1276 additions and 393 deletions

99
CLAUDE.md Normal file
View File

@ -0,0 +1,99 @@
# CLAUDE.md — prole project context
> Loaded automatically by Claude Code (CLI and IntelliJ plugin) as project context.
---
## Dual-cluster GKE architecture
This project uses **two separate GKE Standard clusters** in `us-west3`, both e2-small × 3 nodes:
| Cluster | Context | Role |
|---|---|---|
| `knoe-dev-0` | `gke_plenary-truck-485623-p7_us-west3_knoe-dev-0` | App cluster — Garage, registry, OpenBao, Kong, GitLab, monitoring |
| `knoe-cnpg-0` | `gke_plenary-truck-485623-p7_us-west3_knoe-cnpg-0` | DB cluster — CNPG/PostgreSQL only |
### Resource allocation
| Resource | Cluster | Namespace |
|---|---|---|
| CNPG operator | `knoe-cnpg-0` | `cnpg-system` |
| PostgreSQL cluster (`knoe-db`) | `knoe-cnpg-0` | `knoe-db-0` |
| Barman Cloud plugin | `knoe-cnpg-0` | `cnpg-system` |
| cert-manager | `knoe-cnpg-0` | `cert-manager` |
| Garage (S3 object store) | `knoe-dev-0` | `knoe-system` |
| Registry | `knoe-dev-0` | `knoe-system` |
| OpenBao | `knoe-dev-0` | `knoe-system` |
| Kong API gateway | `knoe-dev-0` | `knoe-system` |
| Monitoring | `knoe-dev-0` | `monitoring` |
**Garage runs ONLY in `knoe-dev-0`.** Do NOT deploy Garage to `knoe-cnpg-0`.
### CNPG backups → GCS (not Garage)
Backups use **GCS with Workload Identity** (not Garage):
- Data bucket: `gs://knoe-0-backups/`
- WAL bucket: `gs://knoe-0-wal/`
- GCP SA: `cnpg-backup@plenary-truck-485623-p7.iam.gserviceaccount.com`
- K8s SA: `cnpg-backup-sa` in `knoe-db-0` (annotated with WI)
- ObjectStore manifest: `k8s/prole/knoe-db-barman-objectstore-gcs.yaml`
Setup script: `etc/init_cnpg_gke.sh` (creates buckets, GCP SA, WI binding, applies CNPG cluster).
---
## install.sh pre-flight checklist
Before running `./install.sh` (especially "Initialization Scripts"), confirm these are correct in `conf/prod/prole.cfg`:
```ini
[Inputs]
init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
init_cluster.db_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-cnpg-0
env_setup.APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
env_setup.DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-cnpg-0
[Global]
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-cnpg-0
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
CNPG_ELIGIBLE_NODES = <comma-separated node names from knoe-cnpg-0>
```
Missing `init_cluster.app_cluster_kubecontext``_cluster_kubecontext("app")` returns `""` → installer falls back to `Global.KUBECONTEXT` (DB cluster) for **both** app and db environments → **Garage deploys to knoe-cnpg-0** (wrong).
### Get current CNPG node names
```bash
kubectl --context=gke_plenary-truck-485623-p7_us-west3_knoe-cnpg-0 get nodes -o name
```
---
## Key scripts
| Script | Purpose |
|---|---|
| `scripts/reset_clusters.sh` | Delete + recreate both GKE clusters (e2-small × 3 each) |
| `scripts/patch_garage_cross_cluster.sh` | Remove garage from DB cluster, configure GCS barman ObjectStore |
| `etc/init_cnpg_gke.sh` | Provision CNPG on GKE with GCS backup via Workload Identity |
| `etc/init_cnpg_backup.sh` | Configure barman plugin + trigger initial backup |
---
## Cluster code constants (`knoe/core/actions.py`)
```python
DEFAULT_APP_CLUSTER_NAME = "knoe-dev-0"
DEFAULT_DB_CLUSTER_NAME = "knoe-cnpg-0" # also: DEFAULT_CNPG_CLUSTER_NAME
DEFAULT_APP_CLUSTER_MACHINE_TYPE = "e2-small"
DEFAULT_DB_CLUSTER_MACHINE_TYPE = "e2-small"
```
---
## GCP project
- Project: `plenary-truck-485623-p7`
- Region: `us-west3`
- Artifact Registry: `us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system`
- CPU quota: 16 vCPUs (all regions) — 2 clusters × 3 × e2-small = 12 vCPUs used

View File

@ -1,13 +1,15 @@
{
"assignments": {
"0": "gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z",
"1": "gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z",
"2": "gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z"
"0": "gke-knoe-dev-0-default-pool-503666dd-d39h",
"1": "gke-knoe-dev-0-default-pool-98ee334c-qpf9",
"2": "gke-knoe-dev-0-default-pool-b26cfdbc-r144"
},
"cluster_name": "knoe-db",
"desired_instances": 3,
"eligible_nodes": [
"gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z"
"gke-knoe-dev-0-default-pool-503666dd-d39h",
"gke-knoe-dev-0-default-pool-98ee334c-qpf9",
"gke-knoe-dev-0-default-pool-b26cfdbc-r144"
],
"metadata": {
"prior_plan_present": true,
@ -15,7 +17,7 @@
"regenerated": true,
"reused": false
},
"plan_hash": "c1213b43098112f3",
"plan_id": "cnpg-placement-c1213b43098112f3",
"plan_hash": "c7f53cb912958603",
"plan_id": "cnpg-placement-c7f53cb912958603",
"schema_version": "v1"
}

View File

@ -1,13 +1,15 @@
{
"assignments": {
"0": "gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z",
"1": "gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z",
"2": "gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z"
"0": "gke-knoe-dev-0-default-pool-503666dd-d39h",
"1": "gke-knoe-dev-0-default-pool-98ee334c-qpf9",
"2": "gke-knoe-dev-0-default-pool-b26cfdbc-r144"
},
"cluster_name": "knoe-db",
"desired_instances": 3,
"eligible_nodes": [
"gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z"
"gke-knoe-dev-0-default-pool-503666dd-d39h",
"gke-knoe-dev-0-default-pool-98ee334c-qpf9",
"gke-knoe-dev-0-default-pool-b26cfdbc-r144"
],
"metadata": {
"prior_plan_present": true,
@ -15,7 +17,7 @@
"regenerated": false,
"reused": true
},
"plan_hash": "c1213b43098112f3",
"plan_id": "cnpg-placement-c1213b43098112f3",
"plan_hash": "c7f53cb912958603",
"plan_id": "cnpg-placement-c7f53cb912958603",
"schema_version": "v1"
}

View File

@ -1,5 +1,5 @@
; Prole Master Configuration File
; Generated by install.py on 2026-04-07 09:23:28
; Generated by install.py on 2026-04-08 23:04:37
; This file is used as input for Ansible deployment and k8s cluster creation.
[User]
@ -56,7 +56,7 @@ init_cluster.cluster_env = dev
init_cluster.deployment_target = prole-dev-cluster
init_cluster.gitops_enabled = false
init_cluster.k3s_server_url = https://myrddin.prole.org:6443
init_cluster.k3s_token = ${PROLE_SECRET:v1:9Vez82_9Cxpex2oY:oWLW5wEBqIzPEIOt3EZnB2JLj4IAORTKMxMxtNwDC58Q266OzM83h2l-PV30e_V4QblEmDmGpQyAg28FkDNBpxlYa6HsNk8HnkCLQIP8MK_j51e3ft6UUnMISGSKjkoSl2wMdLy5k9QhCdoQeOGhxN2BpiEscZBUgpDIQw==}
init_cluster.k3s_token = ${PROLE_SECRET:v1:VehUafohCaBAyMN6:5gV4RGzYRcxlD4_bQg4Z30m9dq6SGuVGFBFF7osArk8RVc1vgNdIW7qtFewIwkskNe9xbSiC2-gvjwKKG5Nv1JJlzhaAR8NNH-_g4PaoGychXb18YL-8oWDFByGMNgokUObDEc3uJ1CNXCkcLtawaruRNwQ_Myz48TT1TQ==}
init_cluster.kerberos_enabled = true
init_cluster.mode = k3d
init_cluster.start_cluster = true
@ -100,7 +100,7 @@ DOCKER_PRELOAD = false
KNOE_DB_USER = root
OPTIONAL_WORKLOADS_MIN_READY_SCHEDULABLE_NODES = 2
PROLE_K3S_SERVER = https://myrddin.prole.org:6443
PROLE_K3S_TOKEN = ${PROLE_SECRET:v1:J-6LlQf6BDV1Szx0:b1D60bQkckSeXNHuMobcF5LXy2-ui6oarFisXzjOjTPX2dvXua_v3O8kBicxjvStyTUX5N1EDnunnm4eGWYElfXtgdoHsF8kc_7KkRiNt37ad7-KVShwKMHPDhNsJA8TQ_L3PjhxpQbrwyCo2PWkXoNlCtjn57qUHeqW1Q==}
PROLE_K3S_TOKEN = ${PROLE_SECRET:v1:aOx6XJ6T1bwAWJqO:U_Q9bFmBqns_qvVdmIiwlHkz7Y_vRxnlbFij7pHoTfJZP8iPblmtfZjdi-Z1xOAuxnkYy6-0j1SMBLsIJWcJ0G1g5laIfqBB9d1iv9VLKeNmnUrGV-xgiv8-aPn2TM3UtN8hFrfaRKfeiTG01iyKLG6IoNngF2xA8rb7KQ==}
PROLE_OPENTOFU_URL = http://127.0.0.1:8080
SERVICE_NAMESPACE = knoe-system
SUPABASE_PV_BASE = /synology/d005
@ -143,7 +143,7 @@ PROLE_SERVICE = /Users/chrisfu/dev/prole/etc
[Ollama]
OLLAMA_HOST = http://fairyland.prole.org:11434
OLLAMA_SERVERS = fairyland.prole.org:11434,k3d.localhost:11434,morgoth.prole.org:11434
OLLAMA_SERVERS = 10.0.0.208:11434,fairyland.prole.org:11434,k3d.localhost:11434,morgoth.prole.org:11434
OLLAMA_SERVER_HOST = fairyland.prole.org
OLLAMA_SERVER_PORT = 11434
@ -173,7 +173,7 @@ MODE = k3d
CLUSTER_ENV = prole-service-cluster
DISPLAY_NAME = prole-service-cluster
K3S_SERVER_URL = https://myrddin.prole.org:6443
K3S_TOKEN = ${PROLE_SECRET:v1:sOSNr3SQq1As2lxH:RM-5UCAufudeBWFaEST-pC4y17Gs0Fu106yX_0VsseCRcjrLQUT4f4uFmUxY8ICYE9zArCyJEtgZNNGoRcybQfBUoTTAAAuGrK1nMuakuRd8j_aPEnH3XHSLgraIJIlizCSpA8miq1U2UJRv6KmJjxzoMmfKpExKqKiWtw==}
K3S_TOKEN = ${PROLE_SECRET:v1:7OALwMwbkRA9fIQG:RKg4R1_K6niZRAf9rCngO_Yvss4jMKsvQ1iVyZoNJa9yMgAOlKBZBKRY74qsnpRgfcjyxAbj9V42crkaEUaExR4Rr64cvAk5C3aUrC8kmBNLnnbSzydGxsF8Vgh4SCIg0xo9RyJF5h9M1r-F2Ua2P6aYZTjRYcW-Ro0xiA==}
MODE = k3s
PIPELINE_URL = http://127.0.0.1:8080

View File

@ -1,5 +1,5 @@
; Prole Master Configuration File
; Generated by install.py on 2026-04-07 09:24:02
; Generated by install.py on 2026-04-08 23:05:25
; This file is used as input for Ansible deployment and k8s cluster creation.
[User]
@ -56,7 +56,7 @@ init_cluster.cluster_env = prod
init_cluster.deployment_target = prole-prod-cluster
init_cluster.gitops_enabled = false
init_cluster.k3s_server_url = https://myrddin.prole.org:6443
init_cluster.k3s_token = ${PROLE_SECRET:v1:8jsa5B6uohhtyjxS:pfwgP6OkFpEcR-f4_3_DkCYwRuWlGIWkdg--1vesoya0OpGR9V6nFFW5DsEs5ARzBFkxkIxbgiNDOXC-t7oa9u578IDg2fO0FX4BXguvhqJi8lFjltEbjP4EL8LaHmkBpEZ3ko9egaNqVc03mFn5Fl5Ndc_uj9a-Ehe1jg==}
init_cluster.k3s_token = ${PROLE_SECRET:v1:EnJkApzkJeOQYGnQ:wDfFl2LAFzu6KmMpucPyIiKVcffzx7J2dPg8HmdN2o-4LSEhZMBjQupAJZuBL0Uy6AiEB1aO4UKjEAy-ahUKU2GbkJBodTdZrjFkUc-KHdn-MINCQ9V4S3NspZLEVNizKYtEuXr9aUpmKdMv-NllrFx-9K8YAfclOPikEw==}
init_cluster.kerberos_enabled = true
init_cluster.mode = k8s
init_cluster.start_cluster = true
@ -87,16 +87,18 @@ supabase_config.pv_node = gandalf.prole.org
[Global]
; Variables used by name in more than one place or assumed global scope
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
ARTIFACT_REGISTRY = us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system
ARTIFACT_REGISTRY_AVAILABLE = true
CLUSTER_ENV = prod
CLUSTER_NAME = knoe-db
CNPG_ELIGIBLE_NODES = gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z
CNPG_ELIGIBLE_NODES = gke-knoe-dev-0-default-pool-503666dd-d39h,gke-knoe-dev-0-default-pool-98ee334c-qpf9,gke-knoe-dev-0-default-pool-b26cfdbc-r144
CNPG_PLACEMENT_PLAN_FILE = /Users/chrisfu/dev/prole/conf/cnpg-placement/knoe-system-knoe-db.json
CNPG_PLACEMENT_PLAN_HASH = c1213b43098112f3
CNPG_PLACEMENT_PLAN_ID = cnpg-placement-c1213b43098112f3
CNPG_STAGE1_NODE = gk3-knoe-dev-0-nap-4jvx077g-1a599f8c-pv5z
CNPG_PLACEMENT_PLAN_HASH = c7f53cb912958603
CNPG_PLACEMENT_PLAN_ID = cnpg-placement-c7f53cb912958603
CNPG_STAGE1_NODE = gke-knoe-dev-0-default-pool-503666dd-d39h
DATABASE_NAMESPACE = knoe-db-18-008-18-009-18-013
DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-cnpg-0
DB_HOST_PORT = 5432
DEPLOYMENT_MODE = k8s
DEPLOYMENT_TARGET = prole-prod-cluster
@ -106,7 +108,7 @@ KUBECONFIG = /Users/chrisfu/.kube/config
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
OPTIONAL_WORKLOADS_MIN_READY_SCHEDULABLE_NODES = 2
PROLE_K3S_SERVER = https://myrddin.prole.org:6443
PROLE_K3S_TOKEN = ${PROLE_SECRET:v1:10g1WjSwQq3MiUNh:XkobEkCjr8C5VHxPEFeXEOTA6CA5olncCUpaRgUNZuwrsWBKR_p0jeo5xAKcsRkYT-s1mrmV4ZaqgP93JfBGnCybgC5kHJe-ch_UVf4rkJKV75kVnHzrSlIRYZdng-o1_rwfLEARxJ2nLHvmNLnLV6n-TZUmSy3PZYfOeg==}
PROLE_K3S_TOKEN = ${PROLE_SECRET:v1:on0oK-bUBCGGiibf:RQSubonBNj8V60-zU10D2FYToViJUNt-WyC23-bb6yQrf72WcvkfFMWsiEQA5OxRWvKM4TaaezLF4WFW484p3S9g0Y3oB_9qv7CSc15ZaOOHTUk2DpNOKahV5fItY8bPjmUKWTgKrShSWl6j2C5SFV7fYwmq6TTOuygFPQ==}
SERVICE_NAMESPACE = knoe-system
SUPABASE_PV_BASE = /synology/d005
SUPABASE_PV_BASE_DIR = /synology/d005
@ -164,9 +166,11 @@ SUPABASE_ENABLED = False
; No configuration values captured yet for this section.
[Initialize Cluster]
DB_BOOT_DISK_SIZE_GB = 50
DB_BOOT_DISK_TYPE = pd-standard
ENVIRONMENT = prod
K3S_SERVER_URL = https://myrddin.prole.org:6443
K3S_TOKEN = ${PROLE_SECRET:v1:ILhG782khL7eyT7y:b3s76jimM0Q7hYH7hnSYM2XOfBXknl9mXvjzfLI4jTz-c4dHLCJXG8WlOiFf8FM0HfoIK-xu7NxSdym2UXBzCkuENFIJ2Y1GwNr9Nt2ForgfuwKaYVfoGnDc0ERp_S6qaKP01uLDeR95hUS56Wtuh3TNPN0gp14qvBjX4g==}
K3S_TOKEN = ${PROLE_SECRET:v1:y0saV3fVIh0T5E6G:ek4jKHO09N2Z5uRO3mwsSgimKpVodGN6zE6d6TtYHMqoFlcJ_AZRUf5dMCqAZ8eE9gydBu_k-ey44CQOAmK8m9KryFU9HDL1ioL_AWGfVJIcG16vgVsaiaVO3h7Eppv__qdCzSqaEgwJqNHC8fTYQtfAXBouk3arsM8_Ew==}
[Dev Cluster (k3d)]
CLUSTER_ENV = dev
@ -178,7 +182,7 @@ MODE = k3d
CLUSTER_ENV = prole-service-cluster
DISPLAY_NAME = prole-service-cluster
K3S_SERVER_URL = https://myrddin.prole.org:6443
K3S_TOKEN = ${PROLE_SECRET:v1:wxpiSOY6t_5_gaiT:vfyzLLUJUYc3SmQ48yDYrTpeWQ2LfVoiEWvzRwDPzm1Mx6a6aW0QjJLK52Rbwe5hVtwwDjcglakhM8eBjFV-ceFaLIl7Tsr7CwARLuPD6NBJQrc1qnHuKTh5DOSMGF_emnP_fWT06-M242VZImmVE_2Q82Cr45Gd91h93g==}
K3S_TOKEN = ${PROLE_SECRET:v1:ts8tzrsAcvLHInVF:xQVZWcdVVVI7tAYqG_048NKLPw25ENxCqtcdiyPb4tGSuyjeqSn9kXoMMR-8TvcBGPpug4dXa3IQZEwIHksQvsWkD-MtbPG5uaehhz7wsLJJ1ILkhd3PGjeyCkFn9zFzLEyPBbJoTRY2I91V_Qck7Vlkkn8zOb3Y9NXuFg==}
MODE = k3s
[GCP]
@ -201,7 +205,7 @@ Configure routing, TLS, and frontdoor domain mappings
Prepare migration artifacts from source environment
Generate OpenTofu plan and apply production pipeline
MODE = k8s
OPENTOFU_VARS = {"admin_user":"knoe_admin","app_database":"knoey","app_user":"knoey_app","backup_bucket":"knoe-0-backups","backup_provider":"gcs","backup_retention_days":14,"bootstrap_admin_email":"admin@knoey.com","cluster_name":"knoe-dev-0","cnpg_cluster_name":"knoe-db","db_instances":3,"db_pgdata_size_gi":100,"db_pgdata_storage_class":"premium-rwo","db_storage_class":"premium-rwo","db_storage_size_gi":100,"db_wal_size_gi":25,"db_wal_storage_class":"premium-rwo","ecosystem_id":0,"environment":"production","frontdoor_host":"knoey.com","gcp_project_id":"plenary-truck-485623-p7","gcp_region":"\u2713 us-west3","google_hosted_domain":"","meta_database":"knoe_meta","migration_continuous_until_cutover":false,"migration_mode":"snapshot-restore","migration_source_database":"knoey","migration_source_environment":"prole.org","migration_source_host":"knoe-local-db.prole.org","migration_source_password_ref":"secretref://local-source-db-password","migration_source_port":5432,"migration_source_user":"replication_user","namespace":"ecosystem-0","oidc_client_id_ref":"secretref://google-oidc-client-id","oidc_client_secret_ref":"secretref://google-oidc-client-secret","oidc_issuer":"https://accounts.google.com","platform_domain":"knoe.dev","postgres_version":"16","project_id":"plenary-truck-485623-p7","region":"\u2713 us-west3","tls_mode":"managed","wal_bucket":"knoe-0-wal","workload_identity_sa":"cnpg-backup"}
OPENTOFU_VARS = {"admin_user":"knoe_admin","app_database":"knoey","app_user":"knoey_app","backup_bucket":"knoe-0-backups","backup_provider":"gcs","backup_retention_days":14,"bootstrap_admin_email":"admin@knoey.com","cluster_name":"knoe-cnpg-0","cnpg_cluster_name":"knoe-db","db_instances":3,"db_pgdata_size_gi":100,"db_pgdata_storage_class":"premium-rwo","db_storage_class":"premium-rwo","db_storage_size_gi":100,"db_wal_size_gi":25,"db_wal_storage_class":"premium-rwo","ecosystem_id":0,"environment":"production","frontdoor_host":"knoey.com","gcp_project_id":"plenary-truck-485623-p7","gcp_region":"\u2713 us-west3","google_hosted_domain":"","meta_database":"knoe_meta","migration_continuous_until_cutover":false,"migration_mode":"snapshot-restore","migration_source_database":"knoey","migration_source_environment":"prole.org","migration_source_host":"knoe-local-db.prole.org","migration_source_password_ref":"secretref://local-source-db-password","migration_source_port":5432,"migration_source_user":"replication_user","namespace":"ecosystem-0","oidc_client_id_ref":"secretref://google-oidc-client-id","oidc_client_secret_ref":"secretref://google-oidc-client-secret","oidc_issuer":"https://accounts.google.com","platform_domain":"knoe.dev","postgres_version":"16","project_id":"plenary-truck-485623-p7","region":"\u2713 us-west3","tls_mode":"managed","wal_bucket":"knoe-0-wal","workload_identity_sa":"cnpg-backup"}
PRODUCTION_CONFIG_YAML = kind: KnoeProductionConfig
metadata:
ecosystemId: 0
@ -212,7 +216,7 @@ spec:
provider: gcp
projectId: plenary-truck-485623-p7
region: "\u2713 us-west3"
clusterName: knoe-dev-0
clusterName: knoe-cnpg-0
vpcMode: managed
vpcName: null
subnetName: null

View File

@ -19,7 +19,7 @@ metadata:
name: knoe-db-daily
namespace: knoe-db-0
spec:
schedule: "0 2 * * *"
schedule: "0 0 2 * * *" # 6-field CNPG cron: sec min hour dom mon dow
backupOwnerReference: self
cluster:
name: knoe-db

View File

@ -29,7 +29,7 @@ spec:
postgresql:
parameters:
shared_buffers: 128MB
shared_buffers: 64MB # ~25% of 256Mi request; restore to 128MB when resources increase
pg_stat_statements.max: '10000'
pg_stat_statements.track: all
shared_preload_libraries:
@ -108,11 +108,11 @@ spec:
resources:
requests:
cpu: "500m"
memory: "512Mi"
cpu: "250m" # halved for e2-small quota-increase window; restore to 500m after increase
memory: "256Mi" # halved for e2-small quota-increase window; restore to 512Mi after increase
limits:
cpu: "2"
memory: "2Gi"
cpu: "1"
memory: "1Gi"
enableSuperuserAccess: true
@ -124,7 +124,7 @@ spec:
resources:
requests:
storage: 50Gi
storageClassName: premium-rwo # pd-ssd; 3×50Gi PGDATA + 3×50Gi WAL = 300Gi total (fits 300GB quota)
storageClassName: premium-rwo # pd-ssd; 3×50Gi PGDATA + 3×50Gi WAL = 300Gi total (fits 300GB SSD quota)
walStorage:
size: 50Gi
@ -134,7 +134,7 @@ spec:
resources:
requests:
storage: 50Gi
storageClassName: premium-rwo
storageClassName: premium-rwo # pd-ssd; restore to premium-rwo after quota increase (matches PGDATA above)
monitoring:
enablePodMonitor: true # GKE managed Prometheus
enablePodMonitor: true # kube-prometheus-stack (Prometheus Operator) is installed before CNPG — PodMonitor CRD is present

View File

@ -0,0 +1,36 @@
# garage-service-ilb.yaml
# Internal LoadBalancer for Garage S3 API — exposes port 3900 to the same GCP VPC.
#
# Purpose: Allow knoe-cnpg-0 (DB cluster) to reach Garage running in knoe-dev-0 (app cluster)
# for Barman Cloud backups, without requiring a service mesh or cross-cluster DNS.
#
# Both clusters share the same VPC (us-west3); an ILB IP is routable between them.
#
# Apply to the app cluster (knoe-dev-0):
# kubectl --context=gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 \
# -n knoe-system apply -f k8s/prole/garage-service-ilb.yaml
#
# Then retrieve the ILB IP with:
# kubectl --context=gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 \
# -n knoe-system get svc garage-s3-ilb \
# -o jsonpath='{.status.loadBalancer.ingress[0].ip}'
apiVersion: v1
kind: Service
metadata:
name: garage-s3-ilb
labels:
app: garage
role: cross-cluster-s3
annotations:
# GCP Internal LoadBalancer — stays within the VPC, no public IP
networking.gke.io/load-balancer-type: "Internal"
spec:
type: LoadBalancer
loadBalancerIP: "" # GCP assigns automatically; pin here after first apply
selector:
app: garage
ports:
- name: s3
port: 3900
targetPort: s3
protocol: TCP

View File

@ -0,0 +1,19 @@
apiVersion: barmancloud.cnpg.io/v1
kind: ObjectStore
metadata:
name: knoe-db-barman-objectstore
# Applied to knoe-cnpg-0 in namespace knoe-db-0.
# Auth: Workload Identity — cnpg-backup-sa (in knoe-db-0) is annotated with
# iam.gke.io/gcp-service-account: cnpg-backup@<project>.iam.gserviceaccount.com
# No static credentials secret required.
spec:
retentionPolicy: 30d
configuration:
destinationPath: gs://knoe-0-backups/
# No googleCredentials block — barman-cloud uses the pod's Workload Identity
# token automatically when the ServiceAccount (cnpg-backup-sa in knoe-db-0)
# carries the iam.gke.io/gcp-service-account annotation set by init_cnpg_gke.sh.
wal:
compression: gzip
data:
compression: gzip

View File

@ -116,11 +116,13 @@ from typing import Callable, Sequence
DEFAULT_APP_CLUSTER_NAME = "knoe-dev-0"
DEFAULT_APP_CLUSTER_MODE = "autopilot"
DEFAULT_APP_CLUSTER_MODE = "standard"
DEFAULT_APP_CLUSTER_MACHINE_TYPE = "e2-small"
DEFAULT_APP_CLUSTER_NODE_COUNT = 3
DEFAULT_DB_CLUSTER_NAME = "knoe-cnpg-0"
DEFAULT_DB_CLUSTER_MODE = "standard"
DEFAULT_DB_CLUSTER_NODE_COUNT = 3
DEFAULT_DB_CLUSTER_MACHINE_TYPE = "e2-standard-2"
DEFAULT_DB_CLUSTER_MACHINE_TYPE = "e2-small"
DEFAULT_DB_BOOT_DISK_TYPE = "pd-standard"
DEFAULT_DB_BOOT_DISK_SIZE_GB = 50
@ -3788,6 +3790,8 @@ class KnoeConsoleInstaller(KnoeInstaller):
inputs["init_cluster.supabase_studio_url"] = "db.knoe.org"
inputs["init_cluster.app_cluster_name"] = DEFAULT_APP_CLUSTER_NAME
inputs["init_cluster.app_cluster_mode"] = DEFAULT_APP_CLUSTER_MODE
inputs["init_cluster.app_cluster_machine_type"] = DEFAULT_APP_CLUSTER_MACHINE_TYPE
inputs["init_cluster.app_cluster_node_count"] = str(DEFAULT_APP_CLUSTER_NODE_COUNT)
inputs["init_cluster.db_cluster_name"] = DEFAULT_DB_CLUSTER_NAME
inputs["init_cluster.db_cluster_mode"] = DEFAULT_DB_CLUSTER_MODE
inputs["init_cluster.db_cluster_node_count"] = str(DEFAULT_DB_CLUSTER_NODE_COUNT)
@ -5285,6 +5289,14 @@ class KnoeConsoleInstaller(KnoeInstaller):
app_cluster_name = self._app_cluster_name()
app_cluster_mode = self._app_cluster_mode()
app_machine_type = (
self._get_input("init_cluster.app_cluster_machine_type", "")
or DEFAULT_APP_CLUSTER_MACHINE_TYPE
).strip() or DEFAULT_APP_CLUSTER_MACHINE_TYPE
app_node_count = int(
(self._get_input("init_cluster.app_cluster_node_count", "") or str(DEFAULT_APP_CLUSTER_NODE_COUNT)).strip()
or str(DEFAULT_APP_CLUSTER_NODE_COUNT)
)
db_cluster_name = self._cnpg_cluster_name()
db_cluster_mode = self._db_cluster_mode()
db_node_count = int(
@ -5324,6 +5336,10 @@ class KnoeConsoleInstaller(KnoeInstaller):
name=app_cluster_name,
mode=app_cluster_mode,
location=str(app_location).strip(),
machine_type=app_machine_type,
node_count=max(1, app_node_count),
boot_disk_type=DEFAULT_DB_BOOT_DISK_TYPE,
boot_disk_size_gb=DEFAULT_DB_BOOT_DISK_SIZE_GB,
)
db_spec = GkeClusterSpec(
name=db_cluster_name,

View File

@ -54,13 +54,13 @@ def _kubectl(
timeout: int = 300,
check: bool = False,
) -> subprocess.CompletedProcess:
return _run(
["kubectl", *args],
env=env,
input_text=input_text,
timeout=timeout,
check=check,
)
# KUBECONTEXT is not a real kubectl env var — extract it and pass as --context flag.
ctx = str((env or {}).get("KUBECONTEXT") or "").strip()
cmd = ["kubectl"]
if ctx:
cmd.extend(["--context", ctx])
cmd.extend(args)
return _run(cmd, env=env, input_text=input_text, timeout=timeout, check=check)
def _helm(
@ -70,7 +70,13 @@ def _helm(
timeout: int = 300,
check: bool = False,
) -> subprocess.CompletedProcess:
return _run(["helm", *args], env=env, timeout=timeout, check=check)
# KUBECONTEXT is not a real helm env var — extract it and pass as --kube-context flag.
ctx = str((env or {}).get("KUBECONTEXT") or "").strip()
cmd = ["helm"]
if ctx:
cmd.extend(["--kube-context", ctx])
cmd.extend(args)
return _run(cmd, env=env, timeout=timeout, check=check)
def _k3d(

View File

@ -68,8 +68,14 @@ def _kubectl(
timeout: int = 30,
check: bool = False,
) -> subprocess.CompletedProcess:
# KUBECONTEXT is not a real kubectl env var — extract it and pass as --context flag.
ctx = str((env or {}).get("KUBECONTEXT") or "").strip()
cmd = ["kubectl"]
if ctx:
cmd.extend(["--context", ctx])
cmd.extend(args)
return subprocess.run(
["kubectl"] + args,
cmd,
env=env,
capture_output=True,
text=True,

View File

@ -98,24 +98,53 @@ def ensure_app_cluster(
if result.returncode == 0 and (result.stdout or "").strip() == spec.name:
return
# Cluster not found — create it as Autopilot.
# NOTE: gcloud container clusters create-auto does NOT support --disk-type
# or --disk-size. Autopilot manages all node infrastructure automatically.
# Autopilot only provisions physical nodes when pods are scheduled, so SSD
# quota is not consumed while the cluster is idle.
create_cmd = [
"gcloud",
"container",
"clusters",
"create-auto",
spec.name,
"--project",
project_id,
"--region",
spec.location,
"--workload-policies=allow-net-admin",
"--quiet",
]
# Cluster not found — create it.
if spec.mode.lower() == "autopilot":
# NOTE: gcloud container clusters create-auto does NOT support --disk-type
# or --disk-size. Autopilot manages all node infrastructure automatically.
# Autopilot only provisions physical nodes when pods are scheduled, so SSD
# quota is not consumed while the cluster is idle.
create_cmd = [
"gcloud",
"container",
"clusters",
"create-auto",
spec.name,
"--project",
project_id,
"--region",
spec.location,
"--workload-policies=allow-net-admin",
"--quiet",
]
else:
# Standard mode: explicit machine type, pd-standard boot disks (no SSD quota),
# VPA for dynamic resource adjustment, Workload Identity for GCS access.
create_cmd = [
"gcloud",
"container",
"clusters",
"create",
spec.name,
"--project",
project_id,
"--region",
spec.location,
"--cluster-version=latest",
"--machine-type",
spec.machine_type,
"--disk-type",
spec.boot_disk_type,
"--disk-size",
str(spec.boot_disk_size_gb),
"--num-nodes",
str(max(1, int(spec.node_count))),
"--enable-vertical-pod-autoscaling",
"--enable-ip-alias",
"--workload-pool",
f"{project_id}.svc.id.goog",
"--quiet",
]
created = _run(create_cmd, log=log)
if created.returncode != 0:
stderr = (created.stderr or "").strip()

View File

@ -38,41 +38,53 @@ def start(
update(namespace=namespace, env=env, log=log, mode=mode)
def update(
*,
namespace: str | None = None,
env: dict | None = None,
log: _LogFn | None = None,
mode: str | None = None,
) -> None:
_detect_mode(mode, env) # mode retained for parity with other owners
ns = _monitoring_namespace(namespace, env)
release = _release_name(env)
chart = str((env or {}).get("MONITORING_CHART") or "prometheus-community/kube-prometheus-stack")
grafana_password = str((env or {}).get("GRAFANA_ADMIN_PASSWORD") or "prole")
def _values_yaml_k8s(grafana_password: str) -> str:
"""Helm values for GKE / cloud-managed Kubernetes (no local PV affinity)."""
return (
"prometheus:\n"
" prometheusSpec:\n"
" storageSpec:\n"
" volumeClaimTemplate:\n"
" spec:\n"
" storageClassName: standard\n"
" accessModes: [ReadWriteOnce]\n"
" resources:\n"
" requests:\n"
" storage: 10Gi\n"
"alertmanager:\n"
" alertmanagerSpec:\n"
" storage:\n"
" volumeClaimTemplate:\n"
" spec:\n"
" storageClassName: standard\n"
" accessModes: [ReadWriteOnce]\n"
" resources:\n"
" requests:\n"
" storage: 5Gi\n"
"grafana:\n"
" adminUser: admin\n"
f" adminPassword: {grafana_password}\n"
" persistence:\n"
" type: sts\n"
" enabled: true\n"
" storageClassName: standard\n"
" accessModes: [ReadWriteOnce]\n"
" size: 5Gi\n"
)
# Derive storage class names from PROLE_MONITORING_DATA_DIR (same convention as init_monitoring.sh)
def _values_yaml_k3s(grafana_password: str, env: dict | None) -> str:
"""Helm values for k3s homelab (local iSCSI PVs pinned to merlin.prole.org)."""
data_dir = str((env or {}).get("PROLE_MONITORING_DATA_DIR") or "/synology/d004").rstrip("/")
volume_id = os.path.basename(data_dir) # e.g. "d004"
sc_prom = f"merlin-local-iscsi-{volume_id}-prometheus"
sc_alert = f"merlin-local-iscsi-{volume_id}-alertmanager"
sc_grafana = f"merlin-local-iscsi-{volume_id}-grafana"
# Primary monitoring node: prometheus and alertmanager must schedule here to bind local PVs
monitoring_node = str((env or {}).get("MONITORING_PRIMARY_NODE") or "merlin.prole.org")
_log(log, "[MONITORING] Ensuring helm repos")
_helm(["repo", "add", "prometheus-community", "https://prometheus-community.github.io/helm-charts"], env=env)
_helm(["repo", "update"], env=env)
_kubectl(["create", "namespace", ns], env=env, timeout=60)
# Nodes with broken kubelet (e.g. pi.prole.org returning 502) must be excluded
# from the node-exporter DaemonSet so helm --wait can succeed.
excluded_nodes = str((env or {}).get("MONITORING_NODE_EXPORTER_EXCLUDE_NODES") or "pi.prole.org")
excluded_list = [n.strip() for n in excluded_nodes.split(",") if n.strip()]
# Node affinity block for components that must land on the monitoring node (local PV binding)
node_affinity_yaml = (
" affinity:\n"
" nodeAffinity:\n"
@ -85,10 +97,9 @@ def update(
f" - {monitoring_node}\n"
)
# Node-exporter: exclude broken nodes
values_yaml = "prometheus-node-exporter:\n"
values = "prometheus-node-exporter:\n"
if excluded_list:
values_yaml += (
values += (
" affinity:\n"
" nodeAffinity:\n"
" requiredDuringSchedulingIgnoredDuringExecution:\n"
@ -99,10 +110,9 @@ def update(
" values:\n"
)
for node in excluded_list:
values_yaml += f" - {node}\n"
values += f" - {node}\n"
# Prometheus: pin to monitoring node + persistent storage
values_yaml += (
values += (
"prometheus:\n"
" prometheusSpec:\n"
+ node_affinity_yaml
@ -115,9 +125,7 @@ def update(
" requests:\n"
" storage: 30Gi\n"
)
# Alertmanager: pin to monitoring node + persistent storage
values_yaml += (
values += (
"alertmanager:\n"
" alertmanagerSpec:\n"
+ node_affinity_yaml
@ -130,9 +138,7 @@ def update(
" requests:\n"
" storage: 5Gi\n"
)
# Grafana: persistent storage (Deployment can float; local PV affinity will pull it to merlin)
values_yaml += (
values += (
"grafana:\n"
" adminUser: admin\n"
f" adminPassword: {grafana_password}\n"
@ -143,6 +149,33 @@ def update(
" accessModes: [ReadWriteOnce]\n"
" size: 10Gi\n"
)
return values
def update(
*,
namespace: str | None = None,
env: dict | None = None,
log: _LogFn | None = None,
mode: str | None = None,
) -> None:
effective_mode = _detect_mode(mode, env)
ns = _monitoring_namespace(namespace, env)
release = _release_name(env)
chart = str((env or {}).get("MONITORING_CHART") or "prometheus-community/kube-prometheus-stack")
grafana_password = str((env or {}).get("GRAFANA_ADMIN_PASSWORD") or "prole")
_log(log, "[MONITORING] Ensuring helm repos")
_helm(["repo", "add", "prometheus-community", "https://prometheus-community.github.io/helm-charts"], env=env)
_helm(["repo", "update"], env=env)
_kubectl(["create", "namespace", ns], env=env, timeout=60)
if effective_mode == "k8s":
values_yaml = _values_yaml_k8s(grafana_password)
else:
# k3s / k3d: homelab local-PV setup
values_yaml = _values_yaml_k3s(grafana_password, env)
tmp_values = tempfile.NamedTemporaryFile(
mode="w", suffix=".yaml", prefix="monitoring-values-", delete=False
@ -169,6 +202,27 @@ def update(
timeout=600,
check=True,
)
# Cross-cluster: install Prometheus Operator CRDs on the DB cluster so that
# CNPG (running on knoe-cnpg-0) can create PodMonitor resources without
# the reconciler looping on "PodMonitor CRD not present".
db_ctx = str((env or {}).get("DB_CLUSTER_KUBECONTEXT") or "").strip()
if db_ctx and effective_mode == "k8s":
_log(log, f"[MONITORING] Installing Prometheus Operator CRDs on DB cluster ({db_ctx})")
# helm show crds is client-side — clear KUBECONTEXT so no --kube-context flag
crd_env = {**(env or {}), "KUBECONTEXT": ""}
crd_res = _helm(["show", "crds", chart], env=crd_env, timeout=60)
if crd_res.returncode == 0 and crd_res.stdout.strip():
db_env = {**(env or {}), "KUBECONTEXT": db_ctx}
_kubectl(
["apply", "--server-side", "-f", "-"],
env=db_env,
input_text=crd_res.stdout,
timeout=120,
)
_log(log, "[MONITORING] Prometheus Operator CRDs installed on DB cluster.")
else:
_log(log, "[MONITORING] WARNING: could not fetch CRDs from chart — skipping DB cluster CRD install.")
finally:
os.unlink(tmp_values.name)

View File

@ -132,6 +132,8 @@ class KnoeInstaller(
"init_cluster.argocd_enabled": "argocd_enabled",
"init_cluster.at_rest_encryption_enabled": "at_rest_encryption_enabled",
"init_cluster.selected_kubectx": "selected_kubectx",
"init_cluster.app_cluster_kubecontext": "app_cluster_kubecontext",
"init_cluster.db_cluster_kubecontext": "db_cluster_kubecontext",
"init_cluster.k3s_server_url": "k3s_server_url",
"init_cluster.k3s_token": "k3s_token",
"ollama_config.server_host": "ollama_server_host",
@ -689,6 +691,8 @@ class KnoeInstaller(
)
self.kubectx_list = tk.Variable(value=self._get_kubectx_list())
self.selected_kubectx = tk.StringVar()
self.app_cluster_kubecontext = tk.StringVar()
self.db_cluster_kubecontext = tk.StringVar()
try:
# Prefer explicit prole.cfg selection; fall back to the dev-cluster section.
saved_ctx = (
@ -700,6 +704,14 @@ class KnoeInstaller(
saved_ctx = (saved_ctx or "").strip()
if saved_ctx:
self.selected_kubectx.set(saved_ctx)
# Pre-populate dual App/DB context selectors from [Inputs]
inputs = self.prole_cfg_data.get("Inputs", {}) or {}
saved_app = str(inputs.get("init_cluster.app_cluster_kubecontext") or "").strip()
saved_db = str(inputs.get("init_cluster.db_cluster_kubecontext") or "").strip()
if saved_app:
self.app_cluster_kubecontext.set(saved_app)
if saved_db:
self.db_cluster_kubecontext.set(saved_db)
except Exception:
pass
self.prod_artifacts_path = tk.StringVar(

View File

@ -197,82 +197,163 @@ class ClusterScreenMixin:
else:
# Service / Prod Section
# Context selection for both Service and Prod (managed via kubectx)
self._canvas_items.append(
values = self._get_kubectx_list()
if selected_env_key == "prod":
# ---- k8s / prod mode: explicit App + DB cluster dropdowns ----
self._canvas_items.append(
ui.canvas_text(
self,
x_label,
y,
"App Cluster:",
fill="black",
font=("SF Pro Text", 12, "bold"),
)
)
app_combo = ttk.Combobox(
self.bg_canvas,
textvariable=self.app_cluster_kubecontext,
values=values,
state="readonly",
width=40,
)
app_combo.bind("<<ComboboxSelected>>", self._on_dual_kubectx_select)
if not self.app_cluster_kubecontext.get() and values:
self.app_cluster_kubecontext.set(values[0])
app_combo_win = self.bg_canvas.create_window(
x_label + 140, y - 6, window=app_combo, anchor="nw"
)
self._canvas_items.append(app_combo_win)
self._overlay_widgets.append(app_combo)
self._app_kubectx_combo = app_combo
y += 34
self._canvas_items.append(
ui.canvas_text(
self,
x_label,
y,
"DB Cluster:",
fill="black",
font=("SF Pro Text", 12, "bold"),
)
)
db_combo = ttk.Combobox(
self.bg_canvas,
textvariable=self.db_cluster_kubecontext,
values=values,
state="readonly",
width=40,
)
db_combo.bind("<<ComboboxSelected>>", self._on_dual_kubectx_select)
if not self.db_cluster_kubecontext.get() and values:
self.db_cluster_kubecontext.set(values[0])
db_combo_win = self.bg_canvas.create_window(
x_label + 140, y - 6, window=db_combo, anchor="nw"
)
self._canvas_items.append(db_combo_win)
self._overlay_widgets.append(db_combo)
self._db_kubectx_combo = db_combo
dual_apply_btn = tk.Button(
self.bg_canvas,
text="Apply",
command=self._on_kubectx_apply_dual,
bg="#F5F5DC",
fg="black",
activebackground="#E5E5D5",
highlightbackground="#F5F5DC",
highlightthickness=0,
relief="flat",
font=("SF Pro Text", 10),
padx=10,
)
dual_apply_win = self.bg_canvas.create_window(
x_label + 450, y - 10, window=dual_apply_btn, anchor="nw"
)
self._canvas_items.append(dual_apply_win)
self._overlay_widgets.append(dual_apply_btn)
self._kubectx_apply_btn = dual_apply_btn
self._kubectx_apply_btn_canvas_window = dual_apply_win
y += 34
else:
# ---- Service / k3s mode: single kubectx combobox ----
self._canvas_items.append(
ui.canvas_text(
self,
x_label,
y,
"Kubernetes Context:",
fill="black",
font=("SF Pro Text", 12, "bold"),
)
)
combo = ttk.Combobox(
self.bg_canvas,
textvariable=self.selected_kubectx,
values=values,
state="readonly",
width=40,
)
combo.bind("<<ComboboxSelected>>", self._on_kubectx_select)
if not self.selected_kubectx.get() and values:
# Default to the current context if it exists in the list; otherwise pick
# a sensible default without switching immediately.
current_ctx = (getattr(self, "_kubectx_applied", "") or "").strip()
if not current_ctx:
current_ctx = (self._current_kubectl_context() or "").strip()
if current_ctx:
self._kubectx_applied = current_ctx
if current_ctx and current_ctx in values:
self.selected_kubectx.set(current_ctx)
elif "prole-k3s" in values:
self.selected_kubectx.set("prole-k3s")
else:
self.selected_kubectx.set(values[0])
combo_win = self.bg_canvas.create_window(
x_label + 270, y - 6, window=combo, anchor="nw"
)
self._canvas_items.append(combo_win)
self._overlay_widgets.append(combo)
# Keep references for UI tests / layout verification.
self._kubectx_combo = combo
self._kubectx_combo_canvas_window = combo_win
ui.canvas_text(
self,
x_label,
y,
"Kubernetes Context:",
fill="black",
font=("SF Pro Text", 12, "bold"),
x_label + 580,
y + 2,
"via kubectx",
fill="#6e6e73",
font=("SF Pro Text", 10),
)
)
values = self._get_kubectx_list()
combo = ttk.Combobox(
self.bg_canvas,
textvariable=self.selected_kubectx,
values=values,
state="readonly",
width=40,
)
combo.bind("<<ComboboxSelected>>", self._on_kubectx_select)
if not self.selected_kubectx.get() and values:
# Default to the current context if it exists in the list; otherwise pick
# a sensible default without switching immediately.
current_ctx = (getattr(self, "_kubectx_applied", "") or "").strip()
if not current_ctx:
current_ctx = (self._current_kubectl_context() or "").strip()
if current_ctx:
self._kubectx_applied = current_ctx
if current_ctx and current_ctx in values:
self.selected_kubectx.set(current_ctx)
elif selected_env_key == "service" and "prole-k3s" in values:
self.selected_kubectx.set("prole-k3s")
else:
self.selected_kubectx.set(values[0])
combo_win = self.bg_canvas.create_window(
x_label + 270, y - 6, window=combo, anchor="nw"
)
self._canvas_items.append(combo_win)
self._overlay_widgets.append(combo)
# Keep references for UI tests / layout verification.
self._kubectx_combo = combo
self._kubectx_combo_canvas_window = combo_win
ui.canvas_text(
self,
x_label + 580,
y + 2,
"via kubectx",
fill="#6e6e73",
font=("SF Pro Text", 10),
)
apply_btn = tk.Button(
self.bg_canvas,
text="Apply",
command=self._on_kubectx_apply,
bg="#F5F5DC",
fg="black",
activebackground="#E5E5D5",
highlightbackground="#F5F5DC",
highlightthickness=0,
relief="flat",
font=("SF Pro Text", 10),
padx=10,
state="disabled",
)
apply_win = self.bg_canvas.create_window(
x_label + 670, y - 10, window=apply_btn, anchor="nw"
)
self._canvas_items.append(apply_win)
self._overlay_widgets.append(apply_btn)
# Keep references for UI tests / layout verification.
self._kubectx_apply_btn = apply_btn
self._kubectx_apply_btn_canvas_window = apply_win
self._update_kubectx_apply_button()
y += 34
apply_btn = tk.Button(
self.bg_canvas,
text="Apply",
command=self._on_kubectx_apply,
bg="#F5F5DC",
fg="black",
activebackground="#E5E5D5",
highlightbackground="#F5F5DC",
highlightthickness=0,
relief="flat",
font=("SF Pro Text", 10),
padx=10,
state="disabled",
)
apply_win = self.bg_canvas.create_window(
x_label + 670, y - 10, window=apply_btn, anchor="nw"
)
self._canvas_items.append(apply_win)
self._overlay_widgets.append(apply_btn)
# Keep references for UI tests / layout verification.
self._kubectx_apply_btn = apply_btn
self._kubectx_apply_btn_canvas_window = apply_win
self._update_kubectx_apply_button()
y += 34
if selected_env_key == "prod":
# Staging directory input
@ -1888,6 +1969,51 @@ class ClusterScreenMixin:
self._verify_k3s_services()
self.show_page("init_cluster")
def _on_dual_kubectx_select(self, *args):
"""Called when App or DB cluster context is selected in prod/k8s mode."""
app_ctx = (self.app_cluster_kubecontext.get() or "").strip()
db_ctx = (self.db_cluster_kubecontext.get() or "").strip()
if app_ctx and db_ctx:
self._set_cluster_env_message(
f"App: {app_ctx} / DB: {db_ctx} — click Apply to save",
"#ff9500",
clear_after_ms=5000,
)
def _on_kubectx_apply_dual(self):
"""Save explicit App + DB cluster contexts to prole.cfg [Inputs] and [Global]."""
app_ctx = (self.app_cluster_kubecontext.get() or "").strip()
db_ctx = (self.db_cluster_kubecontext.get() or "").strip()
if not app_ctx or not db_ctx:
self._set_cluster_env_message(
"Both App Cluster and DB Cluster must be selected.", "#ff3b30"
)
return
# Persist to [Inputs] so they survive installer regeneration.
inputs = self.prole_cfg_data.setdefault("Inputs", {})
inputs["init_cluster.app_cluster_kubecontext"] = app_ctx
inputs["init_cluster.db_cluster_kubecontext"] = db_ctx
inputs["env_setup.APP_CLUSTER_KUBECONTEXT"] = app_ctx
inputs["env_setup.DB_CLUSTER_KUBECONTEXT"] = db_ctx
# Global.KUBECONTEXT = APP cluster — common services (garage, openbao, kong,
# monitoring) live there. DB cluster is always accessed via DB_CLUSTER_KUBECONTEXT.
glob = self.prole_cfg_data.setdefault("Global", {})
glob["KUBECONTEXT"] = app_ctx
glob["APP_CLUSTER_KUBECONTEXT"] = app_ctx
glob["DB_CLUSTER_KUBECONTEXT"] = db_ctx
self._save_prole_cfg()
# Switch ambient kubectl context to the app cluster (matches Global.KUBECONTEXT).
self._switch_kubectx(app_ctx)
self._set_cluster_env_message(
f"Saved — App: {app_ctx} / DB: {db_ctx}", "#34c759", clear_after_ms=4000
)
self.show_page("init_cluster")
def _update_kubectx_apply_button(self):
btn = getattr(self, "_kubectx_apply_btn", None)
if not btn:
@ -3651,7 +3777,7 @@ class ClusterScreenMixin:
collected_lines = []
try:
ns = self._get_service_namespace()
env = self._script_env_for_namespace(ns)
env = self._script_env_for_namespace(ns, cluster_role="app")
env["PROLE_MODE"] = self._deployment_mode()
if not env.get("KUBECONFIG"):
@ -3762,7 +3888,7 @@ class ClusterScreenMixin:
pass
service_ns = self._get_service_namespace()
env = self._script_env_for_namespace(service_ns)
env = self._script_env_for_namespace(service_ns, cluster_role="app")
env["PROLE_MODE"] = self._deployment_mode()
log_path = self._common_services_log_path()
env["COMMON_SERVICES_INIT_LOG"] = str(log_path)

View File

@ -62,6 +62,9 @@ class DatabaseScreenMixin:
x_field = 300
y = 270
# Database operations should target the DB cluster context by default.
self._ensure_db_screen_context()
# Database namespace
self._canvas_items.append(
ui.canvas_text(
@ -541,6 +544,115 @@ class DatabaseScreenMixin:
# Allow only valid k8s characters while typing
return re.match(r"^[a-z0-9-]*$", proposed) is not None
def _resolve_db_kube_context(self) -> str:
cfg = getattr(self, "prole_cfg_data", {}) or {}
glob = cfg.get("Global", {}) or {}
gcp = cfg.get("GCP", {}) or {}
for value in (
self._get_input("init_cluster.db_cluster_kubecontext", ""),
self._get_input("env_setup.DB_CLUSTER_KUBECONTEXT", ""),
glob.get("DB_CLUSTER_KUBECONTEXT", ""),
):
ctx = _safe_str(value)
if ctx:
return ctx
cluster_name = (
self._get_input("init_password.db_cluster_name", "")
or self._get_input("env_setup.DB_CLUSTER_NAME", "")
or glob.get("DB_CLUSTER_NAME", "")
or glob.get("CNPG_CLUSTER_NAME", "")
or "knoe-cnpg-0"
)
cluster_name = _safe_str(cluster_name)
selected_ctx = _safe_str(
self._get_input("init_cluster.selected_kubectx", "")
or glob.get("KUBECONTEXT", "")
)
project_id = _safe_str(gcp.get("project_id") or gcp.get("PROJECT_ID") or "")
region = _safe_str(
self._get_input("init_cluster.db_cluster_region", "")
or gcp.get("region")
or gcp.get("REGION")
or gcp.get("location")
or ""
)
if selected_ctx.startswith("gke_"):
parts = selected_ctx.split("_", 3)
if len(parts) == 4:
if not project_id:
project_id = _safe_str(parts[1])
if not region:
region = _safe_str(parts[2])
if project_id and region and cluster_name:
return f"gke_{project_id}_{region}_{cluster_name}"
if selected_ctx.startswith("gke_") and cluster_name:
parts = selected_ctx.split("_", 3)
if len(parts) == 4:
return f"{parts[0]}_{parts[1]}_{parts[2]}_{cluster_name}"
return selected_ctx
def _db_kubectl_base_cmd(self) -> list[str]:
mode = self._deployment_mode()
base_cmd = self._kubectl_base_cmd(mode=mode)
if mode != "k8s":
return base_cmd
target_ctx = self._resolve_db_kube_context()
if not target_ctx:
return base_cmd
cmd = list(base_cmd)
if "--context" in cmd:
try:
idx = cmd.index("--context")
if idx + 1 < len(cmd):
cmd[idx + 1] = target_ctx
return cmd
except Exception:
pass
cmd.extend(["--context", target_ctx])
return cmd
def _ensure_db_screen_context(self):
target_ctx = self._resolve_db_kube_context()
if not target_ctx:
return
selected_var = getattr(self, "selected_kubectx", None)
if not selected_var or not hasattr(selected_var, "get") or not hasattr(selected_var, "set"):
return
current_ctx = _safe_str(selected_var.get())
app_default_ctx = _safe_str((getattr(self, "prole_cfg_data", {}) or {}).get("Global", {}).get("KUBECONTEXT", ""))
# Keep explicit user context overrides; auto-switch only from blank/default app context.
if current_ctx and current_ctx != app_default_ctx:
return
if current_ctx != target_ctx:
selected_var.set(target_ctx)
try:
inputs = getattr(self, "inputs", None)
if isinstance(inputs, dict):
inputs["init_cluster.selected_kubectx"] = target_ctx
except Exception:
pass
def _db_target_note(self, namespace: str, notice: str = "") -> str:
ctx = self._resolve_db_kube_context() or "(default kubectl context)"
ns = _safe_str(namespace) or "(not set)"
base = f"Target context: {ctx} | Target namespace: {ns}"
return f"{base} | {notice}" if notice else base
def _collect_namespace_rows(
self,
*,
@ -554,7 +666,7 @@ class DatabaseScreenMixin:
kubectl_ok = False
try:
base_cmd = self._kubectl_base_cmd()
base_cmd = self._db_kubectl_base_cmd()
result = subprocess.run(
base_cmd + ["get", "ns", "-o", "json"],
capture_output=True,
@ -598,7 +710,7 @@ class DatabaseScreenMixin:
pods_by_ns = {}
if kubectl_ok:
try:
base_cmd = self._kubectl_base_cmd()
base_cmd = self._db_kubectl_base_cmd()
pods_result = subprocess.run(
base_cmd + ["get", "pods", "--all-namespaces", "-o", "json"],
capture_output=True,
@ -713,8 +825,8 @@ class DatabaseScreenMixin:
note_item = getattr(self, "_db_namespace_note", None)
if note_item and self.bg_canvas.winfo_exists():
msg = notice if notice else ""
self.bg_canvas.itemconfig(note_item, text=msg)
msg = self._db_target_note(current_ns, notice)
self.bg_canvas.itemconfig(note_item, text=msg, fill="#6e6e73")
self.safe_after(update_ui)
finally:
@ -725,7 +837,13 @@ class DatabaseScreenMixin:
def _db_set_status(self, message: str, color: str = "#6e6e73"):
note_item = getattr(self, "_db_namespace_note", None)
if note_item and self.bg_canvas.winfo_exists():
self.bg_canvas.itemconfig(note_item, text=message or "", fill=color)
namespace_var = getattr(self, "db_namespace", None)
ns = _safe_str(namespace_var.get()) if hasattr(namespace_var, "get") else ""
self.bg_canvas.itemconfig(
note_item,
text=self._db_target_note(ns, message or ""),
fill=color,
)
def _db_set_buttons_state(self, state: str):
buttons = getattr(self, "_db_action_buttons", None)
@ -818,7 +936,7 @@ class DatabaseScreenMixin:
self.safe_after(
lambda: self._db_set_status(f"Creating namespace {ns}...", "#1d1d1f")
)
cmd = self._kubectl_base_cmd() + ["create", "namespace", ns]
cmd = self._db_kubectl_base_cmd() + ["create", "namespace", ns]
rc, out = self._run_cmd_capture(cmd)
if rc != 0 and "AlreadyExists" not in out:
self._db_log(out)
@ -856,7 +974,7 @@ class DatabaseScreenMixin:
self.safe_after(
lambda: self._db_set_status(f"Deleting namespace {ns}...", "#1d1d1f")
)
cmd = self._kubectl_base_cmd() + ["delete", "namespace", ns]
cmd = self._db_kubectl_base_cmd() + ["delete", "namespace", ns]
rc, out = self._run_cmd_capture(cmd)
if rc != 0:
self._db_log(out)

View File

@ -427,6 +427,31 @@ class ServicesScreenMixin:
if mode == "k8s":
for _k, _v in self._build_gke_registry_env().items():
env[_k] = _v
# Inject explicit cluster contexts so all kubectl/helm ops use the right cluster.
# Common services (garage, openbao, kong, monitoring) → APP cluster.
# CNPG + backup → DB cluster (cnpg_env below).
_app_ctx = (
self._cluster_kubecontext("app")
or (self.prole_cfg_data.get("Global", {}) or {}).get(
"APP_CLUSTER_KUBECONTEXT", ""
)
).strip()
_db_ctx = (
self._cluster_kubecontext("db")
or (self.prole_cfg_data.get("Global", {}) or {}).get(
"DB_CLUSTER_KUBECONTEXT", ""
)
).strip()
if _app_ctx:
env["APP_CLUSTER_KUBECONTEXT"] = _app_ctx
env["KUBECONTEXT"] = _app_ctx # default = app cluster
if _db_ctx:
env["DB_CLUSTER_KUBECONTEXT"] = _db_ctx
cnpg_env = dict(env)
if _db_ctx:
cnpg_env["KUBECONTEXT"] = _db_ctx
else:
cnpg_env = dict(env)
mode_args = ["--mode", mode] if mode else []
raw_min = (
@ -563,154 +588,8 @@ class ServicesScreenMixin:
"Skipping common services initialization because previous steps failed.\n"
)
# 2. CloudNative-PG initialization (Python)
cnpg_script = "init_cloudnative_pg.sh"
if overall_success:
_select_tab(cnpg_script)
_con(cnpg_script).clear()
_con(cnpg_script).write("Running CloudNative-PG initialization...\n")
log_path = _log_path_for(cnpg_script)
self._record_install_log(log_path)
try:
log_fp = log_path.open("a", encoding="utf-8")
except Exception:
log_fp = None
def _cnpg_log(msg):
line = msg if msg.endswith("\n") else msg + "\n"
_con("init_cloudnative_pg.sh").write(line)
self._process_script_output_line(line)
if log_fp:
try:
log_fp.write(line)
log_fp.flush()
except Exception:
pass
cnpg_ns = env.get("DATABASE_NAMESPACE") or env.get("NAMESPACE", "default")
cnpg_cluster = env.get("CLUSTER_NAME") or env.get("CNPG_CLUSTER_NAME") or "knoe-db"
cnpg_project_root = getattr(self.controller, "project_root", PROJECT_ROOT)
try:
cnpg_initialize(
namespace=cnpg_ns,
cluster_name=cnpg_cluster,
env=env,
project_root=cnpg_project_root,
log=_cnpg_log,
mode=mode,
)
_con(cnpg_script).write(
"\nCloudNative-PG initialization completed successfully.\n"
)
for _lk in ("knoe_db_1", "knoe_db_2", "knoe_db_3"):
self.safe_after(lambda k=_lk: self._set_init_light(k, "green"))
except Exception as _cnpg_exc:
_con(cnpg_script).write(
f"\nERROR: CloudNative-PG initialization failed: {_cnpg_exc}\n"
)
overall_success = False
for _lk in ("knoe_db_1", "knoe_db_2", "knoe_db_3"):
self.safe_after(lambda k=_lk: self._set_init_light(k, "red"))
finally:
if log_fp:
try:
log_fp.close()
except Exception:
pass
else:
_con("init_cloudnative_pg.sh").write(
"Skipping CloudNative-PG initialization because previous steps failed.\n"
)
# 4. init_cnpg_backup.sh start
if overall_success:
script = "init_cnpg_backup.sh"
_select_tab(script)
_con(script).clear()
_con(script).write(f"Running {script} start...\n")
log_path = _log_path_for(script)
self._record_install_log(log_path)
try:
log_fp = log_path.open("a", encoding="utf-8")
except Exception:
log_fp = None
def _backup_line(line):
_con("init_cnpg_backup.sh").write(line)
self._process_script_output_line(line)
if log_fp:
try:
log_fp.write(line)
log_fp.flush()
except Exception:
pass
rc_backup = self.controller.run_script(
script, args=mode_args + ["start"], env=env, on_line=_backup_line
)
if log_fp:
try:
log_fp.close()
except Exception:
pass
if rc_backup != 0:
_con(script).write(f"\nERROR: {script} start failed with code {rc_backup}\n")
overall_success = False
self.safe_after(lambda: self._set_init_light("barman_cloud", "red"))
else:
_con(script).write(f"\n{script} completed successfully.\n")
self.safe_after(lambda: self._set_init_light("barman_cloud", "green"))
else:
_con("init_cnpg_backup.sh").write(
f"Skipping {cnpg_cluster} backup because previous steps failed.\n"
)
# 5. init_kong.sh start
if overall_success:
script = "init_kong.sh"
_select_tab(script)
_con(script).clear()
_con(script).write(f"Running {script} start...\n")
log_path = _log_path_for(script)
self._record_install_log(log_path)
try:
log_fp = log_path.open("a", encoding="utf-8")
except Exception:
log_fp = None
def _kong_line(line):
_con("init_kong.sh").write(line)
self._process_script_output_line(line)
if log_fp:
try:
log_fp.write(line)
log_fp.flush()
except Exception:
pass
kong_ns = env.get("NAMESPACE") or "knoe-db"
rc_kong = self.controller.run_script(
script, args=mode_args + ["-n", kong_ns, "start"], env=env, on_line=_kong_line
)
if log_fp:
try:
log_fp.close()
except Exception:
pass
if rc_kong != 0:
_con(script).write(f"\nERROR: {script} start failed with code {rc_kong}\n")
overall_success = False
else:
_con("init_kong.sh").write(
"Skipping Kong because previous steps failed.\n"
)
# 6. init_monitoring.sh initialize
# 2. init_monitoring.sh initialize — runs before CNPG so PodMonitor CRD is present
# when the CNPG cluster is created (enablePodMonitor: true in knoe-db.yaml).
if overall_success and opt_allowed:
script = "init_monitoring.sh"
_select_tab(script)
@ -801,7 +680,154 @@ class ServicesScreenMixin:
"Skipping Monitoring initialization because previous steps failed.\n"
)
# 7. init_nginx_ingress.sh initialize (k3s only — not used in k3d)
# 3. CloudNative-PG initialization (Python)
cnpg_script = "init_cloudnative_pg.sh"
if overall_success:
_select_tab(cnpg_script)
_con(cnpg_script).clear()
_con(cnpg_script).write("Running CloudNative-PG initialization...\n")
log_path = _log_path_for(cnpg_script)
self._record_install_log(log_path)
try:
log_fp = log_path.open("a", encoding="utf-8")
except Exception:
log_fp = None
def _cnpg_log(msg):
line = msg if msg.endswith("\n") else msg + "\n"
_con("init_cloudnative_pg.sh").write(line)
self._process_script_output_line(line)
if log_fp:
try:
log_fp.write(line)
log_fp.flush()
except Exception:
pass
cnpg_ns = env.get("DATABASE_NAMESPACE") or env.get("NAMESPACE", "default")
cnpg_cluster = env.get("CLUSTER_NAME") or env.get("CNPG_CLUSTER_NAME") or "knoe-db"
cnpg_project_root = getattr(self.controller, "project_root", PROJECT_ROOT)
try:
cnpg_initialize(
namespace=cnpg_ns,
cluster_name=cnpg_cluster,
env=cnpg_env,
project_root=cnpg_project_root,
log=_cnpg_log,
mode=mode,
)
_con(cnpg_script).write(
"\nCloudNative-PG initialization completed successfully.\n"
)
for _lk in ("knoe_db_1", "knoe_db_2", "knoe_db_3"):
self.safe_after(lambda k=_lk: self._set_init_light(k, "green"))
except Exception as _cnpg_exc:
_con(cnpg_script).write(
f"\nERROR: CloudNative-PG initialization failed: {_cnpg_exc}\n"
)
overall_success = False
for _lk in ("knoe_db_1", "knoe_db_2", "knoe_db_3"):
self.safe_after(lambda k=_lk: self._set_init_light(k, "red"))
finally:
if log_fp:
try:
log_fp.close()
except Exception:
pass
else:
_con("init_cloudnative_pg.sh").write(
"Skipping CloudNative-PG initialization because previous steps failed.\n"
)
# 4. init_cnpg_backup.sh start
if overall_success:
script = "init_cnpg_backup.sh"
_select_tab(script)
_con(script).clear()
_con(script).write(f"Running {script} start...\n")
log_path = _log_path_for(script)
self._record_install_log(log_path)
try:
log_fp = log_path.open("a", encoding="utf-8")
except Exception:
log_fp = None
def _backup_line(line):
_con("init_cnpg_backup.sh").write(line)
self._process_script_output_line(line)
if log_fp:
try:
log_fp.write(line)
log_fp.flush()
except Exception:
pass
rc_backup = self.controller.run_script(
script, args=mode_args + ["start"], env=cnpg_env, on_line=_backup_line
)
if log_fp:
try:
log_fp.close()
except Exception:
pass
if rc_backup != 0:
_con(script).write(f"\nERROR: {script} start failed with code {rc_backup}\n")
overall_success = False
self.safe_after(lambda: self._set_init_light("barman_cloud", "red"))
else:
_con(script).write(f"\n{script} completed successfully.\n")
self.safe_after(lambda: self._set_init_light("barman_cloud", "green"))
else:
_con("init_cnpg_backup.sh").write(
f"Skipping {cnpg_cluster} backup because previous steps failed.\n"
)
# 5. init_kong.sh start
if overall_success:
script = "init_kong.sh"
_select_tab(script)
_con(script).clear()
_con(script).write(f"Running {script} start...\n")
log_path = _log_path_for(script)
self._record_install_log(log_path)
try:
log_fp = log_path.open("a", encoding="utf-8")
except Exception:
log_fp = None
def _kong_line(line):
_con("init_kong.sh").write(line)
self._process_script_output_line(line)
if log_fp:
try:
log_fp.write(line)
log_fp.flush()
except Exception:
pass
kong_ns = env.get("NAMESPACE") or "knoe-db"
rc_kong = self.controller.run_script(
script, args=mode_args + ["-n", kong_ns, "start"], env=env, on_line=_kong_line
)
if log_fp:
try:
log_fp.close()
except Exception:
pass
if rc_kong != 0:
_con(script).write(f"\nERROR: {script} start failed with code {rc_kong}\n")
overall_success = False
else:
_con("init_kong.sh").write(
"Skipping Kong because previous steps failed.\n"
)
# 6. init_nginx_ingress.sh initialize (k3s only — not used in k3d)
if overall_success and mode != "k3d":
script = "init_nginx_ingress.sh"
_select_tab(script)
@ -844,7 +870,7 @@ class ServicesScreenMixin:
"Skipping Nginx Ingress because previous steps failed.\n"
)
# 8. init_port_forwards.sh start (k3d only — not used in k3s)
# 7. init_port_forwards.sh start (k3d only — not used in k3s)
if overall_success and mode != "k3s":
script = "init_port_forwards.sh"
_select_tab(script)

View File

@ -1 +1 @@
40
48

View File

@ -1,55 +1,55 @@
Network Discovery Summary:
Primary Router: 10.0.0.1 (Unknown)
DNS Servers: 10.0.0.5, 10.0.0.4, 100.100.100.100
Primary Router: 10.0.0.1 (eero_5d:50:f2)
DNS Servers: 100.100.100.100, 10.0.0.5, 10.0.0.4
Detected Devices:
- 10.0.0.27 [3a:b4:ce:79:b0:55] (3a:b4:ce:79:b0:55): Ports [], Services: []
- 10.0.0.63 [24:fc:e5:51:cf:74] (SamsungElect_51:cf:74): Ports [], Services: []
- 10.0.0.111 [ca:0b:6c:5d:0a:e5] (ca:0b:6c:5d:0a:e5): Ports [], Services: []
- 10.0.0.49 [00:e0:4c:08:0c:75] (RealtekSemic_08:0c:75): Ports [22, 445, 5900, 11434, 88], Services: ['SSH', 'VNC', 'SMB/CIFS (Possible Windows/AD)', 'Ollama', 'Active Directory Related']
- 10.0.0.203 [00:11:32:3b:2f:08] (Synology_3b:2f:08): Ports [22, 80, 443, 2049, 445], Services: ['SSH', 'NFS', 'Web Server', 'SMB/CIFS (Possible Windows/AD)']
- 10.0.0.31 [2e:c1:c1:e9:87:b4] (2e:c1:c1:e9:87:b4): Ports [22, 445, 5900, 11434, 88], Services: ['SSH', 'VNC', 'SMB/CIFS (Possible Windows/AD)', 'Ollama', 'Active Directory Related']
- 10.0.0.189 [00:17:88:a3:2f:cc] (PhilipsLight_a3:2f:cc): Ports [80, 443], Services: ['Web Server']
- 10.0.0.31\ [2e:c1:c1:e9:87:b4] (2e:c1:c1:e9:87:b4): Ports [], Services: []
- 10.0.0.73 [4c:a9:19:b3:12:f8] (TuyaSmart_b3:12:f8): Ports [], Services: []
- 10.0.0.2 [dc:a4:ca:ea:1a:2f] (Apple_ea:1a:2f): Ports [], Services: []
- 10.0.0.1 [9c:57:bc:5d:50:f2] (eero_5d:50:f2): Ports [53], Services: ['DNS']
- 10.0.0.7 [2c:cf:67:ca:93:ce] (RaspberryPi_ca:93:ce): Ports [22], Services: ['SSH']
- 10.0.0.26 [c8:db:26:08:52:f8] (Logitech_08:52:f8): Ports [], Services: []
- 10.0.0.32 [48:a6:b8:a4:8e:cc] (Sonos_a4:8e:cc): Ports [], Services: []
- 10.0.0.33 [ec:b5:fa:b0:76:e4] (PhilipsLight_b0:76:e4): Ports [80, 443], Services: ['Web Server']
- 10.0.0.31\ [2e:c1:c1:e9:87:b4] (2e:c1:c1:e9:87:b4): Ports [], Services: []
- 10.0.0.6 [2c:cf:67:45:8d:97] (RaspberryPi_45:8d:97): Ports [22, 80, 443], Services: ['SSH', 'Web Server']
- 10.0.0.3 [2c:cf:67:8c:2b:47] (RaspberryPi_8c:2b:47): Ports [22, 53, 80, 443, 2049, 445, 5900, 88, 389, 636], Services: ['DNS', 'SSH', 'NFS', 'VNC', 'Web Server', 'SMB/CIFS (Possible Windows/AD)', 'Active Directory Related']
- 10.0.0.4 [b8:27:eb:b3:6f:2b] (RaspberryPiF_b3:6f:2b): Ports [22, 53, 443, 2049], Services: ['DNS', 'SSH', 'NFS', 'Web Server']
- 10.0.0.22 [2c:cf:67:45:8d:97] (RaspberryPi_45:8d:97): Ports [22, 80, 443], Services: ['SSH', 'Web Server']
- 10.0.0.5 [b8:27:eb:88:3a:41] (RaspberryPiF_88:3a:41): Ports [22, 2049, 5900], Services: ['SSH', 'NFS', 'VNC']
- 10.0.0.205 [a8:20:66:28:12:e7] (Apple_28:12:e7): Ports [22, 445, 5900, 88], Services: ['SSH', 'VNC', 'SMB/CIFS (Possible Windows/AD)', 'Active Directory Related']
- 10.0.0.37 [54:07:7d:22:c0:b2] (Netgear_22:c0:b2): Ports [80], Services: ['Web Server']
- 10.0.0.38 [24:fc:e5:51:cf:74] (SamsungElect_51:cf:74): Ports [], Services: []
- 10.0.0.45 [f8:b4:6a:30:a4:1f] (HewlettPacka_30:a4:1f): Ports [80, 443], Services: ['Web Server']
- 10.0.0.48 [0c:fe:45:53:f4:3c] (SonyInteract_53:f4:3c): Ports [], Services: []
- 10.0.0.52 [2c:cf:67:ca:93:ce] (RaspberryPi_ca:93:ce): Ports [22], Services: ['SSH']
- 10.0.0.36 [7e:46:74:2a:b0:b2] (7e:46:74:2a:b0:b2): Ports [], Services: []
- 10.0.0.46 [d4:f7:d5:40:ab:17] (SonyInteract_40:ab:17): Ports [], Services: []
- 10.0.0.107 [b8:06:0d:b7:7c:56] (TuyaSmart_b7:7c:56): Ports [], Services: []
- 10.0.0.73 [4c:a9:19:b3:12:f8] (TuyaSmart_b3:12:f8): Ports [], Services: []
- 10.0.0.111 [ca:0b:6c:5d:0a:e5] (ca:0b:6c:5d:0a:e5): Ports [], Services: []
- 10.0.0.128 [20:c9:d0:94:30:dd] (Apple_94:30:dd): Ports [22, 445, 5900, 88], Services: ['SSH', 'VNC', 'SMB/CIFS (Possible Windows/AD)', 'Active Directory Related']
- 10.0.0.95 [b8:27:eb:88:3a:41] (RaspberryPiF_88:3a:41): Ports [22, 2049, 5900], Services: ['SSH', 'NFS', 'VNC']
- 10.0.0.99 [20:28:bc:f1:dc:ba] (Visionscape_f1:dc:ba): Ports [], Services: []
- 10.0.0.106 [48:a6:b8:a7:50:60] (Sonos_a7:50:60): Ports [], Services: []
- 10.0.0.204 [50:eb:f6:56:f3:26] (ASUSTekCOMPU_56:f3:26): Ports [22, 3389, 445, 5900, 11434], Services: ['SSH', 'VNC', 'SMB/CIFS (Possible Windows/AD)', 'RDP (Windows)', 'Ollama']
- 10.0.0.1 [Unknown] (Unknown): Ports [53], Services: ['DNS']
- 10.0.0.2 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.3 [Unknown] (Unknown): Ports [22, 53, 80, 443, 2049, 445, 5900, 88, 389, 636], Services: ['DNS', 'SSH', 'NFS', 'VNC', 'Web Server', 'SMB/CIFS (Possible Windows/AD)', 'Active Directory Related']
- 10.0.0.4 [Unknown] (Unknown): Ports [22, 53, 443, 2049], Services: ['DNS', 'SSH', 'NFS', 'Web Server']
- 10.0.0.5 [Unknown] (Unknown): Ports [22, 2049, 5900], Services: ['SSH', 'NFS', 'VNC']
- 10.0.0.6 [Unknown] (Unknown): Ports [22, 80, 443], Services: ['SSH', 'Web Server']
- 10.0.0.7 [Unknown] (Unknown): Ports [22], Services: ['SSH']
- 10.0.0.22 [Unknown] (Unknown): Ports [22, 80, 443], Services: ['SSH', 'Web Server']
- 10.0.0.26 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.32 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.37 [Unknown] (Unknown): Ports [80], Services: ['Web Server']
- 10.0.0.45 [Unknown] (Unknown): Ports [80, 443], Services: ['Web Server']
- 10.0.0.46 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.48 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.52 [Unknown] (Unknown): Ports [22], Services: ['SSH']
- 10.0.0.62 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.95 [Unknown] (Unknown): Ports [22, 2049, 5900], Services: ['SSH', 'NFS', 'VNC']
- 10.0.0.99 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.106 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.112 [Unknown] (Unknown): Ports [80, 443], Services: ['Web Server']
- 10.0.0.117 [Unknown] (Unknown): Ports [80], Services: ['Web Server']
- 10.0.0.123 [Unknown] (Unknown): Ports [80], Services: ['Web Server']
- 10.0.0.124 [Unknown] (Unknown): Ports [80], Services: ['Web Server']
- 10.0.0.127 [Unknown] (Unknown): Ports [22], Services: ['SSH']
- 10.0.0.128 [Unknown] (Unknown): Ports [22, 445, 5900, 88], Services: ['SSH', 'VNC', 'SMB/CIFS (Possible Windows/AD)', 'Active Directory Related']
- 10.0.0.130 [Unknown] (Unknown): Ports [53], Services: ['DNS']
- 10.0.0.143 [Unknown] (Unknown): Ports [53], Services: ['DNS']
- 10.0.0.145 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.155 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.170 [Unknown] (Unknown): Ports [], Services: []
- 10.0.0.175 [Unknown] (Unknown): Ports [80], Services: ['Web Server']
- 10.0.0.188 [Unknown] (Unknown): Ports [80, 443], Services: ['Web Server']
- 10.0.0.196 [Unknown] (Unknown): Ports [53], Services: ['DNS']
- 10.0.0.203 [Unknown] (Unknown): Ports [22, 80, 443, 2049, 445], Services: ['SSH', 'NFS', 'Web Server', 'SMB/CIFS (Possible Windows/AD)']
- 10.0.0.205 [Unknown] (Unknown): Ports [22, 445, 5900, 88], Services: ['SSH', 'VNC', 'SMB/CIFS (Possible Windows/AD)', 'Active Directory Related']
- 10.0.0.206 [Unknown] (Unknown): Ports [80], Services: ['Web Server']
Ollama Instances found at: 10.0.0.49, 10.0.0.31, 10.0.0.204
- 10.0.0.123 [3c:ef:8c:96:e3:3c] (ZhejiangDahu_96:e3:3c): Ports [80], Services: ['Web Server']
- 10.0.0.117 [40:f5:20:1e:5e:91] (Espressif_1e:5e:91): Ports [80], Services: ['Web Server']
- 10.0.0.124 [9c:8e:cd:02:73:c4] (AmcrestTechn_02:73:c4): Ports [80], Services: ['Web Server']
- 10.0.0.113 [60:81:10:92:02:4e] (Apple_92:02:4e): Ports [], Services: []
- 10.0.0.125 [a0:60:32:04:73:ba] (AmcrestTechn_04:73:ba): Ports [80], Services: ['Web Server']
- 10.0.0.127 [58:55:ca:4a:60:6e] (Apple_4a:60:6e): Ports [22], Services: ['SSH']
- 10.0.0.130 [f8:bb:bf:a3:84:eb] (eero_a3:84:eb): Ports [53], Services: ['DNS']
- 10.0.0.143 [60:5f:8d:7a:aa:32] (eero_7a:aa:32): Ports [53], Services: ['DNS']
- 10.0.0.155 [62:17:97:b3:06:a1] (62:17:97:b3:06:a1): Ports [], Services: []
- 10.0.0.206 [00:a0:de:a2:0b:ef] (Yamaha_a2:0b:ef): Ports [80], Services: ['Web Server']
- 10.0.0.170 [b0:ee:7b:ca:73:99] (Roku_ca:73:99): Ports [], Services: []
- 10.0.0.175 [28:80:88:e4:10:1a] (Netgear_e4:10:1a): Ports [80], Services: ['Web Server']
- 10.0.0.189 [00:17:88:a3:2f:cc] (PhilipsLight_a3:2f:cc): Ports [80, 443], Services: ['Web Server']
- 10.0.0.188 [00:16:6c:c5:1f:54] (SamsungElect_c5:1f:54): Ports [80, 443], Services: ['Web Server']
- 10.0.0.196 [60:5f:8d:88:08:f2] (eero_88:08:f2): Ports [53], Services: ['DNS']
- 10.0.0.199 [7c:a6:b0:01:a7:33] (7c:a6:b0:01:a7:33): Ports [], Services: []
- 10.0.0.208 [a0:ad:9f:30:84:5f] (ASUSTekCOMPU_30:84:5f): Ports [], Services: []
Ollama Instances found at: 10.0.0.31, 10.0.0.204
Ansible Inventory Summary:
Inventory directory: /Users/chrisfu/dev/prole/infrastructure

View File

@ -0,0 +1,186 @@
#!/usr/bin/env bash
# patch_garage_cross_cluster.sh
#
# One-shot patch for the dual-cluster garage/barman situation:
#
# Problem: install.sh incorrectly deployed Garage into knoe-cnpg-0 (DB cluster)
# because app_cluster_kubecontext was not persisted in conf/prod/prole.cfg.
# CNPG backups should use GCS with Workload Identity, not Garage.
#
# Fix applied by this script:
# 1. Remove Garage from knoe-cnpg-0 (statefulset, service, configmap, PVC)
# 2. Apply cnpg-backup-sa ServiceAccount with Workload Identity annotation
# 3. Apply the GCS barman ObjectStore in knoe-db-0
# 4. Refresh CNPG_ELIGIBLE_NODES in conf/prod/prole.cfg from live knoe-cnpg-0 nodes
#
# Note: GCS buckets and GCP SA must already exist (created by init_cnpg_gke.sh or
# the GCP console). If not, run:
# ./etc/init_cnpg_gke.sh --project plenary-truck-485623-p7 --region us-west3 \
# --cluster knoe-cnpg-0
#
# Usage:
# CONFIRM=true ./scripts/patch_garage_cross_cluster.sh
#
# Overrides:
# GCP_PROJECT (default: plenary-truck-485623-p7)
# GCP_REGION (default: us-west3)
# DB_CLUSTER (default: knoe-cnpg-0)
# APP_CLUSTER (default: knoe-dev-0)
# SERVICE_NS (default: knoe-system)
# DB_NS (default: knoe-db-0)
# CONFIRM REQUIRED: must be "true"
set -euo pipefail
export PATH="/opt/homebrew/share/google-cloud-sdk/bin:$PATH"
GCP_PROJECT="${GCP_PROJECT:-plenary-truck-485623-p7}"
GCP_REGION="${GCP_REGION:-us-west3}"
DB_CLUSTER="${DB_CLUSTER:-knoe-cnpg-0}"
APP_CLUSTER="${APP_CLUSTER:-knoe-dev-0}"
SERVICE_NS="${SERVICE_NS:-knoe-system}"
DB_NS="${DB_NS:-knoe-db-0}"
CONFIRM="${CONFIRM:-false}"
DB_CTX="gke_${GCP_PROJECT}_${GCP_REGION}_${DB_CLUSTER}"
APP_CTX="gke_${GCP_PROJECT}_${GCP_REGION}_${APP_CLUSTER}"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
log() { printf "[%s] %s\n" "$(date +%H:%M:%S)" "$*"; }
die() { log "ERROR: $*" >&2; exit 1; }
# ---------------------------------------------------------------------------
# Pre-flight
# ---------------------------------------------------------------------------
if [[ "${CONFIRM}" != "true" ]]; then
echo ""
echo " This script will:"
echo " 1. Remove Garage from ${DB_CLUSTER}/${SERVICE_NS}"
echo " 2. Apply cnpg-backup-sa (Workload Identity) in ${DB_CLUSTER}/${DB_NS}"
echo " 3. Apply the GCS barman ObjectStore in ${DB_CLUSTER}/${DB_NS}"
echo " 4. Refresh CNPG_ELIGIBLE_NODES in conf/prod/prole.cfg"
echo ""
echo " Set CONFIRM=true to proceed:"
echo " CONFIRM=true ./scripts/patch_garage_cross_cluster.sh"
echo ""
exit 1
fi
for tool in gcloud kubectl; do
command -v "$tool" >/dev/null 2>&1 || die "required tool not found: $tool"
done
command -v gke-gcloud-auth-plugin >/dev/null 2>&1 \
|| die "gke-gcloud-auth-plugin not found — install with: gcloud components install gke-gcloud-auth-plugin"
log "==> patch_garage_cross_cluster (GCS mode)"
log " DB cluster: ${DB_CTX}"
log " App cluster: ${APP_CTX}"
# ---------------------------------------------------------------------------
# Phase 1: Remove Garage from the DB cluster (knoe-cnpg-0)
# ---------------------------------------------------------------------------
log "[phase 1] Removing Garage from ${DB_CLUSTER}/${SERVICE_NS} ..."
for resource in "statefulset/garage" "service/garage" "configmap/garage-config"; do
if kubectl --context="${DB_CTX}" -n "${SERVICE_NS}" get "${resource}" \
>/dev/null 2>&1; then
log " Deleting ${resource} from ${DB_CLUSTER}/${SERVICE_NS}"
kubectl --context="${DB_CTX}" -n "${SERVICE_NS}" delete "${resource}" \
--ignore-not-found=true --wait=false
else
log " ${resource} not found in ${DB_CLUSTER}/${SERVICE_NS} — skipping"
fi
done
if kubectl --context="${DB_CTX}" -n "${SERVICE_NS}" get pvc data-garage-0 \
>/dev/null 2>&1; then
log " Deleting PVC data-garage-0 from ${DB_CLUSTER}/${SERVICE_NS}"
kubectl --context="${DB_CTX}" -n "${SERVICE_NS}" delete pvc data-garage-0 \
--ignore-not-found=true
fi
if kubectl --context="${DB_CTX}" get storageclass garage-hdd >/dev/null 2>&1; then
log " Deleting StorageClass garage-hdd from ${DB_CLUSTER}"
kubectl --context="${DB_CTX}" delete storageclass garage-hdd --ignore-not-found=true
fi
log " Garage removed from ${DB_CLUSTER}."
# ---------------------------------------------------------------------------
# Phase 2: Apply cnpg-backup-sa ServiceAccount with Workload Identity
# ---------------------------------------------------------------------------
log "[phase 2] Applying cnpg-backup-sa (Workload Identity) in ${DB_CLUSTER}/${DB_NS} ..."
GCS_MANIFEST="${REPO_ROOT}/deploy/gcp/gke/knoe-db-backup-gcs.yaml"
[[ -f "${GCS_MANIFEST}" ]] || die "GCS manifest not found: ${GCS_MANIFEST}"
GCP_PROJECT_ID="${GCP_PROJECT}" envsubst '${GCP_PROJECT_ID}' < "${GCS_MANIFEST}" \
| kubectl --context="${DB_CTX}" apply -f -
log " cnpg-backup-sa applied."
# ---------------------------------------------------------------------------
# Phase 3: Apply GCS barman ObjectStore
# ---------------------------------------------------------------------------
log "[phase 3] Applying GCS barman ObjectStore in ${DB_CLUSTER}/${DB_NS} ..."
GCS_OBJ_MANIFEST="${REPO_ROOT}/k8s/prole/knoe-db-barman-objectstore-gcs.yaml"
[[ -f "${GCS_OBJ_MANIFEST}" ]] || die "GCS ObjectStore manifest not found: ${GCS_OBJ_MANIFEST}"
kubectl --context="${DB_CTX}" -n "${DB_NS}" apply -f "${GCS_OBJ_MANIFEST}"
log " GCS ObjectStore applied."
# ---------------------------------------------------------------------------
# Phase 4: Refresh CNPG_ELIGIBLE_NODES in conf/prod/prole.cfg
# ---------------------------------------------------------------------------
log "[phase 4] Refreshing CNPG_ELIGIBLE_NODES from ${DB_CLUSTER} ..."
CNPG_NODES=$(kubectl --context="${DB_CTX}" get nodes \
-o jsonpath='{range .items[*]}{.metadata.name}{","}{end}' 2>/dev/null \
| sed 's/,$//' || true)
PROLE_CFG="${REPO_ROOT}/conf/prod/prole.cfg"
if [[ -n "${CNPG_NODES}" && -f "${PROLE_CFG}" ]]; then
STAGE1_NODE=$(echo "${CNPG_NODES}" | cut -d',' -f1)
log " Updating CNPG_ELIGIBLE_NODES=${CNPG_NODES}"
log " Updating CNPG_STAGE1_NODE=${STAGE1_NODE}"
# Replace CNPG_ELIGIBLE_NODES line
sed -i.bak \
"s|^CNPG_ELIGIBLE_NODES = .*|CNPG_ELIGIBLE_NODES = ${CNPG_NODES}|" \
"${PROLE_CFG}"
# Replace CNPG_STAGE1_NODE line
sed -i.bak \
"s|^CNPG_STAGE1_NODE = .*|CNPG_STAGE1_NODE = ${STAGE1_NODE}|" \
"${PROLE_CFG}"
rm -f "${PROLE_CFG}.bak"
log " prole.cfg updated."
else
log " WARNING: Could not read nodes from ${DB_CLUSTER} — CNPG_ELIGIBLE_NODES not updated."
fi
# ---------------------------------------------------------------------------
# Done
# ---------------------------------------------------------------------------
echo ""
log "==> patch_garage_cross_cluster complete."
echo ""
echo "────────────────────────────────────────────────────────────────"
echo " Verify:"
echo ""
echo " # Garage gone from DB cluster"
echo " kubectl --context=${DB_CTX} -n ${SERVICE_NS} get statefulset garage"
echo ""
echo " # WI service account"
echo " kubectl --context=${DB_CTX} -n ${DB_NS} get sa cnpg-backup-sa -o yaml"
echo ""
echo " # GCS ObjectStore"
echo " kubectl --context=${DB_CTX} -n ${DB_NS} get objectstore knoe-db-barman-objectstore -o yaml"
echo ""
echo " If barman plugin isn't configured yet:"
echo " KUBECONTEXT=${DB_CTX} ./etc/init_cnpg_backup.sh start"
echo "────────────────────────────────────────────────────────────────"

View File

@ -6,8 +6,8 @@
# - Uses pd-balanced boot disks (counts against SSD_TOTAL_GB quota, 300 GB at limit)
#
# Solution: Delete both, recreate with pd-standard boot disks:
# knoe-dev-0 → Autopilot (app workloads: GitLab, platform), pd-standard boot
# knoe-cnpg-0 → Standard (CNPG only, 3× e2-standard-2), pd-standard boot
# knoe-dev-0 → Standard (app workloads: GitLab, platform), e2-small × 3, VPA, pd-standard boot
# knoe-cnpg-0 → Standard (CNPG only, 3× e2-small), pd-standard boot
#
# SSD quota budget after reset:
# Boot disks: pd-standard — does NOT count against SSD_TOTAL_GB
@ -23,12 +23,14 @@
# GCP_REGION (default: us-west3)
# APP_CLUSTER (default: knoe-dev-0)
# DB_CLUSTER (default: knoe-cnpg-0)
# DB_MACHINE_TYPE (default: e2-standard-2)
# DB_MACHINE_TYPE (default: e2-small)
# DB_DISK_TYPE (default: pd-standard)
# DB_DISK_SIZE_GB (default: 50)
# DB_NODES_PER_ZONE (default: 1 → 3 nodes across 3 zones)
# APP_MACHINE_TYPE (default: e2-small)
# APP_DISK_TYPE (default: pd-standard)
# APP_DISK_SIZE_GB (default: 50)
# APP_NODES_PER_ZONE (default: 1 → 3 nodes across 3 zones)
# CONFIRM REQUIRED: must be "true" to allow destructive operations
# DRY_RUN (default: false)
@ -42,12 +44,14 @@ GCP_PROJECT="${GCP_PROJECT:-plenary-truck-485623-p7}"
GCP_REGION="${GCP_REGION:-us-west3}"
APP_CLUSTER="${APP_CLUSTER:-knoe-dev-0}"
DB_CLUSTER="${DB_CLUSTER:-knoe-cnpg-0}"
DB_MACHINE_TYPE="${DB_MACHINE_TYPE:-e2-standard-2}"
DB_MACHINE_TYPE="${DB_MACHINE_TYPE:-e2-small}"
DB_DISK_TYPE="${DB_DISK_TYPE:-pd-standard}"
DB_DISK_SIZE_GB="${DB_DISK_SIZE_GB:-50}"
DB_NODES_PER_ZONE="${DB_NODES_PER_ZONE:-1}" # regional cluster = 3 zones = 3 nodes total
APP_MACHINE_TYPE="${APP_MACHINE_TYPE:-e2-small}"
APP_DISK_TYPE="${APP_DISK_TYPE:-pd-standard}"
APP_DISK_SIZE_GB="${APP_DISK_SIZE_GB:-50}"
APP_NODES_PER_ZONE="${APP_NODES_PER_ZONE:-1}" # regional cluster = 3 zones = 3 nodes total
CONFIRM="${CONFIRM:-false}"
DRY_RUN="${DRY_RUN:-false}"
@ -173,6 +177,48 @@ if [[ "${DRY_RUN}" != "true" ]]; then
log " All deletions complete."
fi
# ---------------------------------------------------------------------------
# Phase 2b: Delete orphaned pd-ssd disks left behind by cluster deletion
# ---------------------------------------------------------------------------
# GKE cluster deletion does not remove PersistentVolume-backed GCP disks.
# Orphaned pd-ssd disks hold SSD quota and prevent reset_clusters.sh from
# proceeding (Phase 3 waits for SSD usage < 50 GB).
log "[phase 2b] Deleting orphaned pd-ssd disks in ${GCP_REGION} ..."
if [[ "${DRY_RUN}" == "true" ]]; then
log " [DRY-RUN] Would list and delete orphaned pd-ssd disks in ${GCP_REGION}"
else
# List all disks in the project, filter to pd-ssd in this region in bash
mapfile -t ssd_disks < <(
gcloud compute disks list \
--project="${GCP_PROJECT}" \
--format="csv[no-heading](name,zone,type)" \
--quiet 2>/dev/null \
| awk -F',' "/${GCP_REGION}/ && /pd-ssd/ {print \$1 \",\" \$2}" \
|| true
)
if [[ ${#ssd_disks[@]} -eq 0 ]]; then
log " No orphaned pd-ssd disks found."
else
for entry in "${ssd_disks[@]}"; do
disk_name="${entry%%,*}"
disk_zone="${entry##*,}"
# Extract just the zone name (may be a full URL)
disk_zone="${disk_zone##*/}"
log " Deleting pd-ssd disk: ${disk_name} (zone: ${disk_zone})"
gcloud compute disks delete "${disk_name}" \
--zone="${disk_zone}" \
--project="${GCP_PROJECT}" \
--quiet 2>/dev/null \
&& log " Deleted ${disk_name}." \
|| log " WARNING: Could not delete ${disk_name} — may already be gone."
done
log " Orphaned pd-ssd disk cleanup complete."
fi
fi
# ---------------------------------------------------------------------------
# Phase 3: Wait for SSD quota to be released
# ---------------------------------------------------------------------------
@ -229,22 +275,26 @@ run_cmd gcloud container clusters create "${DB_CLUSTER}" \
log " ${DB_CLUSTER} created."
# ---------------------------------------------------------------------------
# Phase 5: Create knoe-dev-0 as Autopilot cluster
# Phase 5: Create knoe-dev-0 as Standard cluster
# ---------------------------------------------------------------------------
log "[phase 5] Creating ${APP_CLUSTER} (Autopilot, ${APP_DISK_TYPE}, ${APP_DISK_SIZE_GB}GB) ..."
log "[phase 5] Creating ${APP_CLUSTER} (Standard, ${APP_MACHINE_TYPE}, ${APP_DISK_TYPE}, ${APP_DISK_SIZE_GB}GB, VPA) ..."
run_cmd gcloud container clusters create-auto "${APP_CLUSTER}" \
run_cmd gcloud container clusters create "${APP_CLUSTER}" \
--project="${GCP_PROJECT}" \
--region="${GCP_REGION}" \
--cluster-version=latest \
--workload-policies=allow-net-admin \
--machine-type="${APP_MACHINE_TYPE}" \
--disk-type="${APP_DISK_TYPE}" \
--disk-size="${APP_DISK_SIZE_GB}" \
--num-nodes="${APP_NODES_PER_ZONE}" \
--enable-vertical-pod-autoscaling \
--enable-ip-alias \
--workload-pool="${GCP_PROJECT}.svc.id.goog" \
--quiet
# NOTE: Autopilot manages node infrastructure (disk type, size) automatically.
# Boot disk type cannot be set at Autopilot cluster creation time.
# Autopilot only provisions nodes when pods are scheduled — no quota impact
# while the cluster is idle. Deploy workloads after the SSD quota increase
# (300 GB → 2 TB) is approved to ensure headroom for GitLab node boot disks.
# Standard mode: pd-standard boot disks consume ZERO SSD quota (vs Autopilot's
# pd-balanced 100GB/node which exhausted the full 300GB SSD_TOTAL_GB quota).
# VPA handles dynamic resource adjustment for burstable e2-medium nodes.
log " ${APP_CLUSTER} created."

View File

@ -0,0 +1,92 @@
from __future__ import annotations
from knoe.ui.screens.database import DatabaseScreenMixin
class _Var:
def __init__(self, value: str = ""):
self._value = value
def get(self):
return self._value
def set(self, value):
self._value = value
class _DummyDbScreen(DatabaseScreenMixin):
def __init__(self, *, inputs: dict | None = None, prole_cfg_data: dict | None = None):
self.inputs = inputs or {}
self.prole_cfg_data = prole_cfg_data or {}
self.selected_kubectx = _Var(
(self.inputs.get("init_cluster.selected_kubectx") or "")
)
def _get_input(self, key: str, default=""):
return self.inputs.get(key, default)
def _deployment_mode(self):
return "k8s"
def _kubectl_base_cmd(self, mode: str | None = None):
_ = mode
return ["kubectl", "--context", "gke_proj_us-west3_knoe-dev-0"]
def test_db_kubectl_base_cmd_prefers_db_cluster_context_over_dev_context():
screen = _DummyDbScreen(
prole_cfg_data={
"Global": {
"KUBECONTEXT": "gke_proj_us-west3_knoe-dev-0",
"DB_CLUSTER_KUBECONTEXT": "gke_proj_us-west3_knoe-cnpg-0",
}
}
)
cmd = screen._db_kubectl_base_cmd()
assert cmd == ["kubectl", "--context", "gke_proj_us-west3_knoe-cnpg-0"]
def test_resolve_db_kube_context_derives_cnpg_context_when_only_app_context_exists():
screen = _DummyDbScreen(
prole_cfg_data={
"Global": {
"KUBECONTEXT": "gke_proj_us-west3_knoe-dev-0",
"DB_CLUSTER_NAME": "knoe-cnpg-0",
}
}
)
assert screen._resolve_db_kube_context() == "gke_proj_us-west3_knoe-cnpg-0"
def test_ensure_db_screen_context_keeps_explicit_user_override():
screen = _DummyDbScreen(
inputs={"init_cluster.selected_kubectx": "gke_proj_us-west3_custom-ops"},
prole_cfg_data={
"Global": {
"KUBECONTEXT": "gke_proj_us-west3_knoe-dev-0",
"DB_CLUSTER_KUBECONTEXT": "gke_proj_us-west3_knoe-cnpg-0",
}
},
)
screen._ensure_db_screen_context()
assert screen.selected_kubectx.get() == "gke_proj_us-west3_custom-ops"
def test_db_target_note_includes_context_and_namespace_confirmation():
screen = _DummyDbScreen(
prole_cfg_data={
"Global": {
"DB_CLUSTER_KUBECONTEXT": "gke_proj_us-west3_knoe-cnpg-0",
}
}
)
note = screen._db_target_note("knoe-db")
assert "Target context: gke_proj_us-west3_knoe-cnpg-0" in note
assert "Target namespace: knoe-db" in note