mirror of
https://github.com/dredx/prole.git
synced 2026-09-24 19:04:31 +00:00
Bringing the long-running session-feature branch back into main in one deliberate sweep. The branch carried the cluster work that's been live for weeks (cross-cluster CNPG metrics, Grafana w/ Google OAuth, supabase oauth2-proxy, cluster recovery, pg.0.knoe.dev + per-engineer onboarding, GCS-backed CNPG backups via Workload Identity, the env-contamination guard, the Junie brief queue, the cnpg-grafana CSRF + memory-request fixes from today), while main accumulated Junie's parallel knoe-auth Phase 2 OIDC work (full provider surface: discovery, authorize, token, userinfo, JWKS, RS256 signing, code exchange, session services). Key decision: the two branches did COMPETING rebrands off the same starting point (5ba9b63, 2026-04-27): - claude branch (commit b355855, earlier): org.prole.authority.* → dev.knoe.auth.* (artifact renamed to knoe-auth.jar) - main (commit9daa94b, recent): org.prole.authority.* → dev.knoe.authority.* (kept "authority" artifact name) dev.knoe.auth wins: cluster runs from this name, the Maven artifact is already knoe-auth.jar, and the broader rename is the documented namespace direction (per ~/.claude/projects/-Users-chrisfu-dev-knoe-db/ memory/MEMORY.md). All of main's recent Phase 2 OIDC content was ported from authority/src/.../dev/knoe/authority/ into authority/src/.../dev/knoe/auth/ with package declarations rewritten. == File-level resolution summary == Textual conflicts (4): authority/pom.xml - Took our artifactId="auth" - Took our branch's removal of spring-security-kerberos-client (verified: Junie's Phase 2 OIDC code does not import it; the dep was already-dead config) docs/pipeline-phases.md - Took our branch's "Phase 1 not started" status. Main had a misplaced "✅ Complete" with a knoe-auth-Phase-1 commit ref in the autobuild Phase 1 section — different domain. docs/plans/knoe-auth-round-1.md - Took our branch's dev.knoe.auth file table (vs main's dev.knoe.authority listing). Pure rename mismatch. supabase/helm/knoe-supabase/templates/kong/config.yaml - Took our branch's onboard route + plain dashboard wiring. Main had an oauth2proxy.enabled toggle that put oauth2-proxy as a Kong upstream — but the deployed architecture (commit 25f1b2e) has oauth2-proxy in FRONT of Kong, not behind. Main's wrapper reflected an architecture that was never deployed. - Took our branch's removal of basic-auth from dashboard route (queue #15 brief still tracks the matching values.yaml / kong/deployment.yaml cleanup). Java tree reconciliation (44 file-pairs): 20 dual-path source files + 2 dual-path tests Body-identical between main's authority/ and our branch's auth/ after stripping package decls — main's commit9daa94bwas a pure rebrand. Took our branch's auth/ version for all 22. 8 main-only source files (Phase 2 OIDC), ported into auth/: web/JwksController.java web/OidcAuthorizeController.java web/OidcDiscoveryController.java web/OidcTokenController.java web/OidcUserInfoController.java session/OidcCodeService.java session/OidcTokenService.java session/SessionService.java 12 main-only test files, ported into auth/: HealthControllerTest.java enroll/EnrollValueTypesTest.java enroll/EnrollmentControllerTest.java enroll/TotpServiceTest.java kerberos/KadminClientTest.java kerberos/KerberosSpnegoResultTest.java web/LoginControllerTest.java admin/AdminControllerTest.java user/PrincipalNormalizerTest.java regression/IdentityRegressionTest.java session/OidcCodeServiceTest.java session/SessionServiceTest.java Port mechanics: read main:authority/...<file> via git show, then sed rewrite `package dev.knoe.authority` → `package dev.knoe.auth` and `import dev.knoe.authority` → `import dev.knoe.auth`. Body content unchanged. authority/src/main/java/dev/knoe/authority/ — DELETED (duplicate) authority/src/test/java/dev/knoe/authority/ — DELETED (duplicate) == Verification == - grep -rln '<<<<<<<' across .java/.md/.yaml/.yml/.sh/.xml/.tpl: clean - find authority/src -path '*/dev/knoe/authority*': empty (subtree gone) - grep 'package dev.knoe.authority' across repo: clean - bash -n install.sh deploy.sh etc/preflight_kubecontext.sh: clean - git ls-files -u | wc -l: 0 unmerged paths - helm lint supabase/helm/knoe-supabase: pre-existing failure on studioIngress.enabled undefined in values.yaml (introduced by Junie on main; unrelated to this merge — flagging as follow-up). == Followups (carried into TODO ranked queue or noted here) == - helm lint failure: studioIngress block in values.yaml is missing enable flag; templates/studio/{ingress,oauth2proxy-deployment, oauth2proxy-service}.yaml all reference studioIngress.enabled with no default. Pre-existing on main; not introduced by this merge. - The five Junie briefs filed on this branch are now reachable from main at docs/plans/junie/{02,06,07,13,15}-*.md. Junie can pick them up in any order. - knoe-auth Phase 2 OIDC source (now at dev.knoe.auth.*) is not yet deployed to the cluster. Deployment is its own task. - The branch claude/crazy-bose-fec256 stays in place (worktree at .claude/worktrees/crazy-bose-fec256 may have ongoing context for Claude Code sessions). Safe to delete once next session starts cleanly from main. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
209 lines
11 KiB
Markdown
209 lines
11 KiB
Markdown
# CLAUDE.md — knoe project context
|
||
|
||
> Loaded automatically by Claude Code (CLI and IntelliJ plugin) as project context.
|
||
|
||
---
|
||
|
||
## Repo role: knoe-db platform repo
|
||
|
||
This is `knoe-db` (remote: `git@git-ssh.knoe.dev:knoe.dev/knoe-db.git`, configured as both `origin` and `knoe`). It's the platform's source of truth — `authority/`, `knoe/`, `etc/init_*.sh`, `deploy/gcp/gke/*`, the test pipeline, all live here.
|
||
|
||
Customer deploys are intended to live as **branches** in this repo (e.g. a future `customer/prole.org`), not as separate forks. As of this writing, no customer branch is active — the prole→knoe rebrand has merged into `main` and there is no separate `~/dev/prole` working tree under development. See [`docs/plans/customer-deploy-resync.md`](docs/plans/customer-deploy-resync.md) for the original plan, currently dormant.
|
||
|
||
## Master TODO
|
||
|
||
Single source of truth for unfinished work, including the reality-vs-intent gaps flagged below: **[`docs/TODO.md`](docs/TODO.md)**.
|
||
|
||
## Active plans
|
||
|
||
- [`docs/plans/README.md`](docs/plans/README.md) — Index and conventions for this directory.
|
||
- [`docs/plans/knoe-auth-round-1.md`](docs/plans/knoe-auth-round-1.md) — Identity backbone (Kerberos KDC + invite-OTP + TOTP). **Shipped.**
|
||
- [`docs/plans/deployment-modes.md`](docs/plans/deployment-modes.md) — Four installer modes (`min` / `k3d` / `k3s` / `gke`). **Shipped (Phase 0).**
|
||
- [`docs/pipeline-phases.md`](docs/pipeline-phases.md) — Autobuild & test pipeline phase reference. Phase 0 ✅, Phase 1 next (first task: f-string fix at `knoe/core/ops/cloudnative_pg.py:1372`).
|
||
- [`docs/plans/customer-deploy-resync.md`](docs/plans/customer-deploy-resync.md) — The original plan to converge `~/dev/prole` onto `knoe-db/main` as a customer-deploy branch. **Dormant** (rebrand is now in main, no separate prole tree active).
|
||
|
||
---
|
||
|
||
## Dual-cluster GKE architecture
|
||
|
||
This project uses **two separate GKE Standard clusters** in `us-west3`, both currently provisioned with `e2-standard-2` × 3 nodes (2 vCPU / 8 GB each, ~7.1 GB allocatable). Verify with `gcloud container clusters list`.
|
||
|
||
| Cluster | Context | Role |
|
||
|---|---|---|
|
||
| `knoe-dev-0` | `gke_plenary-truck-485623-p7_us-west3_knoe-dev-0` | App cluster — Garage, registry, OpenBao, Kong, GitLab, monitoring |
|
||
| `knoe-dev-cnpg-0` | `gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0` | DB cluster — CNPG/PostgreSQL only |
|
||
|
||
**Storage quota:** the project has `SSD_TOTAL_GB = 300 GB` in `us-west3`, **fully consumed by CNPG**. All non-CNPG PVCs must use `standard` (pd-standard / HDD) — not `standard-rwo` / `premium-rwo`, which are SSD-backed and will fail to provision with a quota error. `GITLAB_GITALY_STORAGE_CLASS = standard` is set in `conf/gke.cfg` accordingly.
|
||
|
||
### Resource allocation
|
||
|
||
| Resource | Cluster | Namespace |
|
||
|---|---|---|
|
||
| CNPG operator (v1.29.0) | `knoe-dev-cnpg-0` | `cnpg-system` |
|
||
| PostgreSQL cluster (`knoe-db`) | `knoe-dev-cnpg-0` | `knoe-db-0` |
|
||
| Barman Cloud plugin (v0.12.0) | `knoe-dev-cnpg-0` | `cnpg-system` |
|
||
| cert-manager | `knoe-dev-cnpg-0` | `cert-manager` |
|
||
| Garage (S3 object store) | `knoe-dev-0` | `knoe-system` |
|
||
| Registry | `knoe-dev-0` | `knoe-system` |
|
||
| OpenBao | `knoe-dev-0` | `knoe-system` |
|
||
| Kong API gateway | `knoe-dev-0` | `knoe-system` |
|
||
| Monitoring | `knoe-dev-0` | `monitoring` |
|
||
|
||
**Garage runs ONLY in `knoe-dev-0`.** The DB cluster (`knoe-dev-cnpg-0`) has none — was removed 2026-04-29. Do NOT redeploy Garage to the DB cluster (use GCS for backups there).
|
||
|
||
### CNPG backups → GCS
|
||
|
||
Backups use **GCS with Workload Identity**:
|
||
- Data + WAL: `gs://knoe-0-backups/` (single bucket; `knoe-db/base/` and `knoe-db/wals/` prefixes). `gs://knoe-0-wal/` exists but is unused.
|
||
- GCP SA: `cnpg-backup@plenary-truck-485623-p7.iam.gserviceaccount.com` (roles on bucket: `storage.objectAdmin`, `storage.legacyBucketReader`)
|
||
- K8s SA: cluster pods run as **`cnpg-backup-sa`** in `knoe-db-0` (set via `cluster.spec.serviceAccountName`, requires CNPG ≥ v1.29.0). The SA is annotated with `iam.gke.io/gcp-service-account=cnpg-backup@…`. Two `RoleBinding` subjects (`knoe-db` and `knoe-db-barman-cloud`) include `cnpg-backup-sa` so the pod has the same RBAC the auto-generated SA would have had.
|
||
- ObjectStore manifest: [`k8s/knoe/knoe-db-barman-objectstore-gcs.yaml`](k8s/knoe/knoe-db-barman-objectstore-gcs.yaml) — includes `googleCredentials.gkeEnvironment: true` (required by plugin-barman-cloud v0.12.0).
|
||
|
||
Setup script: [`etc/init_cnpg_gke.sh`](etc/init_cnpg_gke.sh) (creates buckets, GCP SA, WI binding, applies CNPG cluster).
|
||
|
||
|
||
---
|
||
|
||
## Service mesh (Cloud Service Mesh / Istio)
|
||
|
||
Both clusters are registered in the **knoe-0** GCP fleet with automatic Cloud Service Mesh management. This is automated in `scripts/reset_clusters.sh` (Phase 7) — no longer requires GCP web console.
|
||
|
||
```bash
|
||
# Check mesh provisioning status (~10 min after cluster creation):
|
||
gcloud container fleet mesh describe --project=plenary-truck-485623-p7
|
||
|
||
# Manual re-registration if needed:
|
||
gcloud container fleet memberships register knoe-dev-0 \
|
||
--gke-cluster=us-west3/knoe-dev-0 \
|
||
--enable-workload-identity \
|
||
--project=plenary-truck-485623-p7
|
||
gcloud container fleet mesh update \
|
||
--management=automatic \
|
||
--memberships=knoe-dev-0 \
|
||
--project=plenary-truck-485623-p7
|
||
```
|
||
|
||
---
|
||
|
||
## install.sh / deploy.sh pre-flight checklist
|
||
|
||
### `conf/gke.cfg` (interactive installer — `./install.sh`)
|
||
|
||
Before running `./install.sh` (especially "Initialization Scripts"), confirm these are correct:
|
||
|
||
```ini
|
||
[Inputs]
|
||
init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
init_cluster.db_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
|
||
env_setup.APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
env_setup.DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
|
||
|
||
[Global]
|
||
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
|
||
CNPG_ELIGIBLE_NODES = <comma-separated node names from knoe-dev-cnpg-0>
|
||
```
|
||
|
||
### `conf/service/prod.cfg` (unattended deploy — `./deploy.sh`)
|
||
|
||
Same cluster context entries are required here too:
|
||
|
||
```ini
|
||
[Inputs]
|
||
init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
init_cluster.db_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
|
||
env_setup.APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
env_setup.DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
|
||
|
||
[Global]
|
||
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
|
||
DB_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
|
||
ARTIFACT_REGISTRY = us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system
|
||
SERVICE_NAMESPACE = knoe-system
|
||
REGISTRY_NAMESPACE = knoe-system
|
||
```
|
||
|
||
**Why these matter:** `Milestone._get_script_env()` (in `knoe/milestone.py`) reads these to set `KUBECONTEXT=app_ctx` for common services and `DB_CLUSTER_KUBECONTEXT=db_ctx` for CNPG ops. Without them, all kubectl calls use the ambient context, which may be the DB cluster.
|
||
|
||
Missing `init_cluster.app_cluster_kubecontext` → `_cluster_kubecontext("app")` returns `""` → installer falls back to `Global.KUBECONTEXT` for **both** app and db environments → **Garage deploys to knoe-dev-cnpg-0** (wrong).
|
||
|
||
> **Env-contamination guard (live):** `deploy.sh` calls
|
||
> [`etc/preflight_kubecontext.sh`](etc/preflight_kubecontext.sh) and refuses
|
||
> to proceed when `kubectl config current-context` doesn't match the
|
||
> `[Global] APP_CLUSTER_KUBECONTEXT` of the active config. `install.sh`
|
||
> prints the inherited context up-front (mode-aware strict gate is the
|
||
> Python TUI's responsibility once the welcome screen records a mode).
|
||
> Bypass with `KNOE_SKIP_KUBECONTEXT_GUARD=true` for deliberate
|
||
> cross-cluster maintenance. **History:** the guard was filed in response
|
||
> to the 2026-04-28 14:00 UTC outage — an `install.sh --mode k3d` run with
|
||
> the shell pointed at GKE replaced the GCS-backed ObjectStore with a
|
||
> Garage-backed one, then Garage filled up and backups silently failed for
|
||
> hours. Closes drift R4 / queue item #1.
|
||
|
||
> **Env-contamination guard (live):** `deploy.sh` calls
|
||
> [`etc/preflight_kubecontext.sh`](etc/preflight_kubecontext.sh) and refuses
|
||
> to proceed when `kubectl config current-context` doesn't match the
|
||
> `[Global] APP_CLUSTER_KUBECONTEXT` of the active config. `install.sh`
|
||
> prints the inherited context up-front (mode-aware strict gate is the
|
||
> Python TUI's responsibility once the welcome screen records a mode).
|
||
> Bypass with `KNOE_SKIP_KUBECONTEXT_GUARD=true` for deliberate
|
||
> cross-cluster maintenance. **History:** the guard was filed in response
|
||
> to the 2026-04-28 14:00 UTC outage — an `install.sh --mode k3d` run with
|
||
> the shell pointed at GKE replaced the GCS-backed ObjectStore with a
|
||
> Garage-backed one, then Garage filled up and backups silently failed for
|
||
> hours. Closes drift R4 / queue item #1.
|
||
|
||
### Get current CNPG node names
|
||
|
||
```bash
|
||
kubectl --context=gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 get nodes -o name
|
||
```
|
||
|
||
---
|
||
|
||
## Key scripts
|
||
|
||
| Script | Purpose |
|
||
|---|---|
|
||
| `scripts/reset_clusters.sh` | Delete + recreate both GKE clusters (`e2-standard-2` × 3 each) |
|
||
| `scripts/patch_garage_cross_cluster.sh` | Remove garage from DB cluster, configure GCS barman ObjectStore |
|
||
| `etc/init_cnpg_gke.sh` | Provision CNPG on GKE with GCS backup via Workload Identity |
|
||
| `etc/init_cnpg_backup.sh` | Configure barman plugin + trigger initial backup |
|
||
|
||
---
|
||
|
||
## Cluster code constants (`knoe/core/actions.py`)
|
||
|
||
```python
|
||
DEFAULT_APP_CLUSTER_NAME = "knoe-dev-0"
|
||
DEFAULT_DB_CLUSTER_NAME = "knoe-dev-cnpg-0"
|
||
DEFAULT_APP_CLUSTER_MACHINE_TYPE = "e2-standard-2"
|
||
DEFAULT_DB_CLUSTER_MACHINE_TYPE = "e2-standard-2"
|
||
```
|
||
|
||
---
|
||
|
||
## GCP project
|
||
|
||
- Project: `plenary-truck-485623-p7`
|
||
- Region: `us-west3`
|
||
- Artifact Registry: `us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system`
|
||
- CPU quota: 16 vCPUs (all regions) — 2 clusters × 3 × `e2-standard-2` = 12 vCPUs used
|
||
- SSD quota: `SSD_TOTAL_GB = 300 GB` — fully consumed by CNPG; all other PVCs must use `standard` (pd-standard / HDD)
|
||
|
||
---
|
||
|
||
## Reality TODOs / Drift Log
|
||
|
||
Quick reference. Each entry links to the master index where context, owner, and rank live.
|
||
|
||
| # | Drift | Where described above | Where tracked |
|
||
|---|---|---|---|
|
||
| _(none currently)_ | | | |
|
||
|
||
**Closed in 2026-04-29 stabilization session:** Garage on DB cluster removed; cluster pods migrated to `cnpg-backup-sa` via CNPG v1.29.0 `spec.serviceAccountName`; both operators restarted clean.
|
||
|
||
**Closed 2026-05-01:** R4 — installer env-contamination guard now live in `deploy.sh` (strict) + `install.sh` (informational notice). Helper at [`etc/preflight_kubecontext.sh`](etc/preflight_kubecontext.sh).
|