prole/CLAUDE.md
chrisfu 00f0ebec07 Merge claude/crazy-bose-fec256 into main
Bringing the long-running session-feature branch back into main in one
deliberate sweep. The branch carried the cluster work that's been live for
weeks (cross-cluster CNPG metrics, Grafana w/ Google OAuth, supabase
oauth2-proxy, cluster recovery, pg.0.knoe.dev + per-engineer onboarding,
GCS-backed CNPG backups via Workload Identity, the env-contamination
guard, the Junie brief queue, the cnpg-grafana CSRF + memory-request
fixes from today), while main accumulated Junie's parallel knoe-auth
Phase 2 OIDC work (full provider surface: discovery, authorize, token,
userinfo, JWKS, RS256 signing, code exchange, session services).

Key decision: the two branches did COMPETING rebrands off the same
starting point (5ba9b63, 2026-04-27):

  - claude branch (commit b355855, earlier): org.prole.authority.* →
                                              dev.knoe.auth.*
                                              (artifact renamed to
                                              knoe-auth.jar)
  - main (commit 9daa94b, recent): org.prole.authority.* →
                                    dev.knoe.authority.*
                                    (kept "authority" artifact name)

dev.knoe.auth wins: cluster runs from this name, the Maven artifact is
already knoe-auth.jar, and the broader rename is the documented
namespace direction (per ~/.claude/projects/-Users-chrisfu-dev-knoe-db/
memory/MEMORY.md). All of main's recent Phase 2 OIDC content was ported
from authority/src/.../dev/knoe/authority/ into
authority/src/.../dev/knoe/auth/ with package declarations rewritten.

== File-level resolution summary ==

Textual conflicts (4):

  authority/pom.xml
    - Took our artifactId="auth"
    - Took our branch's removal of spring-security-kerberos-client
      (verified: Junie's Phase 2 OIDC code does not import it; the dep
      was already-dead config)

  docs/pipeline-phases.md
    - Took our branch's "Phase 1 not started" status. Main had a
      misplaced " Complete" with a knoe-auth-Phase-1 commit ref
      in the autobuild Phase 1 section — different domain.

  docs/plans/knoe-auth-round-1.md
    - Took our branch's dev.knoe.auth file table (vs main's
      dev.knoe.authority listing). Pure rename mismatch.

  supabase/helm/knoe-supabase/templates/kong/config.yaml
    - Took our branch's onboard route + plain dashboard wiring.
      Main had an oauth2proxy.enabled toggle that put oauth2-proxy as
      a Kong upstream — but the deployed architecture (commit 25f1b2e)
      has oauth2-proxy in FRONT of Kong, not behind. Main's wrapper
      reflected an architecture that was never deployed.
    - Took our branch's removal of basic-auth from dashboard route
      (queue #15 brief still tracks the matching values.yaml /
      kong/deployment.yaml cleanup).

Java tree reconciliation (44 file-pairs):

  20 dual-path source files + 2 dual-path tests
    Body-identical between main's authority/ and our branch's auth/
    after stripping package decls — main's commit 9daa94b was a pure
    rebrand. Took our branch's auth/ version for all 22.

  8 main-only source files (Phase 2 OIDC), ported into auth/:
    web/JwksController.java
    web/OidcAuthorizeController.java
    web/OidcDiscoveryController.java
    web/OidcTokenController.java
    web/OidcUserInfoController.java
    session/OidcCodeService.java
    session/OidcTokenService.java
    session/SessionService.java

  12 main-only test files, ported into auth/:
    HealthControllerTest.java
    enroll/EnrollValueTypesTest.java
    enroll/EnrollmentControllerTest.java
    enroll/TotpServiceTest.java
    kerberos/KadminClientTest.java
    kerberos/KerberosSpnegoResultTest.java
    web/LoginControllerTest.java
    admin/AdminControllerTest.java
    user/PrincipalNormalizerTest.java
    regression/IdentityRegressionTest.java
    session/OidcCodeServiceTest.java
    session/SessionServiceTest.java

  Port mechanics: read main:authority/...<file> via git show, then sed
  rewrite `package dev.knoe.authority` → `package dev.knoe.auth` and
  `import dev.knoe.authority` → `import dev.knoe.auth`. Body content
  unchanged.

  authority/src/main/java/dev/knoe/authority/ — DELETED (duplicate)
  authority/src/test/java/dev/knoe/authority/  — DELETED (duplicate)

== Verification ==

- grep -rln '<<<<<<<' across .java/.md/.yaml/.yml/.sh/.xml/.tpl: clean
- find authority/src -path '*/dev/knoe/authority*': empty (subtree gone)
- grep 'package dev.knoe.authority' across repo: clean
- bash -n install.sh deploy.sh etc/preflight_kubecontext.sh: clean
- git ls-files -u | wc -l: 0 unmerged paths
- helm lint supabase/helm/knoe-supabase: pre-existing failure on
  studioIngress.enabled undefined in values.yaml (introduced by Junie
  on main; unrelated to this merge — flagging as follow-up).

== Followups (carried into TODO ranked queue or noted here) ==

  - helm lint failure: studioIngress block in values.yaml is missing
    enable flag; templates/studio/{ingress,oauth2proxy-deployment,
    oauth2proxy-service}.yaml all reference studioIngress.enabled with
    no default. Pre-existing on main; not introduced by this merge.
  - The five Junie briefs filed on this branch are now reachable from
    main at docs/plans/junie/{02,06,07,13,15}-*.md. Junie can pick them
    up in any order.
  - knoe-auth Phase 2 OIDC source (now at dev.knoe.auth.*) is not yet
    deployed to the cluster. Deployment is its own task.
  - The branch claude/crazy-bose-fec256 stays in place (worktree at
    .claude/worktrees/crazy-bose-fec256 may have ongoing context for
    Claude Code sessions). Safe to delete once next session starts
    cleanly from main.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 16:39:10 -07:00

11 KiB
Raw Blame History

CLAUDE.md — knoe project context

Loaded automatically by Claude Code (CLI and IntelliJ plugin) as project context.


Repo role: knoe-db platform repo

This is knoe-db (remote: git@git-ssh.knoe.dev:knoe.dev/knoe-db.git, configured as both origin and knoe). It's the platform's source of truth — authority/, knoe/, etc/init_*.sh, deploy/gcp/gke/*, the test pipeline, all live here.

Customer deploys are intended to live as branches in this repo (e.g. a future customer/prole.org), not as separate forks. As of this writing, no customer branch is active — the prole→knoe rebrand has merged into main and there is no separate ~/dev/prole working tree under development. See docs/plans/customer-deploy-resync.md for the original plan, currently dormant.

Master TODO

Single source of truth for unfinished work, including the reality-vs-intent gaps flagged below: docs/TODO.md.

Active plans


Dual-cluster GKE architecture

This project uses two separate GKE Standard clusters in us-west3, both currently provisioned with e2-standard-2 × 3 nodes (2 vCPU / 8 GB each, ~7.1 GB allocatable). Verify with gcloud container clusters list.

Cluster Context Role
knoe-dev-0 gke_plenary-truck-485623-p7_us-west3_knoe-dev-0 App cluster — Garage, registry, OpenBao, Kong, GitLab, monitoring
knoe-dev-cnpg-0 gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 DB cluster — CNPG/PostgreSQL only

Storage quota: the project has SSD_TOTAL_GB = 300 GB in us-west3, fully consumed by CNPG. All non-CNPG PVCs must use standard (pd-standard / HDD) — not standard-rwo / premium-rwo, which are SSD-backed and will fail to provision with a quota error. GITLAB_GITALY_STORAGE_CLASS = standard is set in conf/gke.cfg accordingly.

Resource allocation

Resource Cluster Namespace
CNPG operator (v1.29.0) knoe-dev-cnpg-0 cnpg-system
PostgreSQL cluster (knoe-db) knoe-dev-cnpg-0 knoe-db-0
Barman Cloud plugin (v0.12.0) knoe-dev-cnpg-0 cnpg-system
cert-manager knoe-dev-cnpg-0 cert-manager
Garage (S3 object store) knoe-dev-0 knoe-system
Registry knoe-dev-0 knoe-system
OpenBao knoe-dev-0 knoe-system
Kong API gateway knoe-dev-0 knoe-system
Monitoring knoe-dev-0 monitoring

Garage runs ONLY in knoe-dev-0. The DB cluster (knoe-dev-cnpg-0) has none — was removed 2026-04-29. Do NOT redeploy Garage to the DB cluster (use GCS for backups there).

CNPG backups → GCS

Backups use GCS with Workload Identity:

  • Data + WAL: gs://knoe-0-backups/ (single bucket; knoe-db/base/ and knoe-db/wals/ prefixes). gs://knoe-0-wal/ exists but is unused.
  • GCP SA: cnpg-backup@plenary-truck-485623-p7.iam.gserviceaccount.com (roles on bucket: storage.objectAdmin, storage.legacyBucketReader)
  • K8s SA: cluster pods run as cnpg-backup-sa in knoe-db-0 (set via cluster.spec.serviceAccountName, requires CNPG ≥ v1.29.0). The SA is annotated with iam.gke.io/gcp-service-account=cnpg-backup@…. Two RoleBinding subjects (knoe-db and knoe-db-barman-cloud) include cnpg-backup-sa so the pod has the same RBAC the auto-generated SA would have had.
  • ObjectStore manifest: k8s/knoe/knoe-db-barman-objectstore-gcs.yaml — includes googleCredentials.gkeEnvironment: true (required by plugin-barman-cloud v0.12.0).

Setup script: etc/init_cnpg_gke.sh (creates buckets, GCP SA, WI binding, applies CNPG cluster).


Service mesh (Cloud Service Mesh / Istio)

Both clusters are registered in the knoe-0 GCP fleet with automatic Cloud Service Mesh management. This is automated in scripts/reset_clusters.sh (Phase 7) — no longer requires GCP web console.

# Check mesh provisioning status (~10 min after cluster creation):
gcloud container fleet mesh describe --project=plenary-truck-485623-p7

# Manual re-registration if needed:
gcloud container fleet memberships register knoe-dev-0 \
  --gke-cluster=us-west3/knoe-dev-0 \
  --enable-workload-identity \
  --project=plenary-truck-485623-p7
gcloud container fleet mesh update \
  --management=automatic \
  --memberships=knoe-dev-0 \
  --project=plenary-truck-485623-p7

install.sh / deploy.sh pre-flight checklist

conf/gke.cfg (interactive installer — ./install.sh)

Before running ./install.sh (especially "Initialization Scripts"), confirm these are correct:

[Inputs]
init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
init_cluster.db_cluster_kubecontext  = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
env_setup.APP_CLUSTER_KUBECONTEXT    = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
env_setup.DB_CLUSTER_KUBECONTEXT     = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0

[Global]
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
DB_CLUSTER_KUBECONTEXT  = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
CNPG_ELIGIBLE_NODES = <comma-separated node names from knoe-dev-cnpg-0>

conf/service/prod.cfg (unattended deploy — ./deploy.sh)

Same cluster context entries are required here too:

[Inputs]
init_cluster.app_cluster_kubecontext = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
init_cluster.db_cluster_kubecontext  = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
env_setup.APP_CLUSTER_KUBECONTEXT    = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
env_setup.DB_CLUSTER_KUBECONTEXT     = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0

[Global]
KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
APP_CLUSTER_KUBECONTEXT = gke_plenary-truck-485623-p7_us-west3_knoe-dev-0
DB_CLUSTER_KUBECONTEXT  = gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0
ARTIFACT_REGISTRY = us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system
SERVICE_NAMESPACE = knoe-system
REGISTRY_NAMESPACE = knoe-system

Why these matter: Milestone._get_script_env() (in knoe/milestone.py) reads these to set KUBECONTEXT=app_ctx for common services and DB_CLUSTER_KUBECONTEXT=db_ctx for CNPG ops. Without them, all kubectl calls use the ambient context, which may be the DB cluster.

Missing init_cluster.app_cluster_kubecontext_cluster_kubecontext("app") returns "" → installer falls back to Global.KUBECONTEXT for both app and db environments → Garage deploys to knoe-dev-cnpg-0 (wrong).

Env-contamination guard (live): deploy.sh calls etc/preflight_kubecontext.sh and refuses to proceed when kubectl config current-context doesn't match the [Global] APP_CLUSTER_KUBECONTEXT of the active config. install.sh prints the inherited context up-front (mode-aware strict gate is the Python TUI's responsibility once the welcome screen records a mode). Bypass with KNOE_SKIP_KUBECONTEXT_GUARD=true for deliberate cross-cluster maintenance. History: the guard was filed in response to the 2026-04-28 14:00 UTC outage — an install.sh --mode k3d run with the shell pointed at GKE replaced the GCS-backed ObjectStore with a Garage-backed one, then Garage filled up and backups silently failed for hours. Closes drift R4 / queue item #1.

Env-contamination guard (live): deploy.sh calls etc/preflight_kubecontext.sh and refuses to proceed when kubectl config current-context doesn't match the [Global] APP_CLUSTER_KUBECONTEXT of the active config. install.sh prints the inherited context up-front (mode-aware strict gate is the Python TUI's responsibility once the welcome screen records a mode). Bypass with KNOE_SKIP_KUBECONTEXT_GUARD=true for deliberate cross-cluster maintenance. History: the guard was filed in response to the 2026-04-28 14:00 UTC outage — an install.sh --mode k3d run with the shell pointed at GKE replaced the GCS-backed ObjectStore with a Garage-backed one, then Garage filled up and backups silently failed for hours. Closes drift R4 / queue item #1.

Get current CNPG node names

kubectl --context=gke_plenary-truck-485623-p7_us-west3_knoe-dev-cnpg-0 get nodes -o name

Key scripts

Script Purpose
scripts/reset_clusters.sh Delete + recreate both GKE clusters (e2-standard-2 × 3 each)
scripts/patch_garage_cross_cluster.sh Remove garage from DB cluster, configure GCS barman ObjectStore
etc/init_cnpg_gke.sh Provision CNPG on GKE with GCS backup via Workload Identity
etc/init_cnpg_backup.sh Configure barman plugin + trigger initial backup

Cluster code constants (knoe/core/actions.py)

DEFAULT_APP_CLUSTER_NAME    = "knoe-dev-0"
DEFAULT_DB_CLUSTER_NAME     = "knoe-dev-cnpg-0"
DEFAULT_APP_CLUSTER_MACHINE_TYPE = "e2-standard-2"
DEFAULT_DB_CLUSTER_MACHINE_TYPE  = "e2-standard-2"

GCP project

  • Project: plenary-truck-485623-p7
  • Region: us-west3
  • Artifact Registry: us-west3-docker.pkg.dev/plenary-truck-485623-p7/knoe-system
  • CPU quota: 16 vCPUs (all regions) — 2 clusters × 3 × e2-standard-2 = 12 vCPUs used
  • SSD quota: SSD_TOTAL_GB = 300 GB — fully consumed by CNPG; all other PVCs must use standard (pd-standard / HDD)

Reality TODOs / Drift Log

Quick reference. Each entry links to the master index where context, owner, and rank live.

# Drift Where described above Where tracked
(none currently)

Closed in 2026-04-29 stabilization session: Garage on DB cluster removed; cluster pods migrated to cnpg-backup-sa via CNPG v1.29.0 spec.serviceAccountName; both operators restarted clean.

Closed 2026-05-01: R4 — installer env-contamination guard now live in deploy.sh (strict) + install.sh (informational notice). Helper at etc/preflight_kubecontext.sh.