mirror of
https://github.com/dredx/prole.git
synced 2026-09-23 10:13:58 +00:00
Bringing the long-running session-feature branch back into main in one deliberate sweep. The branch carried the cluster work that's been live for weeks (cross-cluster CNPG metrics, Grafana w/ Google OAuth, supabase oauth2-proxy, cluster recovery, pg.0.knoe.dev + per-engineer onboarding, GCS-backed CNPG backups via Workload Identity, the env-contamination guard, the Junie brief queue, the cnpg-grafana CSRF + memory-request fixes from today), while main accumulated Junie's parallel knoe-auth Phase 2 OIDC work (full provider surface: discovery, authorize, token, userinfo, JWKS, RS256 signing, code exchange, session services). Key decision: the two branches did COMPETING rebrands off the same starting point (5ba9b63, 2026-04-27): - claude branch (commit b355855, earlier): org.prole.authority.* → dev.knoe.auth.* (artifact renamed to knoe-auth.jar) - main (commit9daa94b, recent): org.prole.authority.* → dev.knoe.authority.* (kept "authority" artifact name) dev.knoe.auth wins: cluster runs from this name, the Maven artifact is already knoe-auth.jar, and the broader rename is the documented namespace direction (per ~/.claude/projects/-Users-chrisfu-dev-knoe-db/ memory/MEMORY.md). All of main's recent Phase 2 OIDC content was ported from authority/src/.../dev/knoe/authority/ into authority/src/.../dev/knoe/auth/ with package declarations rewritten. == File-level resolution summary == Textual conflicts (4): authority/pom.xml - Took our artifactId="auth" - Took our branch's removal of spring-security-kerberos-client (verified: Junie's Phase 2 OIDC code does not import it; the dep was already-dead config) docs/pipeline-phases.md - Took our branch's "Phase 1 not started" status. Main had a misplaced "✅ Complete" with a knoe-auth-Phase-1 commit ref in the autobuild Phase 1 section — different domain. docs/plans/knoe-auth-round-1.md - Took our branch's dev.knoe.auth file table (vs main's dev.knoe.authority listing). Pure rename mismatch. supabase/helm/knoe-supabase/templates/kong/config.yaml - Took our branch's onboard route + plain dashboard wiring. Main had an oauth2proxy.enabled toggle that put oauth2-proxy as a Kong upstream — but the deployed architecture (commit 25f1b2e) has oauth2-proxy in FRONT of Kong, not behind. Main's wrapper reflected an architecture that was never deployed. - Took our branch's removal of basic-auth from dashboard route (queue #15 brief still tracks the matching values.yaml / kong/deployment.yaml cleanup). Java tree reconciliation (44 file-pairs): 20 dual-path source files + 2 dual-path tests Body-identical between main's authority/ and our branch's auth/ after stripping package decls — main's commit9daa94bwas a pure rebrand. Took our branch's auth/ version for all 22. 8 main-only source files (Phase 2 OIDC), ported into auth/: web/JwksController.java web/OidcAuthorizeController.java web/OidcDiscoveryController.java web/OidcTokenController.java web/OidcUserInfoController.java session/OidcCodeService.java session/OidcTokenService.java session/SessionService.java 12 main-only test files, ported into auth/: HealthControllerTest.java enroll/EnrollValueTypesTest.java enroll/EnrollmentControllerTest.java enroll/TotpServiceTest.java kerberos/KadminClientTest.java kerberos/KerberosSpnegoResultTest.java web/LoginControllerTest.java admin/AdminControllerTest.java user/PrincipalNormalizerTest.java regression/IdentityRegressionTest.java session/OidcCodeServiceTest.java session/SessionServiceTest.java Port mechanics: read main:authority/...<file> via git show, then sed rewrite `package dev.knoe.authority` → `package dev.knoe.auth` and `import dev.knoe.authority` → `import dev.knoe.auth`. Body content unchanged. authority/src/main/java/dev/knoe/authority/ — DELETED (duplicate) authority/src/test/java/dev/knoe/authority/ — DELETED (duplicate) == Verification == - grep -rln '<<<<<<<' across .java/.md/.yaml/.yml/.sh/.xml/.tpl: clean - find authority/src -path '*/dev/knoe/authority*': empty (subtree gone) - grep 'package dev.knoe.authority' across repo: clean - bash -n install.sh deploy.sh etc/preflight_kubecontext.sh: clean - git ls-files -u | wc -l: 0 unmerged paths - helm lint supabase/helm/knoe-supabase: pre-existing failure on studioIngress.enabled undefined in values.yaml (introduced by Junie on main; unrelated to this merge — flagging as follow-up). == Followups (carried into TODO ranked queue or noted here) == - helm lint failure: studioIngress block in values.yaml is missing enable flag; templates/studio/{ingress,oauth2proxy-deployment, oauth2proxy-service}.yaml all reference studioIngress.enabled with no default. Pre-existing on main; not introduced by this merge. - The five Junie briefs filed on this branch are now reachable from main at docs/plans/junie/{02,06,07,13,15}-*.md. Junie can pick them up in any order. - knoe-auth Phase 2 OIDC source (now at dev.knoe.auth.*) is not yet deployed to the cluster. Deployment is its own task. - The branch claude/crazy-bose-fec256 stays in place (worktree at .claude/worktrees/crazy-bose-fec256 may have ongoing context for Claude Code sessions). Safe to delete once next session starts cleanly from main. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
97 lines
3.4 KiB
YAML
97 lines
3.4 KiB
YAML
# kube-prometheus-stack (kps) Helm values — DB cluster (knoe-dev-cnpg-0).
|
|
#
|
|
# This file scopes a SECOND kube-prometheus-stack install — companion to
|
|
# `monitoring/kps-values-gke.yaml` (the app-cluster install on knoe-dev-0).
|
|
#
|
|
# Why two stacks: Kubernetes service discovery is cluster-local. The app-
|
|
# cluster Prometheus can't see CNPG pods (they live on knoe-dev-cnpg-0), so
|
|
# every cnpg-grafana dashboard rendered "No data". This stack runs Prometheus
|
|
# locally on the DB cluster, auto-discovers CNPG's PodMonitor, and is queried
|
|
# directly by the app-cluster Grafana via an internal-LB datasource.
|
|
#
|
|
# Topology:
|
|
# knoe-dev-cnpg-0 -[this stack]-> Prometheus + node-exporter + ksm
|
|
# |
|
|
# v
|
|
# ILB (prometheus-cnpg-ilb @ 10.180.x.x:9090)
|
|
# ^
|
|
# | (datasource: cnpg-prometheus)
|
|
# knoe-dev-0 --------------- Grafana (single canonical Grafana)
|
|
#
|
|
# Apply:
|
|
# helm --kube-context=$DB_CLUSTER_KUBECONTEXT upgrade --install kps \
|
|
# prometheus-community/kube-prometheus-stack --version 84.3.0 \
|
|
# --namespace monitoring --create-namespace \
|
|
# -f monitoring/kps-cnpg-values.yaml --wait --timeout 5m
|
|
#
|
|
# Storage: every PVC pinned to `standard-hdd` (pd-standard). The cluster also
|
|
# has `standard-rwo` (pd-balanced, partial-SSD) and `premium-rwo` (pd-ssd) —
|
|
# both would consume the SSD_TOTAL_GB regional quota that's shared with CNPG.
|
|
# Honor the standing rule: monitoring is HDD-only.
|
|
---
|
|
|
|
# One Grafana stays canonical on the app cluster (svc.knoe.dev/grafana).
|
|
# This stack is a metrics-only sidecar.
|
|
grafana:
|
|
enabled: false
|
|
|
|
# Alerting routes through the app-cluster Alertmanager (follow-up: federation).
|
|
# For now, no AM here, no rules either (avoid scrape-rule noise into a void).
|
|
alertmanager:
|
|
enabled: false
|
|
defaultRules:
|
|
create: false
|
|
|
|
# DB-cluster object metrics (pods/PVCs/Jobs etc.) for kube-state dashboards.
|
|
kube-state-metrics:
|
|
enabled: true
|
|
|
|
# Per-DB-node host metrics (CPU/mem/disk/network).
|
|
prometheus-node-exporter:
|
|
enabled: true
|
|
|
|
prometheus:
|
|
prometheusSpec:
|
|
# Short retention — this Prometheus is mostly a query-target. App-cluster
|
|
# Grafana queries it on demand; we don't archive long-form.
|
|
retention: 7d
|
|
|
|
# CRITICAL: select ALL PodMonitors / ServiceMonitors / PrometheusRules,
|
|
# not just helm-release-labelled ones. The CNPG operator-emitted
|
|
# PodMonitor at knoe-db-0/knoe-db doesn't carry our `release=kps` label,
|
|
# but we want to scrape it. Selector with empty `{}` matches everything.
|
|
podMonitorSelectorNilUsesHelmValues: false
|
|
serviceMonitorSelectorNilUsesHelmValues: false
|
|
ruleSelectorNilUsesHelmValues: false
|
|
podMonitorSelector: {}
|
|
serviceMonitorSelector: {}
|
|
|
|
# Tight resource shape — small cluster, low metric cardinality (3 PG
|
|
# pods + ~50 metrics each + 15s scrape ≈ 30k samples/day ≈ 250 MB/day).
|
|
resources:
|
|
requests:
|
|
cpu: 100m
|
|
memory: 512Mi
|
|
limits:
|
|
cpu: 500m
|
|
memory: 1Gi
|
|
|
|
storageSpec:
|
|
volumeClaimTemplate:
|
|
spec:
|
|
accessModes:
|
|
- ReadWriteOnce
|
|
resources:
|
|
requests:
|
|
storage: 20Gi
|
|
storageClassName: standard-hdd # explicit HDD (do not let default fall through)
|
|
|
|
prometheusOperator:
|
|
resources:
|
|
requests:
|
|
cpu: 50m
|
|
memory: 128Mi
|
|
limits:
|
|
cpu: 200m
|
|
memory: 256Mi
|