prole/monitoring/kps-cnpg-values.yaml
chrisfu 00f0ebec07 Merge claude/crazy-bose-fec256 into main
Bringing the long-running session-feature branch back into main in one
deliberate sweep. The branch carried the cluster work that's been live for
weeks (cross-cluster CNPG metrics, Grafana w/ Google OAuth, supabase
oauth2-proxy, cluster recovery, pg.0.knoe.dev + per-engineer onboarding,
GCS-backed CNPG backups via Workload Identity, the env-contamination
guard, the Junie brief queue, the cnpg-grafana CSRF + memory-request
fixes from today), while main accumulated Junie's parallel knoe-auth
Phase 2 OIDC work (full provider surface: discovery, authorize, token,
userinfo, JWKS, RS256 signing, code exchange, session services).

Key decision: the two branches did COMPETING rebrands off the same
starting point (5ba9b63, 2026-04-27):

  - claude branch (commit b355855, earlier): org.prole.authority.* →
                                              dev.knoe.auth.*
                                              (artifact renamed to
                                              knoe-auth.jar)
  - main (commit 9daa94b, recent): org.prole.authority.* →
                                    dev.knoe.authority.*
                                    (kept "authority" artifact name)

dev.knoe.auth wins: cluster runs from this name, the Maven artifact is
already knoe-auth.jar, and the broader rename is the documented
namespace direction (per ~/.claude/projects/-Users-chrisfu-dev-knoe-db/
memory/MEMORY.md). All of main's recent Phase 2 OIDC content was ported
from authority/src/.../dev/knoe/authority/ into
authority/src/.../dev/knoe/auth/ with package declarations rewritten.

== File-level resolution summary ==

Textual conflicts (4):

  authority/pom.xml
    - Took our artifactId="auth"
    - Took our branch's removal of spring-security-kerberos-client
      (verified: Junie's Phase 2 OIDC code does not import it; the dep
      was already-dead config)

  docs/pipeline-phases.md
    - Took our branch's "Phase 1 not started" status. Main had a
      misplaced " Complete" with a knoe-auth-Phase-1 commit ref
      in the autobuild Phase 1 section — different domain.

  docs/plans/knoe-auth-round-1.md
    - Took our branch's dev.knoe.auth file table (vs main's
      dev.knoe.authority listing). Pure rename mismatch.

  supabase/helm/knoe-supabase/templates/kong/config.yaml
    - Took our branch's onboard route + plain dashboard wiring.
      Main had an oauth2proxy.enabled toggle that put oauth2-proxy as
      a Kong upstream — but the deployed architecture (commit 25f1b2e)
      has oauth2-proxy in FRONT of Kong, not behind. Main's wrapper
      reflected an architecture that was never deployed.
    - Took our branch's removal of basic-auth from dashboard route
      (queue #15 brief still tracks the matching values.yaml /
      kong/deployment.yaml cleanup).

Java tree reconciliation (44 file-pairs):

  20 dual-path source files + 2 dual-path tests
    Body-identical between main's authority/ and our branch's auth/
    after stripping package decls — main's commit 9daa94b was a pure
    rebrand. Took our branch's auth/ version for all 22.

  8 main-only source files (Phase 2 OIDC), ported into auth/:
    web/JwksController.java
    web/OidcAuthorizeController.java
    web/OidcDiscoveryController.java
    web/OidcTokenController.java
    web/OidcUserInfoController.java
    session/OidcCodeService.java
    session/OidcTokenService.java
    session/SessionService.java

  12 main-only test files, ported into auth/:
    HealthControllerTest.java
    enroll/EnrollValueTypesTest.java
    enroll/EnrollmentControllerTest.java
    enroll/TotpServiceTest.java
    kerberos/KadminClientTest.java
    kerberos/KerberosSpnegoResultTest.java
    web/LoginControllerTest.java
    admin/AdminControllerTest.java
    user/PrincipalNormalizerTest.java
    regression/IdentityRegressionTest.java
    session/OidcCodeServiceTest.java
    session/SessionServiceTest.java

  Port mechanics: read main:authority/...<file> via git show, then sed
  rewrite `package dev.knoe.authority` → `package dev.knoe.auth` and
  `import dev.knoe.authority` → `import dev.knoe.auth`. Body content
  unchanged.

  authority/src/main/java/dev/knoe/authority/ — DELETED (duplicate)
  authority/src/test/java/dev/knoe/authority/  — DELETED (duplicate)

== Verification ==

- grep -rln '<<<<<<<' across .java/.md/.yaml/.yml/.sh/.xml/.tpl: clean
- find authority/src -path '*/dev/knoe/authority*': empty (subtree gone)
- grep 'package dev.knoe.authority' across repo: clean
- bash -n install.sh deploy.sh etc/preflight_kubecontext.sh: clean
- git ls-files -u | wc -l: 0 unmerged paths
- helm lint supabase/helm/knoe-supabase: pre-existing failure on
  studioIngress.enabled undefined in values.yaml (introduced by Junie
  on main; unrelated to this merge — flagging as follow-up).

== Followups (carried into TODO ranked queue or noted here) ==

  - helm lint failure: studioIngress block in values.yaml is missing
    enable flag; templates/studio/{ingress,oauth2proxy-deployment,
    oauth2proxy-service}.yaml all reference studioIngress.enabled with
    no default. Pre-existing on main; not introduced by this merge.
  - The five Junie briefs filed on this branch are now reachable from
    main at docs/plans/junie/{02,06,07,13,15}-*.md. Junie can pick them
    up in any order.
  - knoe-auth Phase 2 OIDC source (now at dev.knoe.auth.*) is not yet
    deployed to the cluster. Deployment is its own task.
  - The branch claude/crazy-bose-fec256 stays in place (worktree at
    .claude/worktrees/crazy-bose-fec256 may have ongoing context for
    Claude Code sessions). Safe to delete once next session starts
    cleanly from main.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 16:39:10 -07:00

97 lines
3.4 KiB
YAML

# kube-prometheus-stack (kps) Helm values — DB cluster (knoe-dev-cnpg-0).
#
# This file scopes a SECOND kube-prometheus-stack install — companion to
# `monitoring/kps-values-gke.yaml` (the app-cluster install on knoe-dev-0).
#
# Why two stacks: Kubernetes service discovery is cluster-local. The app-
# cluster Prometheus can't see CNPG pods (they live on knoe-dev-cnpg-0), so
# every cnpg-grafana dashboard rendered "No data". This stack runs Prometheus
# locally on the DB cluster, auto-discovers CNPG's PodMonitor, and is queried
# directly by the app-cluster Grafana via an internal-LB datasource.
#
# Topology:
# knoe-dev-cnpg-0 -[this stack]-> Prometheus + node-exporter + ksm
# |
# v
# ILB (prometheus-cnpg-ilb @ 10.180.x.x:9090)
# ^
# | (datasource: cnpg-prometheus)
# knoe-dev-0 --------------- Grafana (single canonical Grafana)
#
# Apply:
# helm --kube-context=$DB_CLUSTER_KUBECONTEXT upgrade --install kps \
# prometheus-community/kube-prometheus-stack --version 84.3.0 \
# --namespace monitoring --create-namespace \
# -f monitoring/kps-cnpg-values.yaml --wait --timeout 5m
#
# Storage: every PVC pinned to `standard-hdd` (pd-standard). The cluster also
# has `standard-rwo` (pd-balanced, partial-SSD) and `premium-rwo` (pd-ssd) —
# both would consume the SSD_TOTAL_GB regional quota that's shared with CNPG.
# Honor the standing rule: monitoring is HDD-only.
---
# One Grafana stays canonical on the app cluster (svc.knoe.dev/grafana).
# This stack is a metrics-only sidecar.
grafana:
enabled: false
# Alerting routes through the app-cluster Alertmanager (follow-up: federation).
# For now, no AM here, no rules either (avoid scrape-rule noise into a void).
alertmanager:
enabled: false
defaultRules:
create: false
# DB-cluster object metrics (pods/PVCs/Jobs etc.) for kube-state dashboards.
kube-state-metrics:
enabled: true
# Per-DB-node host metrics (CPU/mem/disk/network).
prometheus-node-exporter:
enabled: true
prometheus:
prometheusSpec:
# Short retention — this Prometheus is mostly a query-target. App-cluster
# Grafana queries it on demand; we don't archive long-form.
retention: 7d
# CRITICAL: select ALL PodMonitors / ServiceMonitors / PrometheusRules,
# not just helm-release-labelled ones. The CNPG operator-emitted
# PodMonitor at knoe-db-0/knoe-db doesn't carry our `release=kps` label,
# but we want to scrape it. Selector with empty `{}` matches everything.
podMonitorSelectorNilUsesHelmValues: false
serviceMonitorSelectorNilUsesHelmValues: false
ruleSelectorNilUsesHelmValues: false
podMonitorSelector: {}
serviceMonitorSelector: {}
# Tight resource shape — small cluster, low metric cardinality (3 PG
# pods + ~50 metrics each + 15s scrape ≈ 30k samples/day ≈ 250 MB/day).
resources:
requests:
cpu: 100m
memory: 512Mi
limits:
cpu: 500m
memory: 1Gi
storageSpec:
volumeClaimTemplate:
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 20Gi
storageClassName: standard-hdd # explicit HDD (do not let default fall through)
prometheusOperator:
resources:
requests:
cpu: 50m
memory: 128Mi
limits:
cpu: 200m
memory: 256Mi