prole/docs/pipeline-phases.md
chrisfu 00f0ebec07 Merge claude/crazy-bose-fec256 into main
Bringing the long-running session-feature branch back into main in one
deliberate sweep. The branch carried the cluster work that's been live for
weeks (cross-cluster CNPG metrics, Grafana w/ Google OAuth, supabase
oauth2-proxy, cluster recovery, pg.0.knoe.dev + per-engineer onboarding,
GCS-backed CNPG backups via Workload Identity, the env-contamination
guard, the Junie brief queue, the cnpg-grafana CSRF + memory-request
fixes from today), while main accumulated Junie's parallel knoe-auth
Phase 2 OIDC work (full provider surface: discovery, authorize, token,
userinfo, JWKS, RS256 signing, code exchange, session services).

Key decision: the two branches did COMPETING rebrands off the same
starting point (5ba9b63, 2026-04-27):

  - claude branch (commit b355855, earlier): org.prole.authority.* →
                                              dev.knoe.auth.*
                                              (artifact renamed to
                                              knoe-auth.jar)
  - main (commit 9daa94b, recent): org.prole.authority.* →
                                    dev.knoe.authority.*
                                    (kept "authority" artifact name)

dev.knoe.auth wins: cluster runs from this name, the Maven artifact is
already knoe-auth.jar, and the broader rename is the documented
namespace direction (per ~/.claude/projects/-Users-chrisfu-dev-knoe-db/
memory/MEMORY.md). All of main's recent Phase 2 OIDC content was ported
from authority/src/.../dev/knoe/authority/ into
authority/src/.../dev/knoe/auth/ with package declarations rewritten.

== File-level resolution summary ==

Textual conflicts (4):

  authority/pom.xml
    - Took our artifactId="auth"
    - Took our branch's removal of spring-security-kerberos-client
      (verified: Junie's Phase 2 OIDC code does not import it; the dep
      was already-dead config)

  docs/pipeline-phases.md
    - Took our branch's "Phase 1 not started" status. Main had a
      misplaced "✅ Complete" with a knoe-auth-Phase-1 commit ref
      in the autobuild Phase 1 section — different domain.

  docs/plans/knoe-auth-round-1.md
    - Took our branch's dev.knoe.auth file table (vs main's
      dev.knoe.authority listing). Pure rename mismatch.

  supabase/helm/knoe-supabase/templates/kong/config.yaml
    - Took our branch's onboard route + plain dashboard wiring.
      Main had an oauth2proxy.enabled toggle that put oauth2-proxy as
      a Kong upstream — but the deployed architecture (commit 25f1b2e)
      has oauth2-proxy in FRONT of Kong, not behind. Main's wrapper
      reflected an architecture that was never deployed.
    - Took our branch's removal of basic-auth from dashboard route
      (queue #15 brief still tracks the matching values.yaml /
      kong/deployment.yaml cleanup).

Java tree reconciliation (44 file-pairs):

  20 dual-path source files + 2 dual-path tests
    Body-identical between main's authority/ and our branch's auth/
    after stripping package decls — main's commit 9daa94b was a pure
    rebrand. Took our branch's auth/ version for all 22.

  8 main-only source files (Phase 2 OIDC), ported into auth/:
    web/JwksController.java
    web/OidcAuthorizeController.java
    web/OidcDiscoveryController.java
    web/OidcTokenController.java
    web/OidcUserInfoController.java
    session/OidcCodeService.java
    session/OidcTokenService.java
    session/SessionService.java

  12 main-only test files, ported into auth/:
    HealthControllerTest.java
    enroll/EnrollValueTypesTest.java
    enroll/EnrollmentControllerTest.java
    enroll/TotpServiceTest.java
    kerberos/KadminClientTest.java
    kerberos/KerberosSpnegoResultTest.java
    web/LoginControllerTest.java
    admin/AdminControllerTest.java
    user/PrincipalNormalizerTest.java
    regression/IdentityRegressionTest.java
    session/OidcCodeServiceTest.java
    session/SessionServiceTest.java

  Port mechanics: read main:authority/...<file> via git show, then sed
  rewrite `package dev.knoe.authority` → `package dev.knoe.auth` and
  `import dev.knoe.authority` → `import dev.knoe.auth`. Body content
  unchanged.

  authority/src/main/java/dev/knoe/authority/ — DELETED (duplicate)
  authority/src/test/java/dev/knoe/authority/  — DELETED (duplicate)

== Verification ==

- grep -rln '<<<<<<<' across .java/.md/.yaml/.yml/.sh/.xml/.tpl: clean
- find authority/src -path '*/dev/knoe/authority*': empty (subtree gone)
- grep 'package dev.knoe.authority' across repo: clean
- bash -n install.sh deploy.sh etc/preflight_kubecontext.sh: clean
- git ls-files -u | wc -l: 0 unmerged paths
- helm lint supabase/helm/knoe-supabase: pre-existing failure on
  studioIngress.enabled undefined in values.yaml (introduced by Junie
  on main; unrelated to this merge — flagging as follow-up).

== Followups (carried into TODO ranked queue or noted here) ==

  - helm lint failure: studioIngress block in values.yaml is missing
    enable flag; templates/studio/{ingress,oauth2proxy-deployment,
    oauth2proxy-service}.yaml all reference studioIngress.enabled with
    no default. Pre-existing on main; not introduced by this merge.
  - The five Junie briefs filed on this branch are now reachable from
    main at docs/plans/junie/{02,06,07,13,15}-*.md. Junie can pick them
    up in any order.
  - knoe-auth Phase 2 OIDC source (now at dev.knoe.auth.*) is not yet
    deployed to the cluster. Deployment is its own task.
  - The branch claude/crazy-bose-fec256 stays in place (worktree at
    .claude/worktrees/crazy-bose-fec256 may have ongoing context for
    Claude Code sessions). Safe to delete once next session starts
    cleanly from main.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 16:39:10 -07:00

32 KiB
Raw Blame History

Knoe.DB Autobuild & Test Pipeline — Phase Reference

How to use this document Each phase has a Status line, a Commit Summary section (filled in after the phase lands), and a Resumption Checklist so any developer can pick up where we left off without needing context from a previous session. Update the Status and Commit Summary before starting the next phase.


Context

The project has four deployment modes that must each produce a stable, reproducible build:

Mode What it is CI tool
min Single knoe-db container via containerd, no Kubernetes Gitea Actions (bare macOS runner)
k3d k3s-in-Docker local cluster — CNPG + optional Supabase / ArgoCD / Gitea / GitLab Gitea Actions (Docker-enabled runner)
k3s Homelab multi-node k3s on physical hardware Deferred (evaluate after GitLab)
gke Dual GKE clusters on Google Cloud (production) GitLab CI

Primary goal: auto-build on every push to main so we have confidence that builds are stable before we touch anything else.

Secondary goal: once we have stable builds, identify unused code with vulture + coverage diff, then prune toward an instructable codebase — a project a junior developer can read and fully understand.


Phase 0 — Foundation

Status: ✅ Complete
Branch / PRs: landed directly on main; pushed to origin (git-ssh.knoe.dev:knoe.dev/knoe-db)

What this phase does

Fixes the broken test infrastructure so IntelliJ can discover tests and make test actually measures the right code. No new tests are written here — just scaffolding that every subsequent phase depends on.

Deliverables

File Change
pyproject.toml Single source of truth for pytest config + coverage config. Replaces the broken .coveragerc (which pointed to installer/ — a directory that no longer exists)
requirements-test.txt Explicit test dependencies (pytest, coverage, vulture, etc.) separated from runtime deps
.idea/runConfigurations/pytest_all.xml IntelliJ: run all tests with coverage
.idea/runConfigurations/pytest_unit.xml IntelliJ: run unit-only tests (fast, no external deps)
.idea/runConfigurations/pytest_min.xml IntelliJ: run tests tagged @pytest.mark.min
.idea/runConfigurations/pytest_k3d.xml IntelliJ: run tests tagged @pytest.mark.k3d
.idea/runConfigurations/pytest_gke.xml IntelliJ: run tests tagged @pytest.mark.gke
tests/run_tests.sh Fix --cov source from install to knoe

Commit Summary

97575b9 docs: update Phase 0 commit summary and resumption checklist
0052a4d Phase 0: test pipeline foundation — pyproject.toml, IntelliJ run configs, coverage fix, welcome mode selector

The following commits landed on the same branch as part of the Cowork+Code thread consolidation (WIP that had accumulated alongside Phase 0):

837da27 docs: update README; IntelliJ run config picks up Python 3.14 SDK
52fe440 chore(k3s): script and hostprobe updates, temp maintenance scripts
dd8c9d1 chore(scripts): init script updates — gitea, gitlab, kong, monitoring, registry
a754659 chore(deploy): cluster config and k3s manifest updates
4995c86 chore(build): Maven version and authority module updates
ca469aa chore(installer): core Python updates — env, milestones, monitoring
5cd9c12 feat(auth): init scripts and k3s/k8s auth manifests for knoe-auth
2a80df8 feat(auth): land Round 1 — invite-OTP enrollment, kadmin client, GKE manifests
6c72c76 docs(plans): add platform architecture plans — deployment-modes, knoe-auth round 1

Known pre-existing issues (fix in Phase 1, not Phase 0)

knoe/core/ops/cloudnative_pg.py:1372 contains:

f"jsonpath={{.data.{field.replace('.', '\\.')}}}",

Python ≥ 3.12 allows backslashes in f-strings (PEP 701) but this still triggers a SyntaxError on Python 3.14 under certain parse modes. This causes 30 test collection errors — the tests themselves are not broken. Fix: extract the replacement to a variable before the f-string. Tagged as Phase 1 work.

Resumption Checklist

Before picking up work on Phase 1, verify:

  • pytest tests/ collects 441+ tests from the project root
  • pytest -m unit runs and reports a coverage number against knoe/
  • IntelliJ shows the five run configs in the Run/Debug Configurations dropdown
  • python -c "from knoe.ui.screens import KnoeInstaller" succeeds
  • The 30 collection errors in cloudnative_pg.py are logged as Phase 1 work

Off-pipeline work since Phase 0

These commits and live-cluster operations landed between Phase 0 and Phase 1. They aren't part of the autobuild pipeline itself — they're cluster-recovery and Maven build-readiness work that interrupted the Phase 1 sequence. Captured here so future readers see the actual chronology.

2026-04-28 → 2026-04-29: repo cleanup + Maven build readiness + cluster recovery

Commit / event Scope Notes
b355855 refactor(auth): rename to dev.knoe.auth; archive tmp/ post-CNPG cleanup Phase A repo cleanup + Phase C Maven build readiness tmp/ → archive/2026-04-cnpg-and-gitlab-cleanup/ (87 ad-hoc shell scripts); authority/ groupId org.knoe → dev.knoe, artifactId → auth, source tree moved org/prole/authority/ → dev/knoe/auth/; build is now green via make build-auth (mvn -f authority/pom.xml package → authority/target/knoe-auth.jar); 4 k3s manifests renamed prole-auth-*.yaml → knoe-auth-*.yaml.
Phase B live cluster recovery (uncommitted; live-cluster operations) Cluster recovery knoe-db-1 had been crashlooping for 14 days (full WAL volume) — fixed via PVC delete + CNPG re-bootstrap; cluster failover-cycled all 3 pods to refresh stale plugin-barman-cloud sidecars; knoe-dev-cnpg-0 is now 3/3 healthy, primary knoe-db-2.
Garage cluster recovery (uncommitted; live-cluster operations) Backup target recovery Garage on knoe-dev-cnpg-0 was 100% full (LMDB: No space left on device) starting 2026-04-28 14:00 UTC; PVC expanded 29 → 80 GiB, layout capacity 10 → 60 GB, bucket cleared (8382 stale objects + 12 orphan multipart uploads).
GCS migration completed (uncommitted; live-cluster operations) Backup target standardization ObjectStore now gs://knoe-0-backups/ with googleCredentials.gkeEnvironment: true; WI auth initially via auto-generated knoe-db SA annotation; 2 verified data.tar.gz backups in catalog; ScheduledBackup unsuspended.

Suspected root cause of the 14:00 UTC outage: an install.sh --mode k3d run while shell kubectx was set to a GKE context overwrote the GCS-backed ObjectStore with a Garage-backed one, then Garage filled up. Env-contamination guard tracked in docs/TODO.md.

2026-04-29 (later same day): CNPG operator + backup stabilization

After 5 consecutive scheduled backups completed cleanly to GCS, did a small wave to remove the remaining drift:

Event Scope Notes
Operators restarted Refresh stale state barman-cloud (plugin operator) and cnpg-controller-manager rolled — both had been running 15-19 days. Plugin re-discovery confirmed in logs; stale Garage retention-worker endpoint cache cleared.
CNPG operator upgraded v1.28.1 → v1.29.0 Operator upgrade Applied upstream release manifest (releases/cnpg-1.29.0.yaml). v1.29.0 added spec.serviceAccountName field (mutually exclusive with serviceAccountTemplate), required to land the SA migration below. Cluster healthy throughout.
Cluster pod SA migration: knoe-db → cnpg-backup-sa Canonical pattern alignment cluster.spec.serviceAccountName=cnpg-backup-sa applied. CNPG supervised rolling restart cycled all 3 pods. First pod (db-3) failed RBAC check — CNPG only auto-binds roles when it auto-creates the SA; for a pre-existing SA we had to add cnpg-backup-sa as a subject to the existing knoe-db and knoe-db-barman-cloud RoleBindings. After the patch all pods rolled cleanly. 23:03 UTC scheduled backup completed end-to-end with the new SA — verification gate passed.
Garage removed from knoe-dev-cnpg-0 DB-cluster cleanup StatefulSet, both services (garage ClusterIP and garage-s3-ilb LoadBalancer), configmap, PVC, and garage-hdd StorageClass all deleted. ObjectStore untouched (still gs://knoe-0-backups/); backups continued through the Garage tear-down.

Reality-vs-intent gaps surfaced by this work are captured in docs/TODO.md and as > **Reality TODO:** blocks in CLAUDE.md and AGENTS.md. Several gaps closed in this session (Garage on DB cluster, SA alignment); new ones added (redundant [knoe-db-0/knoe-db] IAM binding, init script + patch script need updates to reflect the new canonical end state).


Phase 1 — min Mode Pipeline

Status: 🔲 Not started — ready to begin (deferred during the 2026-04-28 cluster-recovery work above) Depends on: Phase 0 complete ✅ First task: fix knoe/core/ops/cloudnative_pg.py:1372 f-string syntax (30 collection errors, see Phase 0 known issues)

What this phase does

Establishes the first green autobuild. The min pipeline is the simplest possible CI: a bare macOS runner (no Docker, no Kubernetes) running pytest. It exercises the welcome-screen mode selector, the min-mode navigation fast-path, and the init_min.sh script. Trigger: push to main on the Gitea remote.

Architecture

push to main (Gitea)
    │
    ▼
Gitea Actions
    │
    ▼
act_runner  ←── bare macOS (dev machine or Mac mini)
    │            labels: [self-hosted, macos, min]
    ▼            requires: Homebrew + Python 3.11
pytest -m "min or unit"
    │
    ▼
coverage.xml  ──► uploaded as artifact

Runner setup (one-time, not in CI):

# On the macOS runner machine:
brew install gitea-act-runner
act_runner register \
  --instance https://<gitea-host>/  \
  --token    <runner-token-from-gitea-settings> \
  --labels   "self-hosted,macos,min" \
  --name     "knoe-min-runner"
act_runner daemon

Deliverables

File Purpose
.gitea/workflows/ci-min.yml Gitea Actions workflow — push-triggered, runs on bare macOS runner
tests/modes/__init__.py Package marker
tests/modes/test_min_mode.py Mode-specific tests for min path
tests/modes/conftest.py Mode fixtures (mock containerd, mock 1Password CLI)

.gitea/workflows/ci-min.yml

name: CI — min mode

on:
  push:
    branches: [main]

jobs:
  min-unit:
    runs-on: [self-hosted, macos, min]
    timeout-minutes: 10
    steps:
      - uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"

      - name: Install dependencies
        run: |
          pip install -r requirements.txt -r requirements-test.txt          

      - name: Run min + unit tests
        run: |
          pytest -m "min or unit" \
            --cov=knoe \
            --cov-report=xml \
            --cov-report=term-missing \
            -v \
            tests/          
        env:
          PYTHONPATH: ${{ github.workspace }}
          KNOE_MODE: min

      - name: Upload coverage artifact
        uses: actions/upload-artifact@v4
        with:
          name: coverage-min
          path: coverage.xml
          retention-days: 14

Tests to write in tests/modes/test_min_mode.py

Each test name is intentionally human-readable — they become living documentation for junior developers.

Test What it proves
test_welcome_mode_selector_starts_with_no_selection deployment_mode defaults to "" on fresh install — forces explicit choice
test_clicking_min_card_sets_cluster_env_to_min Card click wires cluster_env = "min" correctly
test_clicking_min_card_enables_next_button Welcome Next is disabled until a card is clicked
test_min_nav_env_setup_goes_to_init_password_not_cluster Verifies min fast-path in on_next
test_min_nav_init_scripts_goes_to_security_not_kerberos Min skips Kerberos, ArgoCD, GitOps, Supabase
test_min_nav_prev_from_init_password_returns_to_env_setup Symmetric back-navigation
test_min_nav_prev_from_security_returns_to_init_scripts Symmetric back-navigation
test_normalize_cluster_env_min_returns_min _normalize_cluster_env("min") == "min"
test_deployment_mode_from_env_min_returns_min _deployment_mode_from_env("min") == "min"
test_init_min_sh_runs_without_error subprocess call to init_min.sh initialize exits 0 (macOS only, @pytest.mark.min)

Commit Summary

(Fill in after phase lands)

# example:
# abc1234 Phase 1: add .gitea/workflows/ci-min.yml
# def5678 Phase 1: add tests/modes/test_min_mode.py
# ghi9012 Phase 1: add tests/modes/conftest.py

Resumption Checklist

  • Push to main triggers ci-min workflow in Gitea Actions UI
  • pytest -m "min or unit" exits 0 locally
  • All 10 tests in test_min_mode.py pass
  • Coverage artifact appears in the Gitea Actions run summary
  • Badge shows green on Gitea repo homepage (optional but nice)

Phase 2 — k3d Mode Pipeline

Status: 🔲 Not started
Depends on: Phase 1 complete

What this phase does

Adds the k3d integration pipeline — a Docker-enabled runner that creates a real k3d cluster, deploys CNPG, verifies the database is reachable, then tears it down. This is the first pipeline that proves an actual database cluster starts correctly. Estimated runtime: 8–15 minutes.

Architecture

push to main (Gitea)
    │
    ▼
Gitea Actions
    │
    ├─ job: k3d-unit (fast, no cluster) ─────────────► coverage-k3d-unit.xml
    │
    └─ job: k3d-integration (depends on unit) ──────► coverage-k3d-integration.xml
           │
           ├── k3d cluster create knoe-ci
           ├── apply CNPG operator
           ├── apply knoe-db Cluster CR
           ├── pytest -m "k3d and integration"
           └── k3d cluster delete knoe-ci (always)

Runner setup (one-time):
Same act_runner binary but registered with labels self-hosted,macos,k3d on a Docker-enabled machine (Docker Desktop or Colima).

Deliverables

File Purpose
.gitea/workflows/ci-k3d.yml Two-job workflow: unit then integration
tests/modes/test_k3d_mode.py k3d unit + integration tests
tests/modes/fixtures/knoe-db-test.yaml Minimal CNPG Cluster CR for test use (single instance, tiny storage)
tests/modes/fixtures/cnpg-operator-values.yaml Minimal Helm values for CNPG operator in CI

.gitea/workflows/ci-k3d.yml

name: CI — k3d mode

on:
  push:
    branches: [main]

jobs:
  k3d-unit:
    runs-on: [self-hosted, macos, k3d]
    timeout-minutes: 5
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.11" }
      - run: pip install -r requirements.txt -r requirements-test.txt
      - name: Unit tests (k3d, no cluster)
        run: pytest -m "k3d and unit" --cov=knoe --cov-report=xml tests/
        env:
          PYTHONPATH: ${{ github.workspace }}
          KNOE_MODE: k3d
      - uses: actions/upload-artifact@v4
        with: { name: coverage-k3d-unit, path: coverage.xml }

  k3d-integration:
    runs-on: [self-hosted, macos, k3d]
    needs: k3d-unit
    timeout-minutes: 25
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.11" }
      - run: pip install -r requirements.txt -r requirements-test.txt

      - name: Install k3d + kubectl
        run: |
          brew install k3d kubectl          

      - name: Create test cluster
        run: |
          k3d cluster create knoe-ci \
            --agents 1 \
            --k3s-arg '--disable=traefik@server:0' \
            --wait
          kubectl cluster-info --context k3d-knoe-ci          

      - name: Deploy CNPG operator
        run: |
          kubectl apply --server-side \
            -f https://raw.githubusercontent.com/cloudnative-pg/cloudnative-pg/main/releases/cnpg-1.23.0.yaml
          kubectl wait --for=condition=Available \
            deployment/cnpg-controller-manager \
            -n cnpg-system --timeout=120s          

      - name: Integration tests (live cluster)
        run: |
          pytest -m "k3d and integration" \
            --cov=knoe \
            --cov-report=xml \
            -v \
            tests/          
        env:
          PYTHONPATH: ${{ github.workspace }}
          KNOE_MODE: k3d
          KUBECONFIG: ${{ env.HOME }}/.kube/config

      - name: Tear down cluster
        if: always()
        run: k3d cluster delete knoe-ci

      - uses: actions/upload-artifact@v4
        with: { name: coverage-k3d-integration, path: coverage.xml }

Tests to write in tests/modes/test_k3d_mode.py

Unit tests (@pytest.mark.k3d @pytest.mark.unit):

Test What it proves
test_k3d_cluster_env_normalizes_to_dev _normalize_cluster_env("dev") == "dev"
test_welcome_k3d_card_sets_cluster_env_dev Card click → cluster_env = "dev"
test_k3d_nav_proceeds_through_cluster_screens k3d mode does NOT skip init_cluster
test_k3d_supabase_option_visible Supabase toggle is reachable in k3d nav flow
test_cluster_lifecycle_milestone_uses_k3d_script ClusterLifecycleMilestone calls k3d cluster-create command for dev env

Integration tests (@pytest.mark.k3d @pytest.mark.integration):

Test What it proves
test_k3d_cluster_api_is_reachable kubectl can reach k3d-knoe-ci API server
test_cnpg_operator_crds_registered Cluster CRD exists after operator deploy
test_knoe_db_cluster_pod_starts Apply test CR → at least one PostgreSQL pod reaches Running
test_database_accepts_connections psql SELECT 1 succeeds against the CNPG service
test_cluster_delete_is_clean After delete, no k3d cluster named knoe-ci remains

Commit Summary

(Fill in after phase lands)

Resumption Checklist

  • Push to main triggers both k3d-unit and k3d-integration jobs in Gitea
  • k3d-integration passes (PostgreSQL pod reaches Running state)
  • Both coverage artifacts appear in the run summary
  • k3d cluster list shows no leftover knoe-ci cluster after the run

Phase 3 — GKE Pipeline (GitLab CI, expanded)

Status: 🔲 Not started
Depends on: Phase 2 complete

What this phase does

Replaces the current single-job .gitlab-ci.yml (which only runs install.sh -S) with a proper multi-stage pipeline: lint → unit tests → GKE integration tests → dead-code report → deploy. The unit test gate runs on every push; GKE integration runs on main only.

Architecture

push to any branch
    │
    ├─ stage: lint     → python syntax + imports check
    ├─ stage: test     → pytest -m "unit"  (every push)
    │
push to main only:
    ├─ stage: test     → pytest -m "gke and integration"  (real GKE cluster)
    ├─ stage: test     → vulture dead-code report (allow_failure: true)
    └─ stage: deploy   → ./install.sh -S -c conf/service/prod.cfg

Deliverables

File Purpose
.gitlab-ci.yml Full multi-stage pipeline replacing current single-job version
tests/modes/test_gke_mode.py GKE unit + integration tests
scripts/dead_code_analysis.py Cross-references coverage.json with vulture-report.txt, outputs three-tier report

Required GitLab CI variables (project settings → CI/CD → Variables)

Variable Value Protected Masked
GCP_SA_KEY GCP service account JSON with container.viewer + storage.objectViewer on both clusters ✅ ✅
GKE_REGION us-west3
GKE_APP_CLUSTER knoe-dev-0
GKE_DB_CLUSTER knoe-dev-cnpg-0

.gitlab-ci.yml (full replacement)

stages:
  - lint
  - test
  - report
  - deploy

variables:
  PYTHON_VERSION: "3.11"
  PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"

cache:
  paths: [.cache/pip]

# ── Stage: lint ───────────────────────────────────────────────────────────────

lint:
  stage: lint
  image: python:3.11-slim
  script:
    - pip install -r requirements-test.txt -q
    - python -m py_compile knoe/**/*.py
    - python -c "from knoe.ui.screens import KnoeInstaller"
  rules:
    - if: '$CI_PIPELINE_SOURCE == "push"'

# ── Stage: test (unit — every push) ──────────────────────────────────────────

unit-tests:
  stage: test
  image: python:3.11-slim
  script:
    - pip install -r requirements.txt -r requirements-test.txt -q
    - pytest -m "unit and not integration"
        --cov=knoe
        --cov-report=xml
        --cov-report=term-missing
        -q
        tests/
  coverage: '/TOTAL.*\s+(\d+\%)/'
  artifacts:
    reports:
      coverage_report:
        coverage_format: cobertura
        path: coverage.xml
    paths: [coverage.xml]
    expire_in: 7 days
  rules:
    - if: '$CI_PIPELINE_SOURCE == "push"'

# ── Stage: test (GKE integration — main only) ─────────────────────────────────

gke-integration:
  stage: test
  image: google/cloud-sdk:slim
  timeout: 30 minutes
  before_script:
    - pip install -r requirements.txt -r requirements-test.txt -q
    - echo "$GCP_SA_KEY" | gcloud auth activate-service-account --key-file=-
    - gcloud config set project plenary-truck-485623-p7
    - gcloud container clusters get-credentials $GKE_APP_CLUSTER
        --region $GKE_REGION
    - gcloud container clusters get-credentials $GKE_DB_CLUSTER
        --region $GKE_REGION
  script:
    - pytest -m "gke and integration"
        --cov=knoe
        --cov-report=xml
        -v
        tests/
  coverage: '/TOTAL.*\s+(\d+\%)/'
  artifacts:
    reports:
      coverage_report:
        coverage_format: cobertura
        path: coverage.xml
    paths: [coverage.xml]
    expire_in: 30 days
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'

# ── Stage: report (dead code — main only, never blocks build) ─────────────────

dead-code:
  stage: report
  image: python:3.11-slim
  allow_failure: true
  script:
    - pip install vulture -q
    - vulture knoe/ --min-confidence 80 | tee vulture-report.txt
    - python scripts/dead_code_analysis.py
        --vulture vulture-report.txt
        --coverage coverage.xml
        --output dead-code-report.md
  artifacts:
    paths:
      - vulture-report.txt
      - dead-code-report.md
    expire_in: 30 days
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'

# ── Stage: deploy (main only) ─────────────────────────────────────────────────

deploy-service:
  stage: deploy
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'
  script:
    - ./install.sh -S -c conf/service/prod.cfg
  environment:
    name: production

Tests to write in tests/modes/test_gke_mode.py

Unit tests (@pytest.mark.gke @pytest.mark.unit):

Test What it proves
test_gke_cluster_env_normalizes_to_prod _normalize_cluster_env("prod") == "prod"
test_welcome_gke_card_sets_cluster_env_prod Card click → cluster_env = "prod"
test_gke_split_cluster_detected When app_ctx != db_ctx, split-cluster mode activates
test_garage_not_deployed_to_db_cluster CNPG cluster does not include Garage in GKE mode
test_gke_storage_class_validation standard-rwo raises quota error hint; standard passes

Integration tests (@pytest.mark.gke @pytest.mark.integration):

Test What it proves
test_gke_app_cluster_reachable kubectl can reach knoe-dev-0
test_gke_db_cluster_reachable kubectl can reach knoe-dev-cnpg-0
test_cnpg_cluster_knoe_db_0_running knoe-db cluster in knoe-db-0 namespace has 3 Ready instances
test_barman_backup_schedule_exists ScheduledBackup CR exists in knoe-db-0
test_garage_only_in_app_cluster No Garage pods in knoe-dev-cnpg-0 namespace

Commit Summary

(Fill in after phase lands)

Resumption Checklist

  • GitLab pipeline shows four stages: lint → test → report → deploy
  • unit-tests job passes on every push (not just main)
  • gke-integration passes on main (3 CNPG pods Running)
  • dead-code job produces a dead-code-report.md artifact
  • Deploy stage still works (install.sh -S exits 0)
  • GitLab shows coverage percentage on the merge request widget

Phase 4 — Dead Code Identification

Status: 🔲 Not started
Depends on: Phase 3 complete (all three stable builds achieved)

What this phase does

Runs the first systematic dead code analysis across the full codebase. This is not a cleanup sprint — it is reconnaissance. We generate a prioritised report and review it before deleting anything. The report becomes the input for Phase 5.

Why this matters

The project has undergone several major renames (prole → knoe, various directory restructures). Code written for old layouts is likely still present. knoe/core/actions.py alone is 7,740 lines and almost certainly contains branches that no deployment mode ever exercises.

Tools

Tool Role
vulture Static analysis — finds unused functions, classes, imports
pytest --cov --cov-branch Dynamic analysis — lines/branches never executed during tests
scripts/dead_code_analysis.py Cross-references both outputs to produce a tiered report

Three-tier output format

## 🔴 Definite dead code  (vulture-flagged AND 0% coverage)
- knoe/core/ops/legacy_shell.py:47 — function `_old_prole_exec` (unused, 0% coverage)
- ...

## 🟡 Suspect  (vulture-flagged OR 0% coverage, not both)
- knoe/core/actions.py:4201 — function `_k8s_node_drain` (0% coverage, not flagged by vulture)
- ...

## 🟢 Live  (covered by tests and vulture-clean)
- (omitted from report for brevity)

How to run locally

# 1. Full coverage run across all modes
KNOE_MODE=min  pytest -m min  --cov=knoe --cov-append tests/
KNOE_MODE=k3d  pytest -m k3d  --cov=knoe --cov-append tests/
KNOE_MODE=gke  pytest -m gke  --cov=knoe --cov-append tests/
coverage json  # produces coverage.json

# 2. Vulture scan
vulture knoe/ --min-confidence 70 > vulture-report.txt

# 3. Cross-reference
python scripts/dead_code_analysis.py \
  --vulture vulture-report.txt \
  --coverage coverage.json \
  --output dead-code-report.md

# 4. Review
open dead-code-report.md

Deliverables

File Purpose
scripts/dead_code_analysis.py Cross-reference script (coverage.json + vulture output → tiered report)
dead-code-report.md Generated output — committed to repo for review, not production
docs/dead-code-review.md Human review notes — which 🔴 items are safe to delete vs accidentally flagged

Commit Summary

(Fill in after phase lands)

Resumption Checklist

  • python scripts/dead_code_analysis.py runs without error
  • dead-code-report.md contains all three tiers
  • The 🔴 list has been manually reviewed and each item categorised as "safe to delete" or "keep" in docs/dead-code-review.md
  • No production code deleted yet — this phase is report-only

Phase 5 — Pruning and the Instructable Codebase

Status: 🔲 Not started
Depends on: Phase 4 report reviewed and approved

What this phase does

This is the refactoring sprint. Using the Phase 4 report, we delete dead code in small, test-verified commits, raising the fail_under coverage threshold after each deletion. The end state is a codebase that a junior developer can read from top to bottom and fully understand.

Target metrics

Metric Today (est.) Target
Test coverage (unit) ~40% ≥ 80%
Test coverage (integration) ~10% ≥ 60%
vulture --min-confidence 80 warnings Unknown 0
Lines in knoe/core/actions.py 7,740 < 2,000
Modules with zero test coverage ~17 (knoe/core/ops/) 0
fail_under in pyproject.toml 0 75

Deletion protocol

For each 🔴 item from the Phase 4 report:

  1. Write a test that would fail if the code were still needed (proves it's safe to delete)
  2. Delete the code
  3. Run pytest — all tests pass
  4. Commit with message: prune: remove <module/function> — dead code (vulture + 0% coverage)
  5. Raise fail_under by 1–2 points

This protocol means every deletion is backed by a test. The test suite grows because we prune.

The instructable codebase standard

A module is "instructable" when:

  • Every public function has at least one test whose name reads as a plain English sentence
  • The test file for the module is shorter than the module itself
  • A junior developer can understand the module's purpose from the test names alone without reading the source

Example of instructable test names:

def test_min_mode_only_needs_homebrew_and_1password(): ...
def test_k3d_mode_creates_cnpg_cluster_in_docker(): ...
def test_welcome_card_click_routes_to_correct_nav_flow(): ...
def test_kerberos_is_only_enabled_in_k3s_and_gke_modes(): ...

Commit Summary

(Fill in after each deletion sprint)

Resumption Checklist

  • fail_under has been raised at least once since Phase 4
  • Zero 🔴 items remain in dead-code-report.md
  • All modules in knoe/core/ops/ have at least one test
  • CI passes on all three pipelines (min, k3d, gke)

k3s Phase — Homelab Pipeline

Status: ⏸ Deferred
Trigger: Evaluate after GitLab CI (Phase 3) is stable

Notes for when this is ready

  • Runner: one of the physical k3s nodes (myrddin.prole.org, gandalf.prole.org, or merlin.prole.org) registered as a Gitea act_runner with labels self-hosted,linux,k3s
  • Integration tests will need kubeconfig for the k3s cluster and iSCSI storage access for CNPG
  • Kerberos integration is the unique test target here — knoe-auth deployment with embedded KDC
  • The CNPG_ELIGIBLE_NODES config var maps directly to tests that verify node affinity placement

Appendix A — Pytest Marker Reference

Marker When to use
@pytest.mark.unit No external services, no file system writes, runs anywhere
@pytest.mark.integration Requires a running cluster, database, or network service
@pytest.mark.min Tests the min (containerd) deployment path
@pytest.mark.k3d Tests the k3d (local Docker cluster) deployment path
@pytest.mark.k3s Tests the k3s (homelab) deployment path
@pytest.mark.gke Tests the GKE (production) deployment path

Combining markers is the norm: @pytest.mark.k3d @pytest.mark.integration means "requires a live k3d cluster".


Appendix B — Runner Registration Quick Reference

Gitea act_runner (macOS)

# Install
brew install act-runner

# Register (run once per machine)
act_runner register \
  --instance  https://<gitea-host>/ \
  --token     <token-from-gitea-settings-actions-runners> \
  --labels    "self-hosted,macos,min"   # or min,k3d for the Docker machine \
  --name      "knoe-min-runner"         # human-readable name in Gitea UI

# Start as a service
brew services start act-runner

GitLab runner (existing)

The existing GitLab runner is assumed to have gcloud CLI available. If not:

# On the runner host:
curl https://sdk.cloud.google.com | bash
gcloud components install gke-gcloud-auth-plugin

Appendix C — Coverage Increment Strategy

Rather than setting an ambitious fail_under up front and having CI permanently broken, we raise it in steps as each phase lands:

After phase fail_under
Phase 0 (foundation) 0 (measure only)
Phase 1 (min tests) 15
Phase 2 (k3d tests) 25
Phase 3 (gke tests) 35
Phase 4 (dead code pruned) 50
Phase 5 (full prune) 75

Each increment is a one-line change to pyproject.toml committed at the end of the phase.