Bringing the long-running session-feature branch back into main in one deliberate sweep. The branch carried the cluster work that's been live for weeks (cross-cluster CNPG metrics, Grafana w/ Google OAuth, supabase oauth2-proxy, cluster recovery, pg.0.knoe.dev + per-engineer onboarding, GCS-backed CNPG backups via Workload Identity, the env-contamination guard, the Junie brief queue, the cnpg-grafana CSRF + memory-request fixes from today), while main accumulated Junie's parallel knoe-auth Phase 2 OIDC work (full provider surface: discovery, authorize, token, userinfo, JWKS, RS256 signing, code exchange, session services). Key decision: the two branches did COMPETING rebrands off the same starting point (5ba9b63, 2026-04-27): - claude branch (commit b355855, earlier): org.prole.authority.* → dev.knoe.auth.* (artifact renamed to knoe-auth.jar) - main (commit9daa94b, recent): org.prole.authority.* → dev.knoe.authority.* (kept "authority" artifact name) dev.knoe.auth wins: cluster runs from this name, the Maven artifact is already knoe-auth.jar, and the broader rename is the documented namespace direction (per ~/.claude/projects/-Users-chrisfu-dev-knoe-db/ memory/MEMORY.md). All of main's recent Phase 2 OIDC content was ported from authority/src/.../dev/knoe/authority/ into authority/src/.../dev/knoe/auth/ with package declarations rewritten. == File-level resolution summary == Textual conflicts (4): authority/pom.xml - Took our artifactId="auth" - Took our branch's removal of spring-security-kerberos-client (verified: Junie's Phase 2 OIDC code does not import it; the dep was already-dead config) docs/pipeline-phases.md - Took our branch's "Phase 1 not started" status. Main had a misplaced "✅ Complete" with a knoe-auth-Phase-1 commit ref in the autobuild Phase 1 section — different domain. docs/plans/knoe-auth-round-1.md - Took our branch's dev.knoe.auth file table (vs main's dev.knoe.authority listing). Pure rename mismatch. supabase/helm/knoe-supabase/templates/kong/config.yaml - Took our branch's onboard route + plain dashboard wiring. Main had an oauth2proxy.enabled toggle that put oauth2-proxy as a Kong upstream — but the deployed architecture (commit 25f1b2e) has oauth2-proxy in FRONT of Kong, not behind. Main's wrapper reflected an architecture that was never deployed. - Took our branch's removal of basic-auth from dashboard route (queue #15 brief still tracks the matching values.yaml / kong/deployment.yaml cleanup). Java tree reconciliation (44 file-pairs): 20 dual-path source files + 2 dual-path tests Body-identical between main's authority/ and our branch's auth/ after stripping package decls — main's commit9daa94bwas a pure rebrand. Took our branch's auth/ version for all 22. 8 main-only source files (Phase 2 OIDC), ported into auth/: web/JwksController.java web/OidcAuthorizeController.java web/OidcDiscoveryController.java web/OidcTokenController.java web/OidcUserInfoController.java session/OidcCodeService.java session/OidcTokenService.java session/SessionService.java 12 main-only test files, ported into auth/: HealthControllerTest.java enroll/EnrollValueTypesTest.java enroll/EnrollmentControllerTest.java enroll/TotpServiceTest.java kerberos/KadminClientTest.java kerberos/KerberosSpnegoResultTest.java web/LoginControllerTest.java admin/AdminControllerTest.java user/PrincipalNormalizerTest.java regression/IdentityRegressionTest.java session/OidcCodeServiceTest.java session/SessionServiceTest.java Port mechanics: read main:authority/...<file> via git show, then sed rewrite `package dev.knoe.authority` → `package dev.knoe.auth` and `import dev.knoe.authority` → `import dev.knoe.auth`. Body content unchanged. authority/src/main/java/dev/knoe/authority/ — DELETED (duplicate) authority/src/test/java/dev/knoe/authority/ — DELETED (duplicate) == Verification == - grep -rln '<<<<<<<' across .java/.md/.yaml/.yml/.sh/.xml/.tpl: clean - find authority/src -path '*/dev/knoe/authority*': empty (subtree gone) - grep 'package dev.knoe.authority' across repo: clean - bash -n install.sh deploy.sh etc/preflight_kubecontext.sh: clean - git ls-files -u | wc -l: 0 unmerged paths - helm lint supabase/helm/knoe-supabase: pre-existing failure on studioIngress.enabled undefined in values.yaml (introduced by Junie on main; unrelated to this merge — flagging as follow-up). == Followups (carried into TODO ranked queue or noted here) == - helm lint failure: studioIngress block in values.yaml is missing enable flag; templates/studio/{ingress,oauth2proxy-deployment, oauth2proxy-service}.yaml all reference studioIngress.enabled with no default. Pre-existing on main; not introduced by this merge. - The five Junie briefs filed on this branch are now reachable from main at docs/plans/junie/{02,06,07,13,15}-*.md. Junie can pick them up in any order. - knoe-auth Phase 2 OIDC source (now at dev.knoe.auth.*) is not yet deployed to the cluster. Deployment is its own task. - The branch claude/crazy-bose-fec256 stays in place (worktree at .claude/worktrees/crazy-bose-fec256 may have ongoing context for Claude Code sessions). Safe to delete once next session starts cleanly from main. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
32 KiB
Knoe.DB Autobuild & Test Pipeline — Phase Reference
How to use this document Each phase has a Status line, a Commit Summary section (filled in after the phase lands), and a Resumption Checklist so any developer can pick up where we left off without needing context from a previous session. Update the Status and Commit Summary before starting the next phase.
Context
The project has four deployment modes that must each produce a stable, reproducible build:
| Mode | What it is | CI tool |
|---|---|---|
min |
Single knoe-db container via containerd, no Kubernetes |
Gitea Actions (bare macOS runner) |
k3d |
k3s-in-Docker local cluster — CNPG + optional Supabase / ArgoCD / Gitea / GitLab | Gitea Actions (Docker-enabled runner) |
k3s |
Homelab multi-node k3s on physical hardware | Deferred (evaluate after GitLab) |
gke |
Dual GKE clusters on Google Cloud (production) | GitLab CI |
Primary goal: auto-build on every push to main so we have confidence that builds are stable before we touch anything else.
Secondary goal: once we have stable builds, identify unused code with vulture + coverage diff, then prune toward an instructable codebase — a project a junior developer can read and fully understand.
Phase 0 — Foundation
Status: ✅ Complete
Branch / PRs: landed directly on main; pushed to origin (git-ssh.knoe.dev:knoe.dev/knoe-db)
What this phase does
Fixes the broken test infrastructure so IntelliJ can discover tests and make test actually measures the right code. No new tests are written here — just scaffolding that every subsequent phase depends on.
Deliverables
| File | Change |
|---|---|
pyproject.toml |
Single source of truth for pytest config + coverage config. Replaces the broken .coveragerc (which pointed to installer/ — a directory that no longer exists) |
requirements-test.txt |
Explicit test dependencies (pytest, coverage, vulture, etc.) separated from runtime deps |
.idea/runConfigurations/pytest_all.xml |
IntelliJ: run all tests with coverage |
.idea/runConfigurations/pytest_unit.xml |
IntelliJ: run unit-only tests (fast, no external deps) |
.idea/runConfigurations/pytest_min.xml |
IntelliJ: run tests tagged @pytest.mark.min |
.idea/runConfigurations/pytest_k3d.xml |
IntelliJ: run tests tagged @pytest.mark.k3d |
.idea/runConfigurations/pytest_gke.xml |
IntelliJ: run tests tagged @pytest.mark.gke |
tests/run_tests.sh |
Fix --cov source from install to knoe |
Commit Summary
97575b9 docs: update Phase 0 commit summary and resumption checklist
0052a4d Phase 0: test pipeline foundation — pyproject.toml, IntelliJ run configs, coverage fix, welcome mode selector
The following commits landed on the same branch as part of the Cowork+Code thread consolidation (WIP that had accumulated alongside Phase 0):
837da27 docs: update README; IntelliJ run config picks up Python 3.14 SDK
52fe440 chore(k3s): script and hostprobe updates, temp maintenance scripts
dd8c9d1 chore(scripts): init script updates — gitea, gitlab, kong, monitoring, registry
a754659 chore(deploy): cluster config and k3s manifest updates
4995c86 chore(build): Maven version and authority module updates
ca469aa chore(installer): core Python updates — env, milestones, monitoring
5cd9c12 feat(auth): init scripts and k3s/k8s auth manifests for knoe-auth
2a80df8 feat(auth): land Round 1 — invite-OTP enrollment, kadmin client, GKE manifests
6c72c76 docs(plans): add platform architecture plans — deployment-modes, knoe-auth round 1
Known pre-existing issues (fix in Phase 1, not Phase 0)
knoe/core/ops/cloudnative_pg.py:1372 contains:
f"jsonpath={{.data.{field.replace('.', '\\.')}}}",
Python ≥ 3.12 allows backslashes in f-strings (PEP 701) but this still triggers a SyntaxError on Python 3.14 under certain parse modes. This causes 30 test collection errors — the tests themselves are not broken. Fix: extract the replacement to a variable before the f-string. Tagged as Phase 1 work.
Resumption Checklist
Before picking up work on Phase 1, verify:
pytest tests/collects 441+ tests from the project rootpytest -m unitruns and reports a coverage number againstknoe/- IntelliJ shows the five run configs in the Run/Debug Configurations dropdown
python -c "from knoe.ui.screens import KnoeInstaller"succeeds- The 30 collection errors in
cloudnative_pg.pyare logged as Phase 1 work
Off-pipeline work since Phase 0
These commits and live-cluster operations landed between Phase 0 and Phase 1. They aren't part of the autobuild pipeline itself — they're cluster-recovery and Maven build-readiness work that interrupted the Phase 1 sequence. Captured here so future readers see the actual chronology.
2026-04-28 → 2026-04-29: repo cleanup + Maven build readiness + cluster recovery
| Commit / event | Scope | Notes |
|---|---|---|
b355855 refactor(auth): rename to dev.knoe.auth; archive tmp/ post-CNPG cleanup |
Phase A repo cleanup + Phase C Maven build readiness | tmp/ → archive/2026-04-cnpg-and-gitlab-cleanup/ (87 ad-hoc shell scripts); authority/ groupId org.knoe → dev.knoe, artifactId → auth, source tree moved org/prole/authority/ → dev/knoe/auth/; build is now green via make build-auth (mvn -f authority/pom.xml package → authority/target/knoe-auth.jar); 4 k3s manifests renamed prole-auth-*.yaml → knoe-auth-*.yaml. |
| Phase B live cluster recovery (uncommitted; live-cluster operations) | Cluster recovery | knoe-db-1 had been crashlooping for 14 days (full WAL volume) — fixed via PVC delete + CNPG re-bootstrap; cluster failover-cycled all 3 pods to refresh stale plugin-barman-cloud sidecars; knoe-dev-cnpg-0 is now 3/3 healthy, primary knoe-db-2. |
| Garage cluster recovery (uncommitted; live-cluster operations) | Backup target recovery | Garage on knoe-dev-cnpg-0 was 100% full (LMDB: No space left on device) starting 2026-04-28 14:00 UTC; PVC expanded 29 → 80 GiB, layout capacity 10 → 60 GB, bucket cleared (8382 stale objects + 12 orphan multipart uploads). |
| GCS migration completed (uncommitted; live-cluster operations) | Backup target standardization | ObjectStore now gs://knoe-0-backups/ with googleCredentials.gkeEnvironment: true; WI auth initially via auto-generated knoe-db SA annotation; 2 verified data.tar.gz backups in catalog; ScheduledBackup unsuspended. |
Suspected root cause of the 14:00 UTC outage: an install.sh --mode k3d run while shell kubectx was set to a GKE context overwrote the GCS-backed ObjectStore with a Garage-backed one, then Garage filled up. Env-contamination guard tracked in docs/TODO.md.
2026-04-29 (later same day): CNPG operator + backup stabilization
After 5 consecutive scheduled backups completed cleanly to GCS, did a small wave to remove the remaining drift:
| Event | Scope | Notes |
|---|---|---|
| Operators restarted | Refresh stale state | barman-cloud (plugin operator) and cnpg-controller-manager rolled — both had been running 15-19 days. Plugin re-discovery confirmed in logs; stale Garage retention-worker endpoint cache cleared. |
| CNPG operator upgraded v1.28.1 → v1.29.0 | Operator upgrade | Applied upstream release manifest (releases/cnpg-1.29.0.yaml). v1.29.0 added spec.serviceAccountName field (mutually exclusive with serviceAccountTemplate), required to land the SA migration below. Cluster healthy throughout. |
| Cluster pod SA migration: knoe-db → cnpg-backup-sa | Canonical pattern alignment | cluster.spec.serviceAccountName=cnpg-backup-sa applied. CNPG supervised rolling restart cycled all 3 pods. First pod (db-3) failed RBAC check — CNPG only auto-binds roles when it auto-creates the SA; for a pre-existing SA we had to add cnpg-backup-sa as a subject to the existing knoe-db and knoe-db-barman-cloud RoleBindings. After the patch all pods rolled cleanly. 23:03 UTC scheduled backup completed end-to-end with the new SA — verification gate passed. |
Garage removed from knoe-dev-cnpg-0 |
DB-cluster cleanup | StatefulSet, both services (garage ClusterIP and garage-s3-ilb LoadBalancer), configmap, PVC, and garage-hdd StorageClass all deleted. ObjectStore untouched (still gs://knoe-0-backups/); backups continued through the Garage tear-down. |
Reality-vs-intent gaps surfaced by this work are captured in docs/TODO.md and as > **Reality TODO:** blocks in CLAUDE.md and AGENTS.md. Several gaps closed in this session (Garage on DB cluster, SA alignment); new ones added (redundant [knoe-db-0/knoe-db] IAM binding, init script + patch script need updates to reflect the new canonical end state).
Phase 1 — min Mode Pipeline
Status: 🔲 Not started — ready to begin (deferred during the 2026-04-28 cluster-recovery work above)
Depends on: Phase 0 complete ✅
First task: fix knoe/core/ops/cloudnative_pg.py:1372 f-string syntax (30 collection errors, see Phase 0 known issues)
What this phase does
Establishes the first green autobuild. The min pipeline is the simplest possible CI: a bare macOS runner (no Docker, no Kubernetes) running pytest. It exercises the welcome-screen mode selector, the min-mode navigation fast-path, and the init_min.sh script. Trigger: push to main on the Gitea remote.
Architecture
push to main (Gitea)
│
▼
Gitea Actions
│
▼
act_runner ←── bare macOS (dev machine or Mac mini)
│ labels: [self-hosted, macos, min]
▼ requires: Homebrew + Python 3.11
pytest -m "min or unit"
│
▼
coverage.xml ──► uploaded as artifact
Runner setup (one-time, not in CI):
# On the macOS runner machine:
brew install gitea-act-runner
act_runner register \
--instance https://<gitea-host>/ \
--token <runner-token-from-gitea-settings> \
--labels "self-hosted,macos,min" \
--name "knoe-min-runner"
act_runner daemon
Deliverables
| File | Purpose |
|---|---|
.gitea/workflows/ci-min.yml |
Gitea Actions workflow — push-triggered, runs on bare macOS runner |
tests/modes/__init__.py |
Package marker |
tests/modes/test_min_mode.py |
Mode-specific tests for min path |
tests/modes/conftest.py |
Mode fixtures (mock containerd, mock 1Password CLI) |
.gitea/workflows/ci-min.yml
name: CI — min mode
on:
push:
branches: [main]
jobs:
min-unit:
runs-on: [self-hosted, macos, min]
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install dependencies
run: |
pip install -r requirements.txt -r requirements-test.txt
- name: Run min + unit tests
run: |
pytest -m "min or unit" \
--cov=knoe \
--cov-report=xml \
--cov-report=term-missing \
-v \
tests/
env:
PYTHONPATH: ${{ github.workspace }}
KNOE_MODE: min
- name: Upload coverage artifact
uses: actions/upload-artifact@v4
with:
name: coverage-min
path: coverage.xml
retention-days: 14
Tests to write in tests/modes/test_min_mode.py
Each test name is intentionally human-readable — they become living documentation for junior developers.
| Test | What it proves |
|---|---|
test_welcome_mode_selector_starts_with_no_selection |
deployment_mode defaults to "" on fresh install — forces explicit choice |
test_clicking_min_card_sets_cluster_env_to_min |
Card click wires cluster_env = "min" correctly |
test_clicking_min_card_enables_next_button |
Welcome Next is disabled until a card is clicked |
test_min_nav_env_setup_goes_to_init_password_not_cluster |
Verifies min fast-path in on_next |
test_min_nav_init_scripts_goes_to_security_not_kerberos |
Min skips Kerberos, ArgoCD, GitOps, Supabase |
test_min_nav_prev_from_init_password_returns_to_env_setup |
Symmetric back-navigation |
test_min_nav_prev_from_security_returns_to_init_scripts |
Symmetric back-navigation |
test_normalize_cluster_env_min_returns_min |
_normalize_cluster_env("min") == "min" |
test_deployment_mode_from_env_min_returns_min |
_deployment_mode_from_env("min") == "min" |
test_init_min_sh_runs_without_error |
subprocess call to init_min.sh initialize exits 0 (macOS only, @pytest.mark.min) |
Commit Summary
(Fill in after phase lands)
# example:
# abc1234 Phase 1: add .gitea/workflows/ci-min.yml
# def5678 Phase 1: add tests/modes/test_min_mode.py
# ghi9012 Phase 1: add tests/modes/conftest.py
Resumption Checklist
- Push to
maintriggersci-minworkflow in Gitea Actions UI pytest -m "min or unit"exits 0 locally- All 10 tests in
test_min_mode.pypass - Coverage artifact appears in the Gitea Actions run summary
- Badge shows green on Gitea repo homepage (optional but nice)
Phase 2 — k3d Mode Pipeline
Status: 🔲 Not started
Depends on: Phase 1 complete
What this phase does
Adds the k3d integration pipeline — a Docker-enabled runner that creates a real k3d cluster, deploys CNPG, verifies the database is reachable, then tears it down. This is the first pipeline that proves an actual database cluster starts correctly. Estimated runtime: 8–15 minutes.
Architecture
push to main (Gitea)
│
▼
Gitea Actions
│
├─ job: k3d-unit (fast, no cluster) ─────────────► coverage-k3d-unit.xml
│
└─ job: k3d-integration (depends on unit) ──────► coverage-k3d-integration.xml
│
├── k3d cluster create knoe-ci
├── apply CNPG operator
├── apply knoe-db Cluster CR
├── pytest -m "k3d and integration"
└── k3d cluster delete knoe-ci (always)
Runner setup (one-time):
Same act_runner binary but registered with labels self-hosted,macos,k3d on a Docker-enabled machine (Docker Desktop or Colima).
Deliverables
| File | Purpose |
|---|---|
.gitea/workflows/ci-k3d.yml |
Two-job workflow: unit then integration |
tests/modes/test_k3d_mode.py |
k3d unit + integration tests |
tests/modes/fixtures/knoe-db-test.yaml |
Minimal CNPG Cluster CR for test use (single instance, tiny storage) |
tests/modes/fixtures/cnpg-operator-values.yaml |
Minimal Helm values for CNPG operator in CI |
.gitea/workflows/ci-k3d.yml
name: CI — k3d mode
on:
push:
branches: [main]
jobs:
k3d-unit:
runs-on: [self-hosted, macos, k3d]
timeout-minutes: 5
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.11" }
- run: pip install -r requirements.txt -r requirements-test.txt
- name: Unit tests (k3d, no cluster)
run: pytest -m "k3d and unit" --cov=knoe --cov-report=xml tests/
env:
PYTHONPATH: ${{ github.workspace }}
KNOE_MODE: k3d
- uses: actions/upload-artifact@v4
with: { name: coverage-k3d-unit, path: coverage.xml }
k3d-integration:
runs-on: [self-hosted, macos, k3d]
needs: k3d-unit
timeout-minutes: 25
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.11" }
- run: pip install -r requirements.txt -r requirements-test.txt
- name: Install k3d + kubectl
run: |
brew install k3d kubectl
- name: Create test cluster
run: |
k3d cluster create knoe-ci \
--agents 1 \
--k3s-arg '--disable=traefik@server:0' \
--wait
kubectl cluster-info --context k3d-knoe-ci
- name: Deploy CNPG operator
run: |
kubectl apply --server-side \
-f https://raw.githubusercontent.com/cloudnative-pg/cloudnative-pg/main/releases/cnpg-1.23.0.yaml
kubectl wait --for=condition=Available \
deployment/cnpg-controller-manager \
-n cnpg-system --timeout=120s
- name: Integration tests (live cluster)
run: |
pytest -m "k3d and integration" \
--cov=knoe \
--cov-report=xml \
-v \
tests/
env:
PYTHONPATH: ${{ github.workspace }}
KNOE_MODE: k3d
KUBECONFIG: ${{ env.HOME }}/.kube/config
- name: Tear down cluster
if: always()
run: k3d cluster delete knoe-ci
- uses: actions/upload-artifact@v4
with: { name: coverage-k3d-integration, path: coverage.xml }
Tests to write in tests/modes/test_k3d_mode.py
Unit tests (@pytest.mark.k3d @pytest.mark.unit):
| Test | What it proves |
|---|---|
test_k3d_cluster_env_normalizes_to_dev |
_normalize_cluster_env("dev") == "dev" |
test_welcome_k3d_card_sets_cluster_env_dev |
Card click → cluster_env = "dev" |
test_k3d_nav_proceeds_through_cluster_screens |
k3d mode does NOT skip init_cluster |
test_k3d_supabase_option_visible |
Supabase toggle is reachable in k3d nav flow |
test_cluster_lifecycle_milestone_uses_k3d_script |
ClusterLifecycleMilestone calls k3d cluster-create command for dev env |
Integration tests (@pytest.mark.k3d @pytest.mark.integration):
| Test | What it proves |
|---|---|
test_k3d_cluster_api_is_reachable |
kubectl can reach k3d-knoe-ci API server |
test_cnpg_operator_crds_registered |
Cluster CRD exists after operator deploy |
test_knoe_db_cluster_pod_starts |
Apply test CR → at least one PostgreSQL pod reaches Running |
test_database_accepts_connections |
psql SELECT 1 succeeds against the CNPG service |
test_cluster_delete_is_clean |
After delete, no k3d cluster named knoe-ci remains |
Commit Summary
(Fill in after phase lands)
Resumption Checklist
- Push to
maintriggers bothk3d-unitandk3d-integrationjobs in Gitea k3d-integrationpasses (PostgreSQL pod reaches Running state)- Both coverage artifacts appear in the run summary
k3d cluster listshows no leftoverknoe-cicluster after the run
Phase 3 — GKE Pipeline (GitLab CI, expanded)
Status: 🔲 Not started
Depends on: Phase 2 complete
What this phase does
Replaces the current single-job .gitlab-ci.yml (which only runs install.sh -S) with a proper multi-stage pipeline: lint → unit tests → GKE integration tests → dead-code report → deploy. The unit test gate runs on every push; GKE integration runs on main only.
Architecture
push to any branch
│
├─ stage: lint → python syntax + imports check
├─ stage: test → pytest -m "unit" (every push)
│
push to main only:
├─ stage: test → pytest -m "gke and integration" (real GKE cluster)
├─ stage: test → vulture dead-code report (allow_failure: true)
└─ stage: deploy → ./install.sh -S -c conf/service/prod.cfg
Deliverables
| File | Purpose |
|---|---|
.gitlab-ci.yml |
Full multi-stage pipeline replacing current single-job version |
tests/modes/test_gke_mode.py |
GKE unit + integration tests |
scripts/dead_code_analysis.py |
Cross-references coverage.json with vulture-report.txt, outputs three-tier report |
Required GitLab CI variables (project settings → CI/CD → Variables)
| Variable | Value | Protected | Masked |
|---|---|---|---|
GCP_SA_KEY |
GCP service account JSON with container.viewer + storage.objectViewer on both clusters |
✅ | ✅ |
GKE_REGION |
us-west3 |
||
GKE_APP_CLUSTER |
knoe-dev-0 |
||
GKE_DB_CLUSTER |
knoe-dev-cnpg-0 |
.gitlab-ci.yml (full replacement)
stages:
- lint
- test
- report
- deploy
variables:
PYTHON_VERSION: "3.11"
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
cache:
paths: [.cache/pip]
# ── Stage: lint ───────────────────────────────────────────────────────────────
lint:
stage: lint
image: python:3.11-slim
script:
- pip install -r requirements-test.txt -q
- python -m py_compile knoe/**/*.py
- python -c "from knoe.ui.screens import KnoeInstaller"
rules:
- if: '$CI_PIPELINE_SOURCE == "push"'
# ── Stage: test (unit — every push) ──────────────────────────────────────────
unit-tests:
stage: test
image: python:3.11-slim
script:
- pip install -r requirements.txt -r requirements-test.txt -q
- pytest -m "unit and not integration"
--cov=knoe
--cov-report=xml
--cov-report=term-missing
-q
tests/
coverage: '/TOTAL.*\s+(\d+\%)/'
artifacts:
reports:
coverage_report:
coverage_format: cobertura
path: coverage.xml
paths: [coverage.xml]
expire_in: 7 days
rules:
- if: '$CI_PIPELINE_SOURCE == "push"'
# ── Stage: test (GKE integration — main only) ─────────────────────────────────
gke-integration:
stage: test
image: google/cloud-sdk:slim
timeout: 30 minutes
before_script:
- pip install -r requirements.txt -r requirements-test.txt -q
- echo "$GCP_SA_KEY" | gcloud auth activate-service-account --key-file=-
- gcloud config set project plenary-truck-485623-p7
- gcloud container clusters get-credentials $GKE_APP_CLUSTER
--region $GKE_REGION
- gcloud container clusters get-credentials $GKE_DB_CLUSTER
--region $GKE_REGION
script:
- pytest -m "gke and integration"
--cov=knoe
--cov-report=xml
-v
tests/
coverage: '/TOTAL.*\s+(\d+\%)/'
artifacts:
reports:
coverage_report:
coverage_format: cobertura
path: coverage.xml
paths: [coverage.xml]
expire_in: 30 days
rules:
- if: '$CI_COMMIT_BRANCH == "main"'
# ── Stage: report (dead code — main only, never blocks build) ─────────────────
dead-code:
stage: report
image: python:3.11-slim
allow_failure: true
script:
- pip install vulture -q
- vulture knoe/ --min-confidence 80 | tee vulture-report.txt
- python scripts/dead_code_analysis.py
--vulture vulture-report.txt
--coverage coverage.xml
--output dead-code-report.md
artifacts:
paths:
- vulture-report.txt
- dead-code-report.md
expire_in: 30 days
rules:
- if: '$CI_COMMIT_BRANCH == "main"'
# ── Stage: deploy (main only) ─────────────────────────────────────────────────
deploy-service:
stage: deploy
rules:
- if: '$CI_COMMIT_BRANCH == "main"'
script:
- ./install.sh -S -c conf/service/prod.cfg
environment:
name: production
Tests to write in tests/modes/test_gke_mode.py
Unit tests (@pytest.mark.gke @pytest.mark.unit):
| Test | What it proves |
|---|---|
test_gke_cluster_env_normalizes_to_prod |
_normalize_cluster_env("prod") == "prod" |
test_welcome_gke_card_sets_cluster_env_prod |
Card click → cluster_env = "prod" |
test_gke_split_cluster_detected |
When app_ctx != db_ctx, split-cluster mode activates |
test_garage_not_deployed_to_db_cluster |
CNPG cluster does not include Garage in GKE mode |
test_gke_storage_class_validation |
standard-rwo raises quota error hint; standard passes |
Integration tests (@pytest.mark.gke @pytest.mark.integration):
| Test | What it proves |
|---|---|
test_gke_app_cluster_reachable |
kubectl can reach knoe-dev-0 |
test_gke_db_cluster_reachable |
kubectl can reach knoe-dev-cnpg-0 |
test_cnpg_cluster_knoe_db_0_running |
knoe-db cluster in knoe-db-0 namespace has 3 Ready instances |
test_barman_backup_schedule_exists |
ScheduledBackup CR exists in knoe-db-0 |
test_garage_only_in_app_cluster |
No Garage pods in knoe-dev-cnpg-0 namespace |
Commit Summary
(Fill in after phase lands)
Resumption Checklist
- GitLab pipeline shows four stages: lint → test → report → deploy
unit-testsjob passes on every push (not just main)gke-integrationpasses on main (3 CNPG pods Running)dead-codejob produces adead-code-report.mdartifact- Deploy stage still works (
install.sh -Sexits 0) - GitLab shows coverage percentage on the merge request widget
Phase 4 — Dead Code Identification
Status: 🔲 Not started
Depends on: Phase 3 complete (all three stable builds achieved)
What this phase does
Runs the first systematic dead code analysis across the full codebase. This is not a cleanup sprint — it is reconnaissance. We generate a prioritised report and review it before deleting anything. The report becomes the input for Phase 5.
Why this matters
The project has undergone several major renames (prole → knoe, various directory restructures). Code written for old layouts is likely still present. knoe/core/actions.py alone is 7,740 lines and almost certainly contains branches that no deployment mode ever exercises.
Tools
| Tool | Role |
|---|---|
vulture |
Static analysis — finds unused functions, classes, imports |
pytest --cov --cov-branch |
Dynamic analysis — lines/branches never executed during tests |
scripts/dead_code_analysis.py |
Cross-references both outputs to produce a tiered report |
Three-tier output format
## 🔴 Definite dead code (vulture-flagged AND 0% coverage)
- knoe/core/ops/legacy_shell.py:47 — function `_old_prole_exec` (unused, 0% coverage)
- ...
## 🟡 Suspect (vulture-flagged OR 0% coverage, not both)
- knoe/core/actions.py:4201 — function `_k8s_node_drain` (0% coverage, not flagged by vulture)
- ...
## 🟢 Live (covered by tests and vulture-clean)
- (omitted from report for brevity)
How to run locally
# 1. Full coverage run across all modes
KNOE_MODE=min pytest -m min --cov=knoe --cov-append tests/
KNOE_MODE=k3d pytest -m k3d --cov=knoe --cov-append tests/
KNOE_MODE=gke pytest -m gke --cov=knoe --cov-append tests/
coverage json # produces coverage.json
# 2. Vulture scan
vulture knoe/ --min-confidence 70 > vulture-report.txt
# 3. Cross-reference
python scripts/dead_code_analysis.py \
--vulture vulture-report.txt \
--coverage coverage.json \
--output dead-code-report.md
# 4. Review
open dead-code-report.md
Deliverables
| File | Purpose |
|---|---|
scripts/dead_code_analysis.py |
Cross-reference script (coverage.json + vulture output → tiered report) |
dead-code-report.md |
Generated output — committed to repo for review, not production |
docs/dead-code-review.md |
Human review notes — which 🔴 items are safe to delete vs accidentally flagged |
Commit Summary
(Fill in after phase lands)
Resumption Checklist
python scripts/dead_code_analysis.pyruns without errordead-code-report.mdcontains all three tiers- The 🔴 list has been manually reviewed and each item categorised as "safe to delete" or "keep" in
docs/dead-code-review.md - No production code deleted yet — this phase is report-only
Phase 5 — Pruning and the Instructable Codebase
Status: 🔲 Not started
Depends on: Phase 4 report reviewed and approved
What this phase does
This is the refactoring sprint. Using the Phase 4 report, we delete dead code in small, test-verified commits, raising the fail_under coverage threshold after each deletion. The end state is a codebase that a junior developer can read from top to bottom and fully understand.
Target metrics
| Metric | Today (est.) | Target |
|---|---|---|
| Test coverage (unit) | ~40% | ≥ 80% |
| Test coverage (integration) | ~10% | ≥ 60% |
vulture --min-confidence 80 warnings |
Unknown | 0 |
Lines in knoe/core/actions.py |
7,740 | < 2,000 |
| Modules with zero test coverage | ~17 (knoe/core/ops/) |
0 |
fail_under in pyproject.toml |
0 | 75 |
Deletion protocol
For each 🔴 item from the Phase 4 report:
- Write a test that would fail if the code were still needed (proves it's safe to delete)
- Delete the code
- Run
pytest— all tests pass - Commit with message:
prune: remove <module/function> — dead code (vulture + 0% coverage) - Raise
fail_underby 1–2 points
This protocol means every deletion is backed by a test. The test suite grows because we prune.
The instructable codebase standard
A module is "instructable" when:
- Every public function has at least one test whose name reads as a plain English sentence
- The test file for the module is shorter than the module itself
- A junior developer can understand the module's purpose from the test names alone without reading the source
Example of instructable test names:
def test_min_mode_only_needs_homebrew_and_1password(): ...
def test_k3d_mode_creates_cnpg_cluster_in_docker(): ...
def test_welcome_card_click_routes_to_correct_nav_flow(): ...
def test_kerberos_is_only_enabled_in_k3s_and_gke_modes(): ...
Commit Summary
(Fill in after each deletion sprint)
Resumption Checklist
fail_underhas been raised at least once since Phase 4- Zero 🔴 items remain in
dead-code-report.md - All modules in
knoe/core/ops/have at least one test - CI passes on all three pipelines (min, k3d, gke)
k3s Phase — Homelab Pipeline
Status: ⏸ Deferred
Trigger: Evaluate after GitLab CI (Phase 3) is stable
Notes for when this is ready
- Runner: one of the physical k3s nodes (
myrddin.prole.org,gandalf.prole.org, ormerlin.prole.org) registered as a Gitea act_runner with labelsself-hosted,linux,k3s - Integration tests will need
kubeconfigfor the k3s cluster and iSCSI storage access for CNPG - Kerberos integration is the unique test target here —
knoe-authdeployment with embedded KDC - The
CNPG_ELIGIBLE_NODESconfig var maps directly to tests that verify node affinity placement
Appendix A — Pytest Marker Reference
| Marker | When to use |
|---|---|
@pytest.mark.unit |
No external services, no file system writes, runs anywhere |
@pytest.mark.integration |
Requires a running cluster, database, or network service |
@pytest.mark.min |
Tests the min (containerd) deployment path |
@pytest.mark.k3d |
Tests the k3d (local Docker cluster) deployment path |
@pytest.mark.k3s |
Tests the k3s (homelab) deployment path |
@pytest.mark.gke |
Tests the GKE (production) deployment path |
Combining markers is the norm: @pytest.mark.k3d @pytest.mark.integration means "requires a live k3d cluster".
Appendix B — Runner Registration Quick Reference
Gitea act_runner (macOS)
# Install
brew install act-runner
# Register (run once per machine)
act_runner register \
--instance https://<gitea-host>/ \
--token <token-from-gitea-settings-actions-runners> \
--labels "self-hosted,macos,min" # or min,k3d for the Docker machine \
--name "knoe-min-runner" # human-readable name in Gitea UI
# Start as a service
brew services start act-runner
GitLab runner (existing)
The existing GitLab runner is assumed to have gcloud CLI available. If not:
# On the runner host:
curl https://sdk.cloud.google.com | bash
gcloud components install gke-gcloud-auth-plugin
Appendix C — Coverage Increment Strategy
Rather than setting an ambitious fail_under up front and having CI permanently broken, we raise it in steps as each phase lands:
| After phase | fail_under |
|---|---|
| Phase 0 (foundation) | 0 (measure only) |
| Phase 1 (min tests) | 15 |
| Phase 2 (k3d tests) | 25 |
| Phase 3 (gke tests) | 35 |
| Phase 4 (dead code pruned) | 50 |
| Phase 5 (full prune) | 75 |
Each increment is a one-line change to pyproject.toml committed at the end of the phase.