chore: purge dead credentials and harden .gitignore

- Delete tracked knoe-db/.prole_user_password (dead secret, rotated
  in ansible since the original commit)
- Remove postgres-password.txt and build/ copy from working tree
- Add explicit ignore rules for *_password, postgres-password*,
  .prole_user_password, *-secret variants

Add docs/plans/customer-deploy-resync.md with the upstream/downstream
migration plan and update CLAUDE.md to reflect prole's role as a
customer deploy of knoe-db.
This commit is contained in:
chrisfu 2026-04-27 13:33:29 -07:00
parent 34de8a5f34
commit 553612bdff
5 changed files with 274 additions and 1 deletions

10
.gitignore vendored
View File

@ -45,7 +45,17 @@
*.kubeconfig*
etc/secrets/
*-password.txt
# Catches both the legacy literal `.prole_user_password` (was committed
# to history under prole-db/ and knoe-db/) and any future variant.
.prole_user_password
**/.prole_user_password
*_password
*.password
postgres-password*
mssql-password*
*secret.yaml
*-secret.yaml
*-secret
.vault_pass
/conf/prole.cfg
/secrets/

View File

@ -4,6 +4,21 @@
---
## Repo role: customer deploy of `knoe-db`
This working tree (`prole`) is being converged into a **customer-deploy branch** of the upstream platform repo `knoe-db` (remote: `git@git-ssh.knoe.dev:knoe.dev/knoe-db.git`, configured here as the `knoe` remote). Platform code — `authority/`, `knoe/`, `etc/init_*.sh`, `deploy/gcp/gke/*`, the test pipeline — is not maintained in this repo; its source of truth is `knoe-db/main`. Customer-specific divergence (config, branding, on-prem `prole.org` manifests) is what stays on the customer branch.
**Before changing anything in this tree, check whether the change belongs upstream.** If it's platform code, land it on `knoe-db/main` and rebase the customer branch.
## Active plans
- [`docs/plans/customer-deploy-resync.md`](docs/plans/customer-deploy-resync.md) — The active plan for converging `prole` onto `knoe-db/main` as a customer-deploy branch. Inventory + cherry-pick + rebase cadence. **Read this first** if you are about to touch anything in this tree.
- [`docs/plans/README.md`](docs/plans/README.md) — Index and conventions for this directory.
For platform-level plans (knoe-auth, deployment modes, etc.), read [`knoe-db/docs/plans/`](../knoe-db/docs/plans/) (path relative to your `~/dev/` root). Those documents are the architectural reference for the shipped platform features and are versioned with the upstream code, not here.
---
## Dual-cluster GKE architecture
This project uses **two separate GKE Standard clusters** in `us-west3`, both currently provisioned with `e2-standard-2` × 3 nodes (2 vCPU / 8 GB each, ~7.1 GB allocatable). Verify with `gcloud container clusters list``knoe/core/actions.py` still carries an `e2-small` default for the app cluster but the live cluster is `e2-standard-2`.

24
docs/plans/README.md Normal file
View File

@ -0,0 +1,24 @@
# Plans (prole — customer deploy)
This directory holds **plans specific to the prole.org customer deployment** of the knoe.dev platform.
## What lives here
| File | Purpose |
| --- | --- |
| [`customer-deploy-resync.md`](customer-deploy-resync.md) | The active plan to converge this working tree onto `knoe-db/main` as a long-lived customer branch. |
## What does **not** live here
Platform-level plans live upstream in `~/dev/knoe-db/docs/plans/`:
- [`knoe-auth-round-1.md`](../../../knoe-db/docs/plans/knoe-auth-round-1.md) — Kerberos identity + invite-OTP enrollment + TOTP. Architectural reference.
- [`deployment-modes.md`](../../../knoe-db/docs/plans/deployment-modes.md) — Four-mode installer + welcome-screen selector. Architectural reference.
(Paths above are relative to your `~/dev/` root. If you're reading this on the file server, navigate to the `knoe-db` checkout.)
## Why the split
`prole` is being converged into a customer-deploy branch in `knoe-db`, not a separate fork. Platform docs travel with the platform code (in `knoe-db`); customer-deploy-specific docs travel with the customer branch (here). When the customer branch is published in `knoe-db`, this directory will move with it.
See the upstream `~/dev/knoe-db/docs/plans/README.md` for the full conventions on writing plans, plan statuses, and how customer deploys are structured.

View File

@ -0,0 +1,225 @@
# Resync `prole` as a Customer Deploy of `knoe-db`
**Status:** Active plan. Not yet executed.
**Owner:** chrisfu
**Audience:** Jr/mid engineer joining the project. No prior knowledge of the prole↔knoe-db split assumed.
---
## 1. Context
This working tree (`~/dev/prole`) and `~/dev/knoe-db` started life as the same project. Over time `knoe-db` has become the canonical upstream — the "Knoe" infrastructure platform, hosted at `git@git-ssh.knoe.dev:knoe.dev/knoe-db.git` — while `prole` has continued as a separate fork on GitHub with its own diverging history.
The strategic decision the team has made is: **stop forking, start branching.** `knoe-db` is the platform's source of truth. Customer deployments — currently `prole.org`, eventually others — live as **branches in `knoe-db`**, not as separate repositories. A customer branch carries only the customer-specific divergence (config files, branding, on-prem manifests, kubeconfig handling). Platform changes always land on `knoe-db/main` and customer branches rebase or merge from main on a regular cadence.
This plan describes how to bring `~/dev/prole` into that shape.
### Why a branch and not a fork
- **One pull-request review surface for platform changes.** A change that affects all deployments goes through MR review on `knoe-db/main` once. In a fork model the same change has to be re-reviewed (or, more often, silently re-implemented) per fork.
- **Customer-specific divergence is observable.** A `git diff main..prole` is the complete answer to "what's special about the prole.org deploy?" In a fork model that question requires a manual cross-repo diff that quickly goes stale.
- **Resync is mechanical.** `git rebase main` or `git merge main` updates the deploy branch. In a fork model resync is a recurring tax that engineers underinvest in until the forks have drifted past easy reconciliation — which is roughly where prole is today.
- **Customer-specific platform changes get attention.** If a customer needs a platform-level change to support their deploy, it lands on `main` (and is therefore available to all customers) instead of being hidden in a fork that never makes it back.
### What's already done
- prole has `knoe-db` configured as a git remote named `knoe`:
```
knoe git@git-ssh.knoe.dev:knoe.dev/knoe-db.git (fetch)
knoe git@git-ssh.knoe.dev:knoe.dev/knoe-db.git (push)
origin git@github.com:dredx/prole.git (fetch)
origin git@github.com:dredx/prole.git (push)
```
So the network plumbing for `git fetch knoe` is in place. `origin` is the legacy GitHub fork.
- `knoe-db` already has `docs/plans/` with the architectural references for `knoe-auth Round 1` and the four-mode installer — see [`knoe-db/docs/plans/`](../../../knoe-db/docs/plans/) (path relative to your `~/dev/` root).
### What's not yet done
- Most of `prole/main` has not seen `knoe-db/main`'s recent work (the prole→knoe rebrand, the welcome-mode-selector Phase 0, the fully implemented `authority/` module, the 1Password CLI cutover, etc.).
- Prole still tracks `infrastructure/` (the prole.org-specific Ansible tree) which `knoe-db` excised in commit `5a769f2`.
- Prole still tracks the old `prole-*` directory names that knoe-db has renamed (`prole-app/`, `prole-tools-app/`, `prole.sh`, `prole.spec`, `prole_requirements.txt`).
- Prole's working tree carries credentials and runtime artifacts that should never have been near git in the first place — `postgres-password.txt`, `prole-k3s.kubeconfig*`, `secrets/`, `ssh-keys/`. None are committed (verified with `git ls-files`), but they exist in the working copy and are a data-leak risk.
---
## 2. Target shape
When this plan is done:
- The `prole` working directory tracks a branch in `knoe-db` — name to be confirmed, proposed `customer/prole.org` or simply `prole`.
- `git remote` looks like:
```
origin git@git-ssh.knoe.dev:knoe.dev/knoe-db.git
```
The legacy GitHub `origin` is either retired entirely or kept as `archive` for historical reference.
- `git diff origin/main` (i.e. `knoe-db/main`) shows **only** prole-specific divergence:
- Customer-branded config (`conf/prole.cfg`, host names, realm `PROLE.LOCAL`, etc.)
- On-prem manifests scoped to `prole.org` infrastructure (`deploy/opentofu/k3s/manifests/prole/*`)
- Customer-specific docs (`docs/plans/customer-deploy-resync.md` — this file — and any future prole-only plans)
- All tracked platform code (`authority/`, `knoe/`, `etc/init_*.sh`, `deploy/gcp/gke/*`, the test pipeline) is **byte-identical** to `knoe-db/main`. Changes to those files happen on `knoe-db/main` first and reach prole via merge/rebase.
- Working-tree credentials and runtime files are removed, gitignored, and replaced with the 1Password-driven equivalents already used in `knoe-db`.
### What "customer-specific" actually means here
It's a short list, and the discipline is to keep it short:
| Type | Example | Where it lives on the branch |
| --- | --- | --- |
| Customer config | `conf/prole.cfg` (cluster contexts, realm, hostnames) | `conf/` |
| Customer realm | `PROLE.LOCAL` Kerberos manifests | `deploy/opentofu/k3s/manifests/prole/prole-kdc-*.yaml` |
| Customer cluster glue | k3s `prole.org` host inventory, anything host-pinned to `merlin/myrddin/pi.prole.org` | `deploy/opentofu/k3s/manifests/prole/` |
| Customer docs | `docs/plans/customer-deploy-resync.md`, any future prole-only specs | `docs/plans/` |
Anything that doesn't fit one of those four lines is probably platform code that should land on `main`.
---
## 3. Migration plan
The migration is in four phases. Each phase is reviewable on its own; do not collapse them.
> **All git work in this plan should run in Claude Code, not in Cowork.** Cowork's bash sandbox has no network access and cannot reach `git-ssh.knoe.dev`. Open a Code session in `~/dev/prole` and drive the commands from there.
### Phase 1 — Snapshot and inventory
The goal of this phase is to produce a complete account of what's currently special about prole, *before* anyone starts deleting things. Output is a file checked into the working tree (`docs/plans/resync-inventory.md`, net-new) listing every file that differs from `knoe-db/main`, classified as one of:
- **PLATFORM-DRIFT** — file should match `knoe-db/main`. Discrepancy is just stale prole. Action: take knoe-db's version.
- **CUSTOMER-SPECIFIC** — file is genuinely prole-only. Action: keep, plan to land on the customer branch.
- **WORKING-TREE-NOISE** — credentials, build artifacts, kubeconfigs. Action: gitignore upstream if not already; remove from working tree; never commit.
- **AMBIGUOUS** — needs human review. Action: discuss before acting.
Concrete commands (run in Code in `~/dev/prole`):
```bash
# fetch knoe-db
git fetch knoe
# list all files that exist in prole but not in knoe-db/main
git diff --name-only --diff-filter=D knoe/main..HEAD > /tmp/prole-only.txt
# list all files that exist in both but differ
git diff --name-only knoe/main..HEAD | grep -v '^docs/plans/customer-deploy-resync.md$' > /tmp/prole-modified.txt
# list all files that exist in knoe-db/main but not in prole
git diff --name-only --diff-filter=A knoe/main..HEAD > /tmp/knoe-db-new.txt
```
Walk through each list and classify. Expected hot spots based on the current state:
- `infrastructure/` — entirely PLATFORM-DRIFT (excised upstream by commit `5a769f2`).
- `prole-app/`, `prole-auth/`, `prole-mssql-db/`, `prole-net/`, `prole-tools-app/`, `prole.sh`, `prole.spec`, `prole_requirements.txt` — PLATFORM-DRIFT (renamed upstream during the prole→knoe rebrand).
- `authority/src/main/java/org/prole/authority/HealthController.java` etc. — PLATFORM-DRIFT (stale Spring Boot skeleton; upstream has the full module).
- `deploy/opentofu/k3s/manifests/prole/*` — likely CUSTOMER-SPECIFIC (the `prole.org` realm and host pinning belongs on the customer branch).
- `conf/prole.cfg`, `conf/prole.silent.all.cfg` — CUSTOMER-SPECIFIC.
- `postgres-password.txt`, `prole-k3s.kubeconfig*`, `secrets/`, `ssh-keys/` — WORKING-TREE-NOISE. These should not be tracked anywhere.
Output of phase 1 is the inventory file plus a short summary commit on a fresh branch (proposed name `resync/phase-1-inventory`).
### Phase 2 — Bring `prole/main` to `knoe-db/main`
> ⚠️ **Read this whole phase before running any of it.** This is where the divergence collapses; recovery from a botched merge is much harder than waiting an hour to do it right.
Approach: do **not** try to merge upstream into prole's diverged main. Instead:
1. Create a backup tag at the current state: `git tag prole-pre-resync-$(date +%Y%m%d)`. Push it to `origin` (the GitHub fork) for safety.
2. Create a fresh customer-deploy branch off `knoe-db/main`:
```bash
git checkout -b customer/prole.org knoe/main
```
3. Cherry-pick **only** the CUSTOMER-SPECIFIC commits/files from the inventory. The cleanest way is a one-pass apply:
```bash
# for each file in the customer-specific list, copy from old prole main
git checkout prole-pre-resync-<date> -- <path>
git add <path>
```
Commit in coherent groups (e.g. one commit per logical scope: "config", "k3s manifests", "customer docs").
4. Push the new branch to the `knoe` remote:
```bash
git push knoe customer/prole.org
```
5. Open an MR in `knoe-db` for the customer branch (it doesn't merge to main — it's a long-lived branch — but the MR makes it visible and reviewable).
After this phase, `customer/prole.org` is byte-identical to `main` except for the (small) customer-specific delta.
### Phase 3 — Retire the GitHub fork and clean the working tree
1. Update the local working tree's `origin` remote to point at `knoe`:
```bash
git remote rename origin archive
git remote rename knoe origin
git fetch origin
```
Or, if you want to delete the GitHub fork outright, drop the `archive` remote. Either is fine; the GitHub copy is no longer authoritative.
2. Switch the working tree to the new branch:
```bash
git checkout customer/prole.org
```
3. Delete working-tree noise (the WORKING-TREE-NOISE bucket from phase 1). For each file:
- Confirm it's not tracked: `git ls-files --error-unmatch <path>`.
- Confirm it's covered by `.gitignore` (or add it).
- Delete it: `rm -rf <path>`.
- Special case: `postgres-password.txt`. Treat the contents as compromised because it was sitting in a working tree synced to who-knows-where. **Rotate the password through the standard process**, do not just delete the file.
4. Replicate `knoe-db`'s 1Password CLI workflow: `etc/` already has the 1Password integration after `knoe-db` commit `b421f49`. Set up the 1Password CLI and verify `op signin` works before relying on the new flow.
### Phase 4 — Establish the resync cadence
Going forward, prole resyncs from `knoe-db/main` on a deliberate cadence (proposal: weekly, or after any commit on `main` that touches a file the customer branch carries a divergent copy of).
Mechanism — pick one:
- **Rebase** `customer/prole.org` onto `knoe-db/main`. Linear history, clean diff against main, but rewrites the customer-branch SHAs on every resync. Choose this if no other tooling has hardcoded references to customer-branch commits.
- **Merge** `knoe-db/main` into `customer/prole.org`. Preserves history, no SHA rewriting, but the `git log --oneline` gets messy and `git diff main..customer/prole.org` includes merge artifacts.
Recommendation: **rebase**. Customer divergence is small enough to keep clean, and the simpler diff is worth more than the merge-commit history.
Set up CI on `knoe-db` to run the customer-deploy smoke test against the rebased branch nightly. Failure of that smoke test is an early signal that platform changes have broken the customer deploy and need a customer-side fix.
---
## 4. Out of scope
These are real items but they are **not** part of the resync. Each becomes its own plan when its time comes.
- Migrating the *running* prole.org cluster to the new branch's deployment shape. Resync is a code-level operation; the live cluster is a separate migration governed by `etc/init_*.sh` and the installer.
- Bringing prole's existing Supabase work back into knoe-db/main. Recent prole commits include Supabase fixes (`fix(supabase): ...`) that aren't in upstream. Some of that may be platform code that should land on `main`; some may be prole-specific. Sort that out as a follow-up MR.
- The active "knoe-db Init:0/1 on myrddin" debug. That's a runtime-state problem on the prole.org cluster, unrelated to the source-code resync. Track in its own ticket.
- Onboarding additional customer deploys. Once one customer branch is healthy, the second is mostly mechanical, but not part of this plan's scope.
---
## 5. Verification
The resync is done when all of these are true.
1. `git remote -v` in `~/dev/prole` shows `origin` pointing at `knoe-db`. The legacy GitHub remote is either gone or named `archive`.
2. `git status` is clean on branch `customer/prole.org`.
3. `git diff knoe/main` (or `git diff origin/main` post-rename) shows **only** files in the four categories listed in §2 ("What 'customer-specific' actually means"). No `infrastructure/`, no `prole-app/`, no stale Java skeletons.
4. `git ls-files` does not list `postgres-password.txt`, `prole-k3s.kubeconfig*`, or any file under `secrets/` or `ssh-keys/`.
5. The Postgres password that was previously sitting in `postgres-password.txt` has been rotated.
6. `./install.sh` or `./knoe.sh install` runs end-to-end on the customer branch in `min` and `k3d` modes (the cheap modes to verify locally).
7. `customer/prole.org` is pushed to `knoe-db` and visible in the upstream repo's branch list.
8. `docs/plans/customer-deploy-resync.md` (this file) and `docs/plans/resync-inventory.md` (created in phase 1) are committed on the customer branch.
---
## 6. Risks and call-outs
- **Don't run any of this with the prole.org cluster in a degraded state if the cluster recovery depends on a script that the resync deletes.** The current `knoe-db` Init:0/1 debug is a live ops problem; finish that first, or be very careful that the scripts you're about to delete from prole are not the ones the cluster needs to come back.
- **The Postgres password file is a real incident.** A credential file in a working tree synced via Dropbox / iCloud / network home directory has a non-trivial blast radius. Treat as compromised, rotate, do not just delete.
- **The GitHub `origin` may have copies of secrets in its history.** Before retiring the fork, audit its commit history for committed credentials (`git log -p | grep -iE 'password|secret|token|key'` is a starting point). If anything is in there, rotate before retiring. GitHub's "delete repository" does not retroactively un-disclose anything that was scraped while the repo was up.
---
## 7. Glossary
**Upstream / downstream** — Upstream is the canonical project (`knoe-db`); downstream is a deployment of it (`prole.org` customer branch). Changes flow from upstream to downstream, never the other way without an MR back.
**Customer deploy** — A specific deployment of the knoe.dev platform for one organization. Carries customer-specific config (hostnames, realm, branding) on top of the platform's main branch. Lives as a long-lived branch in the platform repo, not a fork.
**Long-lived branch** — A branch that's not intended to merge back to main. It's continuously rebased onto main as platform work lands, but its own divergence stays as a divergence forever.
**Resync** — The act of bringing a customer branch up to a newer `main`. A periodic, deliberate operation; not a continuous one.
**Phase 0** — The most recent platform-side rebrand and test-pipeline foundation work (commit `0052a4d` in `knoe-db`). Context here only because it's what `customer/prole.org` is rebasing onto.

View File

@ -1 +0,0 @@
${OPENBAO:kv/prole/<namespace>/db#password}