From d6586cf1d938383cfc625a3753dc836fc1fda31f Mon Sep 17 00:00:00 2001 From: chrisfu Date: Mon, 11 May 2026 01:24:34 -0700 Subject: [PATCH] kdc: backport init_kdc.sh trust fixes + add reset-repeatable Junie brief Mirrors the knoe-db commit `ff7546d` patches into the prole copy of `etc/init_kdc.sh` so a re-run of `install.sh --mode k3s --reset` from this repo produces a working cross-realm trust without manual cluster surgery. The k3s cluster is provisioned from this repo, so the source fix must live here (knoe-db remains canonical for GKE). Changes to etc/init_kdc.sh: 1. Create BOTH cross-realm krbtgts in MIT, not just the outbound one. The inbound `krbtgt/@` (issued by Samba, decrypted here) was missing entirely; without it, MIT cannot decrypt inbound TGTs and the trust never carries traffic. 2. Pin both cross-realm krbtgts to RC4 (`arcfour-hmac:normal`). AES keys depend on salt, and Samba's `+UPN` salt does not match MIT's `+`; RC4 has no salt so both sides converge from the password alone. Matches the already-pinned Samba side (commit `ad1eced`). 3. Replace the broken "remote kadmin to Samba" reciprocal-trust block with a documented no-op pointing at `infrastructure/playbooks/kerberos_trust_setup.yml`. Samba AD does not accept additions over MIT's kadmin protocol; the block always failed with "Missing parameters in krb5.conf required for kadmin client". 4. Switch the KDC data volume from emptyDir to a PVC (claimName `knoe-kdc-data`, parameterized by `$PROLE_KDC_STORAGE_SIZE` and `$PROLE_KDC_STORAGE_CLASS`). State now survives pod restarts. Adds Junie brief `docs/plans/junie/kdc-trust-reset-repeatable.md` with four TDD acceptance criteria for an end-to-end --reset run on the prole k3s cluster. Co-Authored-By: Claude Sonnet 4.6 --- .../plans/junie/kdc-trust-reset-repeatable.md | 224 ++++++++++++++++++ etc/init_kdc.sh | 63 ++++- 2 files changed, 278 insertions(+), 9 deletions(-) create mode 100644 docs/plans/junie/kdc-trust-reset-repeatable.md diff --git a/docs/plans/junie/kdc-trust-reset-repeatable.md b/docs/plans/junie/kdc-trust-reset-repeatable.md new file mode 100644 index 0000000..4a13e52 --- /dev/null +++ b/docs/plans/junie/kdc-trust-reset-repeatable.md @@ -0,0 +1,224 @@ +# Junie brief: end-to-end k3s reset/deploy that produces a working KNOE.LOCAL↔PROLE.ORG Kerberos trust + +## Why + +A full afternoon on 2026-05-10/11 went into wiring up the +PROLE.ORG ↔ KNOE.LOCAL cross-realm Kerberos trust on the prole k3s +cluster (myrddin/merlin/gandalf). The Samba side now works +(`infrastructure/playbooks/kerberos_trust_setup.yml`, commits +`5cece40`…`ad1eced`), but the in-cluster KDC pod was running for realm +**`PROLE.LOCAL`** — a pre-rebrand artifact — and could not decrypt +referral TGTs from Samba. Every "Server not found in Kerberos +database" / "Decrypt integrity check failed" / "PROCESS_TGS" symptom +we chased was downstream of that. + +Root cause: the originally-deployed pod inherited +`PROLE_KDC_REALM=PROLE.LOCAL` from a stale `etc/init_kdc.sh` shell +default. The default was fixed in commit `4fe8647` today, and a set +of larger patches to `etc/init_kdc.sh` followed. + +What's **not** done: running `install.sh --reset` end-to-end against +the k3s cluster to confirm the patches converge to a working trust +from a blank slate. That's this brief. + +## What's already landed (do not re-do) + +| Commit | Subject | +|---|---| +| `4fe8647` | `etc/init_kdc.sh:327` shell default `PROLE.LOCAL` → `KNOE.LOCAL` | +| `5cece40`…`227f490` | `infrastructure/playbooks/kerberos_trust_setup.yml` — full Samba-side trust account provisioning | +| `ad1eced` | `kerberos_trust_setup.yml` — `msDS-SupportedEncryptionTypes` 28 → 4 (RC4 only) | +| _(also pending in this branch)_ | `etc/init_kdc.sh` — both cross-realm krbtgts created in MIT with RC4, KDC volume switched from emptyDir to PVC, broken remote-kadmin block replaced with documentation pointing at the playbook | + +The matching source fix also landed in `~/dev/knoe-db` (commit +`ff7546d`) so the GKE / canonical installer stays in sync. **This +brief runs against the prole k3s cluster from this repo (`~/dev/prole`) +— do NOT run it against the knoe-db repo.** + +## What changes + +Nothing more in source. Your job is to **run** the end-to-end +reset-and-deploy from `~/dev/prole` and verify it succeeds without +manual cluster surgery. + +## Execution path + +> **Sequencing constraint:** `install.sh --reset` for `--mode k3s` runs +> the `k3s_reset.yml` Ansible playbook, which **wipes and reinstalls +> k3s on all three k3s nodes (myrddin, merlin, gandalf)**. Schedule +> accordingly — anything currently running in that cluster goes away. +> +> Before starting, capture the current state so you can roll back if +> needed: +> +> ```bash +> kubectl --context=default -n knoe-system get pods,svc,pvc,configmap > /tmp/pre-reset.txt +> kubectl --context=default -n knoe-system get deploy authority-knoe-auth -o yaml > /tmp/pre-deploy.yaml || true +> ``` + +```bash +# Run from a control node with kubectl access to the prole k3s cluster +# (myrddin itself is fine — it's the k3s server) + +cd ~/dev/prole +git pull # pick up the trust-fix commits if you don't have them locally + +# 1. RESET — wipes k3s on all three nodes +./install.sh --mode k3s --reset + +# 2. DEPLOY — re-provisions services including the KDC. +# Watch the init_kdc.sh output for: +# "Applying Knoe KDC '...' in namespace '...' (realm KNOE.LOCAL)" +# "Creating outbound trust principal krbtgt/PROLE.ORG@KNOE.LOCAL" +# "Creating inbound trust principal krbtgt/KNOE.LOCAL@PROLE.ORG" +# "Note: Samba-side trust account is provisioned out-of-band by ..." +./install.sh --mode k3s + +# 3. Sync the Samba-side trust account to the (possibly new) cluster +# Secret value. +ANSIBLE_VAULT_PASSWORD_FILE=$PWD/.vault_pass \ + ansible-playbook infrastructure/playbooks/kerberos_trust_setup.yml +``` + +## TDD acceptance criteria + +All four must pass. Capture the output of each and paste into your +MR description. + +### AC1 — KDC pod has realm `KNOE.LOCAL` + +```bash +kubectl --context=default -n knoe-system exec deploy/auth -- \ + cat /etc/krb5.conf | grep default_realm +# EXPECT: default_realm = KNOE.LOCAL +``` + +### AC2 — Both cross-realm krbtgts exist in MIT + +```bash +kubectl --context=default -n knoe-system exec deploy/auth -- \ + kadmin.local -q 'listprincs' | grep -E '^krbtgt/' +# EXPECT (at minimum): +# krbtgt/KNOE.LOCAL@KNOE.LOCAL +# krbtgt/KNOE.LOCAL@PROLE.ORG +# krbtgt/PROLE.ORG@KNOE.LOCAL +``` + +### AC3 — KDC data PVC is bound + +```bash +kubectl --context=default -n knoe-system get pvc knoe-kdc-data +# EXPECT: STATUS=Bound +``` + +If you see `Pending`, look at the StorageClass: + +```bash +kubectl get sc +kubectl --context=default -n knoe-system describe pvc knoe-kdc-data +``` + +The default k3s StorageClass `local-path` should be sufficient. If +your cluster doesn't have a default SC, re-run install with +`PROLE_KDC_STORAGE_CLASS=local-path` in the environment. + +### AC4 — End-to-end cross-realm ticket flow + +```bash +# Pre-stage a test SPN in KNOE.LOCAL so we have something to kvno +# against beyond just the krbtgt pair (no real KNOE.LOCAL services +# are deployed yet that aren't already covered by AC2). +kubectl --context=default -n knoe-system exec deploy/auth -- \ + kadmin.local -q 'addprinc -randkey test/myrddin.prole.org@KNOE.LOCAL' + +# From myrddin (which has a PROLE.ORG client config in /etc/krb5.conf +# and a working /etc/krb5.conf KNOE.LOCAL block courtesy of the +# kerberos_trust_setup.yml playbook): +kdestroy +kinit chrisfu@PROLE.ORG # enter your AD password +kvno krbtgt/KNOE.LOCAL@PROLE.ORG # cross-realm referral +kvno test/myrddin.prole.org@KNOE.LOCAL # full service ticket +klist +# EXPECT: both kvno calls succeed with no errors. klist shows three +# entries: +# chrisfu@PROLE.ORG (initial TGT) +# krbtgt/KNOE.LOCAL@PROLE.ORG (cross-realm TGT) +# test/myrddin.prole.org@KNOE.LOCAL (service ticket) + +# Cleanup +kubectl --context=default -n knoe-system exec deploy/auth -- \ + kadmin.local -q 'delprinc -force test/myrddin.prole.org@KNOE.LOCAL' +``` + +## If any AC fails + +Surface the failure with diagnostics rather than papering over it: + +| AC | If it fails, capture | +|---|---| +| AC1 | `kubectl get deploy -A | grep -i auth` (in case the live pod is named something other than `authority-knoe-auth`); the pod's full env: `kubectl -n knoe-system get deploy authority-knoe-auth -o yaml | grep -A2 PROLE_KDC_REALM`; the pod logs from the init phase: `kubectl logs deploy/auth` | +| AC2 | The init script's stdout from the pod logs: `kubectl logs deploy/auth` | +| AC3 | `kubectl get sc` + `kubectl describe pvc knoe-kdc-data`; the deployment events: `kubectl describe deploy authority-knoe-auth` | +| AC4 | `KRB5_TRACE=/dev/stdout kvno test/myrddin.prole.org@KNOE.LOCAL 2>&1`; `kubectl logs deploy/auth --tail=100` | + +**Do not** manually `kadmin.local addprinc` to paper over a missing +principal — that's exactly the trap of the previous attempt. If a +principal is missing, the bug is in `etc/init_kdc.sh` (or in its +environment) and we want to know. + +Also be alert for any leftover pods named `auth-*` or +`authority-prole-auth-*` from the pre-rebrand era. The `--reset` path +should wipe them, but if `kubectl get pods -n knoe-system` shows +multiple deployments still alive, delete the stale ones explicitly: + +```bash +kubectl -n knoe-system delete deploy auth authority-prole-auth --ignore-not-found +``` + +## Out of scope + +- GKE (`--mode gke`) deployment — different realm (`KNOE.DEV`), different + cluster, separate brief if needed. +- Image rebuild — every config the KDC pod consumes is generated by + `etc/init_kdc.sh` and applied via `kubectl apply`. No image work + required. +- The OIDC issuer (`auth.knoe.dev`) — that runs on GKE; this brief + is about the in-cluster k3s KDC only. +- HA storage for the KDC PVC — `local-path` (k3s default + StorageClass) is fine for now. A future task can move to longhorn + / synology iSCSI. + +## Commit shape + +One commit covering this brief plus any `etc/init_kdc.sh` follow-ups +needed during execution. Title: + +``` +kdc: verify install.sh --reset converges to a working cross-realm trust +``` + +Body should embed the AC1-AC4 outputs as proof of acceptance. + +## Reference: the cascade that led here + +The hand-debugging session on 2026-05-10/11 walked these layers, in +order, before reaching the root cause. Preserved so the next person +who lands at "PROCESS_TGS error" doesn't repeat it: + +1. "Cannot find KDC for realm" → workstation `/etc/krb5.conf` didn't + know about KNOE.LOCAL. Fixed by editing the workstation conf via + `infrastructure/playbooks/workstation_kerberos.yml`. +2. "Server not found" → kadmin.local writes were going to the wrong + pod. Service `auth` routed to `authority-knoe-auth-...`, not the + `auth-...` pod we'd been editing. +3. "Decrypt integrity check failed" → kvno mismatch between Samba's + re-keyed `krbtgt_KNOE.LOCAL` (kvno=4) and MIT's freshly-created + `krbtgt/KNOE.LOCAL@PROLE.ORG` (kvno=1). +4. "PROCESS_TGS" → MIT pod's realm was `PROLE.LOCAL`, not `KNOE.LOCAL`, + so even matching kvnos failed because the realm-of-decryption + didn't match the realm-of-issue. +5. `default_realm = PROLE.LOCAL` baked into the Deployment env → + came from the prole `etc/init_kdc.sh:327` shell default never being + updated post-rebrand. **THIS IS THE ROOT.** + +All of (1)–(4) above are downstream symptoms of (5). diff --git a/etc/init_kdc.sh b/etc/init_kdc.sh index 0390628..95bb096 100755 --- a/etc/init_kdc.sh +++ b/etc/init_kdc.sh @@ -567,6 +567,13 @@ EOF host_net_block=" hostNetwork: true" dns_policy_block=" dnsPolicy: ClusterFirstWithHostNet" fi + # PVC storage class: explicit when $PROLE_KDC_STORAGE_CLASS is set; + # otherwise leave blank so the cluster's default StorageClass picks + # the binder (k3s "local-path", GKE "standard", etc.). + local storage_class_block="" + if [[ -n "${PROLE_KDC_STORAGE_CLASS:-}" ]]; then + storage_class_block=" storageClassName: ${PROLE_KDC_STORAGE_CLASS}" + fi cat < + UPN) does not match MIT's + # ( + ). RC4 derives keys from + # the password alone, so both sides converge with no salt fight. + # ------------------------------------------------------------------ + + # Outbound: KNOE.LOCAL → PROLE.ORG (issued here, decrypted by Samba) if ! kadmin.local -q "get_principal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}" >/dev/null 2>&1; then - echo "Creating trust principal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}..." - kadmin.local -q "addprinc -pw \${shared_pw} krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}" + echo "Creating outbound trust principal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}..." + kadmin.local -q "addprinc -pw \${shared_pw} -e arcfour-hmac:normal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}" fi - if [[ -n "\${PROLE_KDC_TRUST_ADMIN:-}" && -n "\${PROLE_KDC_TRUST_PASSWORD:-}" ]]; then - echo "Creating reciprocal trust principal in \${PROLE_KDC_TRUST_REALM}..." - kadmin -r "\${PROLE_KDC_TRUST_REALM}" -p "\${PROLE_KDC_TRUST_ADMIN}" -w "\${PROLE_KDC_TRUST_PASSWORD}" \ - -q "addprinc -pw \${shared_pw} krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}" || true - else - echo "WARN: Missing PROLE_KDC_TRUST_ADMIN/PROLE_KDC_TRUST_PASSWORD; skipping external trust principal." + # Inbound: PROLE.ORG → KNOE.LOCAL (issued by Samba, decrypted here) + if ! kadmin.local -q "get_principal krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}" >/dev/null 2>&1; then + echo "Creating inbound trust principal krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}..." + kadmin.local -q "addprinc -pw \${shared_pw} -e arcfour-hmac:normal krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}" fi + + # NOTE: The Samba-side trust account (user "krbtgt_\${PROLE_KDC_REALM}" + # in PROLE.ORG with UPN/SPN krbtgt/\${PROLE_KDC_REALM}) is provisioned + # OUT-OF-BAND by this repo's Ansible playbook: + # infrastructure/playbooks/kerberos_trust_setup.yml + # Earlier versions of this script tried to use a remote "kadmin" + # client to write that principal into Samba, but Samba AD does not + # accept additions over MIT's kadmin protocol — it always failed + # with "Missing parameters in krb5.conf required for kadmin client". + # Run the playbook once after this KDC comes up: + # ANSIBLE_VAULT_PASSWORD_FILE=\$PWD/.vault_pass \\ + # ansible-playbook infrastructure/playbooks/kerberos_trust_setup.yml + echo "Note: Samba-side trust account is provisioned out-of-band by" + echo " infrastructure/playbooks/kerberos_trust_setup.yml" fi # Start daemons. Keep kadmind in PID 1; run krb5kdc in background and verify it binds. @@ -809,7 +840,21 @@ ${dns_policy_block} configMap: name: knoe-kdc-config - name: knoe-kdc-data - emptyDir: {} + persistentVolumeClaim: + claimName: knoe-kdc-data +--- +apiVersion: v1 +kind: PersistentVolumeClaim +metadata: + name: knoe-kdc-data + namespace: ${KDC_NAMESPACE} +spec: + accessModes: + - ReadWriteOnce + resources: + requests: + storage: ${PROLE_KDC_STORAGE_SIZE:-1Gi} +${storage_class_block} --- apiVersion: v1 kind: Service