prole/docs/plans/junie/kdc-trust-reset-repeatable.md
chrisfu d6586cf1d9 kdc: backport init_kdc.sh trust fixes + add reset-repeatable Junie brief
Mirrors the knoe-db commit `ff7546d` patches into the prole copy of
`etc/init_kdc.sh` so a re-run of `install.sh --mode k3s --reset` from
this repo produces a working cross-realm trust without manual cluster
surgery. The k3s cluster is provisioned from this repo, so the source
fix must live here (knoe-db remains canonical for GKE).

Changes to etc/init_kdc.sh:

1. Create BOTH cross-realm krbtgts in MIT, not just the outbound one.
   The inbound `krbtgt/<REALM>@<TRUST_REALM>` (issued by Samba,
   decrypted here) was missing entirely; without it, MIT cannot
   decrypt inbound TGTs and the trust never carries traffic.

2. Pin both cross-realm krbtgts to RC4 (`arcfour-hmac:normal`). AES
   keys depend on salt, and Samba's `<remote_realm>+UPN` salt does
   not match MIT's `<local_realm>+<principal-no-realm>`; RC4 has no
   salt so both sides converge from the password alone. Matches the
   already-pinned Samba side (commit `ad1eced`).

3. Replace the broken "remote kadmin to Samba" reciprocal-trust block
   with a documented no-op pointing at
   `infrastructure/playbooks/kerberos_trust_setup.yml`. Samba AD
   does not accept additions over MIT's kadmin protocol; the block
   always failed with "Missing parameters in krb5.conf required for
   kadmin client".

4. Switch the KDC data volume from emptyDir to a PVC
   (claimName `knoe-kdc-data`, parameterized by
   `$PROLE_KDC_STORAGE_SIZE` and `$PROLE_KDC_STORAGE_CLASS`).
   State now survives pod restarts.

Adds Junie brief `docs/plans/junie/kdc-trust-reset-repeatable.md`
with four TDD acceptance criteria for an end-to-end --reset run on
the prole k3s cluster.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-11 01:24:34 -07:00

9.0 KiB
Raw Blame History

Junie brief: end-to-end k3s reset/deploy that produces a working KNOE.LOCAL↔PROLE.ORG Kerberos trust

Why

A full afternoon on 2026-05-10/11 went into wiring up the PROLE.ORG ↔ KNOE.LOCAL cross-realm Kerberos trust on the prole k3s cluster (myrddin/merlin/gandalf). The Samba side now works (infrastructure/playbooks/kerberos_trust_setup.yml, commits 5cece40…ad1eced), but the in-cluster KDC pod was running for realm PROLE.LOCAL — a pre-rebrand artifact — and could not decrypt referral TGTs from Samba. Every "Server not found in Kerberos database" / "Decrypt integrity check failed" / "PROCESS_TGS" symptom we chased was downstream of that.

Root cause: the originally-deployed pod inherited PROLE_KDC_REALM=PROLE.LOCAL from a stale etc/init_kdc.sh shell default. The default was fixed in commit 4fe8647 today, and a set of larger patches to etc/init_kdc.sh followed.

What's not done: running install.sh --reset end-to-end against the k3s cluster to confirm the patches converge to a working trust from a blank slate. That's this brief.

What's already landed (do not re-do)

Commit Subject
4fe8647 etc/init_kdc.sh:327 shell default PROLE.LOCAL → KNOE.LOCAL
5cece40…227f490 infrastructure/playbooks/kerberos_trust_setup.yml — full Samba-side trust account provisioning
ad1eced kerberos_trust_setup.yml — msDS-SupportedEncryptionTypes 28 → 4 (RC4 only)
(also pending in this branch) etc/init_kdc.sh — both cross-realm krbtgts created in MIT with RC4, KDC volume switched from emptyDir to PVC, broken remote-kadmin block replaced with documentation pointing at the playbook

The matching source fix also landed in ~/dev/knoe-db (commit ff7546d) so the GKE / canonical installer stays in sync. This brief runs against the prole k3s cluster from this repo (~/dev/prole) — do NOT run it against the knoe-db repo.

What changes

Nothing more in source. Your job is to run the end-to-end reset-and-deploy from ~/dev/prole and verify it succeeds without manual cluster surgery.

Execution path

Sequencing constraint: install.sh --reset for --mode k3s runs the k3s_reset.yml Ansible playbook, which wipes and reinstalls k3s on all three k3s nodes (myrddin, merlin, gandalf). Schedule accordingly — anything currently running in that cluster goes away.

Before starting, capture the current state so you can roll back if needed:

kubectl --context=default -n knoe-system get pods,svc,pvc,configmap > /tmp/pre-reset.txt
kubectl --context=default -n knoe-system get deploy authority-knoe-auth -o yaml > /tmp/pre-deploy.yaml || true
# Run from a control node with kubectl access to the prole k3s cluster
# (myrddin itself is fine — it's the k3s server)

cd ~/dev/prole
git pull   # pick up the trust-fix commits if you don't have them locally

# 1. RESET — wipes k3s on all three nodes
./install.sh --mode k3s --reset

# 2. DEPLOY — re-provisions services including the KDC.
#    Watch the init_kdc.sh output for:
#      "Applying Knoe KDC '...' in namespace '...' (realm KNOE.LOCAL)"
#      "Creating outbound trust principal krbtgt/PROLE.ORG@KNOE.LOCAL"
#      "Creating inbound trust principal krbtgt/KNOE.LOCAL@PROLE.ORG"
#      "Note: Samba-side trust account is provisioned out-of-band by ..."
./install.sh --mode k3s

# 3. Sync the Samba-side trust account to the (possibly new) cluster
#    Secret value.
ANSIBLE_VAULT_PASSWORD_FILE=$PWD/.vault_pass \
  ansible-playbook infrastructure/playbooks/kerberos_trust_setup.yml

TDD acceptance criteria

All four must pass. Capture the output of each and paste into your MR description.

AC1 — KDC pod has realm KNOE.LOCAL

kubectl --context=default -n knoe-system exec deploy/auth -- \
  cat /etc/krb5.conf | grep default_realm
# EXPECT: default_realm = KNOE.LOCAL

AC2 — Both cross-realm krbtgts exist in MIT

kubectl --context=default -n knoe-system exec deploy/auth -- \
  kadmin.local -q 'listprincs' | grep -E '^krbtgt/'
# EXPECT (at minimum):
#   krbtgt/KNOE.LOCAL@KNOE.LOCAL
#   krbtgt/KNOE.LOCAL@PROLE.ORG
#   krbtgt/PROLE.ORG@KNOE.LOCAL

AC3 — KDC data PVC is bound

kubectl --context=default -n knoe-system get pvc knoe-kdc-data
# EXPECT: STATUS=Bound

If you see Pending, look at the StorageClass:

kubectl get sc
kubectl --context=default -n knoe-system describe pvc knoe-kdc-data

The default k3s StorageClass local-path should be sufficient. If your cluster doesn't have a default SC, re-run install with PROLE_KDC_STORAGE_CLASS=local-path in the environment.

AC4 — End-to-end cross-realm ticket flow

# Pre-stage a test SPN in KNOE.LOCAL so we have something to kvno
# against beyond just the krbtgt pair (no real KNOE.LOCAL services
# are deployed yet that aren't already covered by AC2).
kubectl --context=default -n knoe-system exec deploy/auth -- \
  kadmin.local -q 'addprinc -randkey test/myrddin.prole.org@KNOE.LOCAL'

# From myrddin (which has a PROLE.ORG client config in /etc/krb5.conf
# and a working /etc/krb5.conf KNOE.LOCAL block courtesy of the
# kerberos_trust_setup.yml playbook):
kdestroy
kinit chrisfu@PROLE.ORG       # enter your AD password
kvno krbtgt/KNOE.LOCAL@PROLE.ORG       # cross-realm referral
kvno test/myrddin.prole.org@KNOE.LOCAL # full service ticket
klist
# EXPECT: both kvno calls succeed with no errors. klist shows three
#         entries:
#           chrisfu@PROLE.ORG                  (initial TGT)
#           krbtgt/KNOE.LOCAL@PROLE.ORG        (cross-realm TGT)
#           test/myrddin.prole.org@KNOE.LOCAL  (service ticket)

# Cleanup
kubectl --context=default -n knoe-system exec deploy/auth -- \
  kadmin.local -q 'delprinc -force test/myrddin.prole.org@KNOE.LOCAL'

If any AC fails

Surface the failure with diagnostics rather than papering over it:

AC If it fails, capture
AC1 `kubectl get deploy -A
AC2 The init script's stdout from the pod logs: kubectl logs deploy/auth
AC3 kubectl get sc + kubectl describe pvc knoe-kdc-data; the deployment events: kubectl describe deploy authority-knoe-auth
AC4 KRB5_TRACE=/dev/stdout kvno test/myrddin.prole.org@KNOE.LOCAL 2>&1; kubectl logs deploy/auth --tail=100

Do not manually kadmin.local addprinc to paper over a missing principal — that's exactly the trap of the previous attempt. If a principal is missing, the bug is in etc/init_kdc.sh (or in its environment) and we want to know.

Also be alert for any leftover pods named auth-* or authority-prole-auth-* from the pre-rebrand era. The --reset path should wipe them, but if kubectl get pods -n knoe-system shows multiple deployments still alive, delete the stale ones explicitly:

kubectl -n knoe-system delete deploy auth authority-prole-auth --ignore-not-found

Out of scope

  • GKE (--mode gke) deployment — different realm (KNOE.DEV), different cluster, separate brief if needed.
  • Image rebuild — every config the KDC pod consumes is generated by etc/init_kdc.sh and applied via kubectl apply. No image work required.
  • The OIDC issuer (auth.knoe.dev) — that runs on GKE; this brief is about the in-cluster k3s KDC only.
  • HA storage for the KDC PVC — local-path (k3s default StorageClass) is fine for now. A future task can move to longhorn / synology iSCSI.

Commit shape

One commit covering this brief plus any etc/init_kdc.sh follow-ups needed during execution. Title:

kdc: verify install.sh --reset converges to a working cross-realm trust

Body should embed the AC1-AC4 outputs as proof of acceptance.

Reference: the cascade that led here

The hand-debugging session on 2026-05-10/11 walked these layers, in order, before reaching the root cause. Preserved so the next person who lands at "PROCESS_TGS error" doesn't repeat it:

  1. "Cannot find KDC for realm" → workstation /etc/krb5.conf didn't know about KNOE.LOCAL. Fixed by editing the workstation conf via infrastructure/playbooks/workstation_kerberos.yml.
  2. "Server not found" → kadmin.local writes were going to the wrong pod. Service auth routed to authority-knoe-auth-..., not the auth-... pod we'd been editing.
  3. "Decrypt integrity check failed" → kvno mismatch between Samba's re-keyed krbtgt_KNOE.LOCAL (kvno=4) and MIT's freshly-created krbtgt/KNOE.LOCAL@PROLE.ORG (kvno=1).
  4. "PROCESS_TGS" → MIT pod's realm was PROLE.LOCAL, not KNOE.LOCAL, so even matching kvnos failed because the realm-of-decryption didn't match the realm-of-issue.
  5. default_realm = PROLE.LOCAL baked into the Deployment env → came from the prole etc/init_kdc.sh:327 shell default never being updated post-rebrand. THIS IS THE ROOT.

All of (1)–(4) above are downstream symptoms of (5).