kdc: backport init_kdc.sh trust fixes + add reset-repeatable Junie brief

Mirrors the knoe-db commit `ff7546d` patches into the prole copy of
`etc/init_kdc.sh` so a re-run of `install.sh --mode k3s --reset` from
this repo produces a working cross-realm trust without manual cluster
surgery. The k3s cluster is provisioned from this repo, so the source
fix must live here (knoe-db remains canonical for GKE).

Changes to etc/init_kdc.sh:

1. Create BOTH cross-realm krbtgts in MIT, not just the outbound one.
   The inbound `krbtgt/<REALM>@<TRUST_REALM>` (issued by Samba,
   decrypted here) was missing entirely; without it, MIT cannot
   decrypt inbound TGTs and the trust never carries traffic.

2. Pin both cross-realm krbtgts to RC4 (`arcfour-hmac:normal`). AES
   keys depend on salt, and Samba's `<remote_realm>+UPN` salt does
   not match MIT's `<local_realm>+<principal-no-realm>`; RC4 has no
   salt so both sides converge from the password alone. Matches the
   already-pinned Samba side (commit `ad1eced`).

3. Replace the broken "remote kadmin to Samba" reciprocal-trust block
   with a documented no-op pointing at
   `infrastructure/playbooks/kerberos_trust_setup.yml`. Samba AD
   does not accept additions over MIT's kadmin protocol; the block
   always failed with "Missing parameters in krb5.conf required for
   kadmin client".

4. Switch the KDC data volume from emptyDir to a PVC
   (claimName `knoe-kdc-data`, parameterized by
   `$PROLE_KDC_STORAGE_SIZE` and `$PROLE_KDC_STORAGE_CLASS`).
   State now survives pod restarts.

Adds Junie brief `docs/plans/junie/kdc-trust-reset-repeatable.md`
with four TDD acceptance criteria for an end-to-end --reset run on
the prole k3s cluster.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
chrisfu 2026-05-11 01:24:34 -07:00
parent ad1ecedb5e
commit d6586cf1d9
2 changed files with 278 additions and 9 deletions

View File

@ -0,0 +1,224 @@
# Junie brief: end-to-end k3s reset/deploy that produces a working KNOE.LOCAL↔PROLE.ORG Kerberos trust
## Why
A full afternoon on 2026-05-10/11 went into wiring up the
PROLE.ORG ↔ KNOE.LOCAL cross-realm Kerberos trust on the prole k3s
cluster (myrddin/merlin/gandalf). The Samba side now works
(`infrastructure/playbooks/kerberos_trust_setup.yml`, commits
`5cece40`…`ad1eced`), but the in-cluster KDC pod was running for realm
**`PROLE.LOCAL`** — a pre-rebrand artifact — and could not decrypt
referral TGTs from Samba. Every "Server not found in Kerberos
database" / "Decrypt integrity check failed" / "PROCESS_TGS" symptom
we chased was downstream of that.
Root cause: the originally-deployed pod inherited
`PROLE_KDC_REALM=PROLE.LOCAL` from a stale `etc/init_kdc.sh` shell
default. The default was fixed in commit `4fe8647` today, and a set
of larger patches to `etc/init_kdc.sh` followed.
What's **not** done: running `install.sh --reset` end-to-end against
the k3s cluster to confirm the patches converge to a working trust
from a blank slate. That's this brief.
## What's already landed (do not re-do)
| Commit | Subject |
|---|---|
| `4fe8647` | `etc/init_kdc.sh:327` shell default `PROLE.LOCAL``KNOE.LOCAL` |
| `5cece40`…`227f490` | `infrastructure/playbooks/kerberos_trust_setup.yml` — full Samba-side trust account provisioning |
| `ad1eced` | `kerberos_trust_setup.yml``msDS-SupportedEncryptionTypes` 28 → 4 (RC4 only) |
| _(also pending in this branch)_ | `etc/init_kdc.sh` — both cross-realm krbtgts created in MIT with RC4, KDC volume switched from emptyDir to PVC, broken remote-kadmin block replaced with documentation pointing at the playbook |
The matching source fix also landed in `~/dev/knoe-db` (commit
`ff7546d`) so the GKE / canonical installer stays in sync. **This
brief runs against the prole k3s cluster from this repo (`~/dev/prole`)
— do NOT run it against the knoe-db repo.**
## What changes
Nothing more in source. Your job is to **run** the end-to-end
reset-and-deploy from `~/dev/prole` and verify it succeeds without
manual cluster surgery.
## Execution path
> **Sequencing constraint:** `install.sh --reset` for `--mode k3s` runs
> the `k3s_reset.yml` Ansible playbook, which **wipes and reinstalls
> k3s on all three k3s nodes (myrddin, merlin, gandalf)**. Schedule
> accordingly — anything currently running in that cluster goes away.
>
> Before starting, capture the current state so you can roll back if
> needed:
>
> ```bash
> kubectl --context=default -n knoe-system get pods,svc,pvc,configmap > /tmp/pre-reset.txt
> kubectl --context=default -n knoe-system get deploy authority-knoe-auth -o yaml > /tmp/pre-deploy.yaml || true
> ```
```bash
# Run from a control node with kubectl access to the prole k3s cluster
# (myrddin itself is fine — it's the k3s server)
cd ~/dev/prole
git pull # pick up the trust-fix commits if you don't have them locally
# 1. RESET — wipes k3s on all three nodes
./install.sh --mode k3s --reset
# 2. DEPLOY — re-provisions services including the KDC.
# Watch the init_kdc.sh output for:
# "Applying Knoe KDC '...' in namespace '...' (realm KNOE.LOCAL)"
# "Creating outbound trust principal krbtgt/PROLE.ORG@KNOE.LOCAL"
# "Creating inbound trust principal krbtgt/KNOE.LOCAL@PROLE.ORG"
# "Note: Samba-side trust account is provisioned out-of-band by ..."
./install.sh --mode k3s
# 3. Sync the Samba-side trust account to the (possibly new) cluster
# Secret value.
ANSIBLE_VAULT_PASSWORD_FILE=$PWD/.vault_pass \
ansible-playbook infrastructure/playbooks/kerberos_trust_setup.yml
```
## TDD acceptance criteria
All four must pass. Capture the output of each and paste into your
MR description.
### AC1 — KDC pod has realm `KNOE.LOCAL`
```bash
kubectl --context=default -n knoe-system exec deploy/auth -- \
cat /etc/krb5.conf | grep default_realm
# EXPECT: default_realm = KNOE.LOCAL
```
### AC2 — Both cross-realm krbtgts exist in MIT
```bash
kubectl --context=default -n knoe-system exec deploy/auth -- \
kadmin.local -q 'listprincs' | grep -E '^krbtgt/'
# EXPECT (at minimum):
# krbtgt/KNOE.LOCAL@KNOE.LOCAL
# krbtgt/KNOE.LOCAL@PROLE.ORG
# krbtgt/PROLE.ORG@KNOE.LOCAL
```
### AC3 — KDC data PVC is bound
```bash
kubectl --context=default -n knoe-system get pvc knoe-kdc-data
# EXPECT: STATUS=Bound
```
If you see `Pending`, look at the StorageClass:
```bash
kubectl get sc
kubectl --context=default -n knoe-system describe pvc knoe-kdc-data
```
The default k3s StorageClass `local-path` should be sufficient. If
your cluster doesn't have a default SC, re-run install with
`PROLE_KDC_STORAGE_CLASS=local-path` in the environment.
### AC4 — End-to-end cross-realm ticket flow
```bash
# Pre-stage a test SPN in KNOE.LOCAL so we have something to kvno
# against beyond just the krbtgt pair (no real KNOE.LOCAL services
# are deployed yet that aren't already covered by AC2).
kubectl --context=default -n knoe-system exec deploy/auth -- \
kadmin.local -q 'addprinc -randkey test/myrddin.prole.org@KNOE.LOCAL'
# From myrddin (which has a PROLE.ORG client config in /etc/krb5.conf
# and a working /etc/krb5.conf KNOE.LOCAL block courtesy of the
# kerberos_trust_setup.yml playbook):
kdestroy
kinit chrisfu@PROLE.ORG # enter your AD password
kvno krbtgt/KNOE.LOCAL@PROLE.ORG # cross-realm referral
kvno test/myrddin.prole.org@KNOE.LOCAL # full service ticket
klist
# EXPECT: both kvno calls succeed with no errors. klist shows three
# entries:
# chrisfu@PROLE.ORG (initial TGT)
# krbtgt/KNOE.LOCAL@PROLE.ORG (cross-realm TGT)
# test/myrddin.prole.org@KNOE.LOCAL (service ticket)
# Cleanup
kubectl --context=default -n knoe-system exec deploy/auth -- \
kadmin.local -q 'delprinc -force test/myrddin.prole.org@KNOE.LOCAL'
```
## If any AC fails
Surface the failure with diagnostics rather than papering over it:
| AC | If it fails, capture |
|---|---|
| AC1 | `kubectl get deploy -A | grep -i auth` (in case the live pod is named something other than `authority-knoe-auth`); the pod's full env: `kubectl -n knoe-system get deploy authority-knoe-auth -o yaml | grep -A2 PROLE_KDC_REALM`; the pod logs from the init phase: `kubectl logs deploy/auth` |
| AC2 | The init script's stdout from the pod logs: `kubectl logs deploy/auth` |
| AC3 | `kubectl get sc` + `kubectl describe pvc knoe-kdc-data`; the deployment events: `kubectl describe deploy authority-knoe-auth` |
| AC4 | `KRB5_TRACE=/dev/stdout kvno test/myrddin.prole.org@KNOE.LOCAL 2>&1`; `kubectl logs deploy/auth --tail=100` |
**Do not** manually `kadmin.local addprinc` to paper over a missing
principal — that's exactly the trap of the previous attempt. If a
principal is missing, the bug is in `etc/init_kdc.sh` (or in its
environment) and we want to know.
Also be alert for any leftover pods named `auth-*` or
`authority-prole-auth-*` from the pre-rebrand era. The `--reset` path
should wipe them, but if `kubectl get pods -n knoe-system` shows
multiple deployments still alive, delete the stale ones explicitly:
```bash
kubectl -n knoe-system delete deploy auth authority-prole-auth --ignore-not-found
```
## Out of scope
- GKE (`--mode gke`) deployment — different realm (`KNOE.DEV`), different
cluster, separate brief if needed.
- Image rebuild — every config the KDC pod consumes is generated by
`etc/init_kdc.sh` and applied via `kubectl apply`. No image work
required.
- The OIDC issuer (`auth.knoe.dev`) — that runs on GKE; this brief
is about the in-cluster k3s KDC only.
- HA storage for the KDC PVC — `local-path` (k3s default
StorageClass) is fine for now. A future task can move to longhorn
/ synology iSCSI.
## Commit shape
One commit covering this brief plus any `etc/init_kdc.sh` follow-ups
needed during execution. Title:
```
kdc: verify install.sh --reset converges to a working cross-realm trust
```
Body should embed the AC1-AC4 outputs as proof of acceptance.
## Reference: the cascade that led here
The hand-debugging session on 2026-05-10/11 walked these layers, in
order, before reaching the root cause. Preserved so the next person
who lands at "PROCESS_TGS error" doesn't repeat it:
1. "Cannot find KDC for realm" → workstation `/etc/krb5.conf` didn't
know about KNOE.LOCAL. Fixed by editing the workstation conf via
`infrastructure/playbooks/workstation_kerberos.yml`.
2. "Server not found" → kadmin.local writes were going to the wrong
pod. Service `auth` routed to `authority-knoe-auth-...`, not the
`auth-...` pod we'd been editing.
3. "Decrypt integrity check failed" → kvno mismatch between Samba's
re-keyed `krbtgt_KNOE.LOCAL` (kvno=4) and MIT's freshly-created
`krbtgt/KNOE.LOCAL@PROLE.ORG` (kvno=1).
4. "PROCESS_TGS" → MIT pod's realm was `PROLE.LOCAL`, not `KNOE.LOCAL`,
so even matching kvnos failed because the realm-of-decryption
didn't match the realm-of-issue.
5. `default_realm = PROLE.LOCAL` baked into the Deployment env →
came from the prole `etc/init_kdc.sh:327` shell default never being
updated post-rebrand. **THIS IS THE ROOT.**
All of (1)(4) above are downstream symptoms of (5).

View File

@ -567,6 +567,13 @@ EOF
host_net_block=" hostNetwork: true"
dns_policy_block=" dnsPolicy: ClusterFirstWithHostNet"
fi
# PVC storage class: explicit when $PROLE_KDC_STORAGE_CLASS is set;
# otherwise leave blank so the cluster's default StorageClass picks
# the binder (k3s "local-path", GKE "standard", etc.).
local storage_class_block=""
if [[ -n "${PROLE_KDC_STORAGE_CLASS:-}" ]]; then
storage_class_block=" storageClassName: ${PROLE_KDC_STORAGE_CLASS}"
fi
cat <<EOF | kubectl apply -n "$KDC_NAMESPACE" -f -
apiVersion: v1
kind: ConfigMap
@ -696,18 +703,42 @@ data:
if [[ -n "\${PROLE_KDC_TRUST_REALM:-}" && "\${PROLE_KDC_TRUST_REALM}" != "\${PROLE_KDC_REALM}" ]]; then
shared_pw="\${PROLE_KDC_TRUST_SHARED_PASSWORD:-\${PROLE_KDC_MASTER_PASSWORD}}"
# ------------------------------------------------------------------
# Cross-realm krbtgt principals — RC4 only.
#
# Both directions of the trust live as their own krbtgt principal,
# each keyed to the same shared password. We pin RC4 (arcfour-hmac)
# because AES key derivation requires a salt, and Samba's salt
# convention (<remote_realm> + UPN) does not match MIT's
# (<local_realm> + <principal-no-realm>). RC4 derives keys from
# the password alone, so both sides converge with no salt fight.
# ------------------------------------------------------------------
# Outbound: KNOE.LOCAL → PROLE.ORG (issued here, decrypted by Samba)
if ! kadmin.local -q "get_principal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}" >/dev/null 2>&1; then
echo "Creating trust principal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}..."
kadmin.local -q "addprinc -pw \${shared_pw} krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}"
echo "Creating outbound trust principal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}..."
kadmin.local -q "addprinc -pw \${shared_pw} -e arcfour-hmac:normal krbtgt/\${PROLE_KDC_TRUST_REALM}@\${PROLE_KDC_REALM}"
fi
if [[ -n "\${PROLE_KDC_TRUST_ADMIN:-}" && -n "\${PROLE_KDC_TRUST_PASSWORD:-}" ]]; then
echo "Creating reciprocal trust principal in \${PROLE_KDC_TRUST_REALM}..."
kadmin -r "\${PROLE_KDC_TRUST_REALM}" -p "\${PROLE_KDC_TRUST_ADMIN}" -w "\${PROLE_KDC_TRUST_PASSWORD}" \
-q "addprinc -pw \${shared_pw} krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}" || true
else
echo "WARN: Missing PROLE_KDC_TRUST_ADMIN/PROLE_KDC_TRUST_PASSWORD; skipping external trust principal."
# Inbound: PROLE.ORG → KNOE.LOCAL (issued by Samba, decrypted here)
if ! kadmin.local -q "get_principal krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}" >/dev/null 2>&1; then
echo "Creating inbound trust principal krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}..."
kadmin.local -q "addprinc -pw \${shared_pw} -e arcfour-hmac:normal krbtgt/\${PROLE_KDC_REALM}@\${PROLE_KDC_TRUST_REALM}"
fi
# NOTE: The Samba-side trust account (user "krbtgt_\${PROLE_KDC_REALM}"
# in PROLE.ORG with UPN/SPN krbtgt/\${PROLE_KDC_REALM}) is provisioned
# OUT-OF-BAND by this repo's Ansible playbook:
# infrastructure/playbooks/kerberos_trust_setup.yml
# Earlier versions of this script tried to use a remote "kadmin"
# client to write that principal into Samba, but Samba AD does not
# accept additions over MIT's kadmin protocol — it always failed
# with "Missing parameters in krb5.conf required for kadmin client".
# Run the playbook once after this KDC comes up:
# ANSIBLE_VAULT_PASSWORD_FILE=\$PWD/.vault_pass \\
# ansible-playbook infrastructure/playbooks/kerberos_trust_setup.yml
echo "Note: Samba-side trust account is provisioned out-of-band by"
echo " infrastructure/playbooks/kerberos_trust_setup.yml"
fi
# Start daemons. Keep kadmind in PID 1; run krb5kdc in background and verify it binds.
@ -809,7 +840,21 @@ ${dns_policy_block}
configMap:
name: knoe-kdc-config
- name: knoe-kdc-data
emptyDir: {}
persistentVolumeClaim:
claimName: knoe-kdc-data
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: knoe-kdc-data
namespace: ${KDC_NAMESPACE}
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: ${PROLE_KDC_STORAGE_SIZE:-1Gi}
${storage_class_block}
---
apiVersion: v1
kind: Service