Mirrors the knoe-db commit `ff7546d` patches into the prole copy of `etc/init_kdc.sh` so a re-run of `install.sh --mode k3s --reset` from this repo produces a working cross-realm trust without manual cluster surgery. The k3s cluster is provisioned from this repo, so the source fix must live here (knoe-db remains canonical for GKE). Changes to etc/init_kdc.sh: 1. Create BOTH cross-realm krbtgts in MIT, not just the outbound one. The inbound `krbtgt/<REALM>@<TRUST_REALM>` (issued by Samba, decrypted here) was missing entirely; without it, MIT cannot decrypt inbound TGTs and the trust never carries traffic. 2. Pin both cross-realm krbtgts to RC4 (`arcfour-hmac:normal`). AES keys depend on salt, and Samba's `<remote_realm>+UPN` salt does not match MIT's `<local_realm>+<principal-no-realm>`; RC4 has no salt so both sides converge from the password alone. Matches the already-pinned Samba side (commit `ad1eced`). 3. Replace the broken "remote kadmin to Samba" reciprocal-trust block with a documented no-op pointing at `infrastructure/playbooks/kerberos_trust_setup.yml`. Samba AD does not accept additions over MIT's kadmin protocol; the block always failed with "Missing parameters in krb5.conf required for kadmin client". 4. Switch the KDC data volume from emptyDir to a PVC (claimName `knoe-kdc-data`, parameterized by `$PROLE_KDC_STORAGE_SIZE` and `$PROLE_KDC_STORAGE_CLASS`). State now survives pod restarts. Adds Junie brief `docs/plans/junie/kdc-trust-reset-repeatable.md` with four TDD acceptance criteria for an end-to-end --reset run on the prole k3s cluster. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
9.0 KiB
Junie brief: end-to-end k3s reset/deploy that produces a working KNOE.LOCAL↔PROLE.ORG Kerberos trust
Why
A full afternoon on 2026-05-10/11 went into wiring up the
PROLE.ORG ↔ KNOE.LOCAL cross-realm Kerberos trust on the prole k3s
cluster (myrddin/merlin/gandalf). The Samba side now works
(infrastructure/playbooks/kerberos_trust_setup.yml, commits
5cece40…ad1eced), but the in-cluster KDC pod was running for realm
PROLE.LOCAL — a pre-rebrand artifact — and could not decrypt
referral TGTs from Samba. Every "Server not found in Kerberos
database" / "Decrypt integrity check failed" / "PROCESS_TGS" symptom
we chased was downstream of that.
Root cause: the originally-deployed pod inherited
PROLE_KDC_REALM=PROLE.LOCAL from a stale etc/init_kdc.sh shell
default. The default was fixed in commit 4fe8647 today, and a set
of larger patches to etc/init_kdc.sh followed.
What's not done: running install.sh --reset end-to-end against
the k3s cluster to confirm the patches converge to a working trust
from a blank slate. That's this brief.
What's already landed (do not re-do)
| Commit | Subject |
|---|---|
4fe8647 |
etc/init_kdc.sh:327 shell default PROLE.LOCAL → KNOE.LOCAL |
5cece40…227f490 |
infrastructure/playbooks/kerberos_trust_setup.yml — full Samba-side trust account provisioning |
ad1eced |
kerberos_trust_setup.yml — msDS-SupportedEncryptionTypes 28 → 4 (RC4 only) |
| (also pending in this branch) | etc/init_kdc.sh — both cross-realm krbtgts created in MIT with RC4, KDC volume switched from emptyDir to PVC, broken remote-kadmin block replaced with documentation pointing at the playbook |
The matching source fix also landed in ~/dev/knoe-db (commit
ff7546d) so the GKE / canonical installer stays in sync. This
brief runs against the prole k3s cluster from this repo (~/dev/prole)
— do NOT run it against the knoe-db repo.
What changes
Nothing more in source. Your job is to run the end-to-end
reset-and-deploy from ~/dev/prole and verify it succeeds without
manual cluster surgery.
Execution path
Sequencing constraint:
install.sh --resetfor--mode k3sruns thek3s_reset.ymlAnsible playbook, which wipes and reinstalls k3s on all three k3s nodes (myrddin, merlin, gandalf). Schedule accordingly — anything currently running in that cluster goes away.Before starting, capture the current state so you can roll back if needed:
kubectl --context=default -n knoe-system get pods,svc,pvc,configmap > /tmp/pre-reset.txt kubectl --context=default -n knoe-system get deploy authority-knoe-auth -o yaml > /tmp/pre-deploy.yaml || true
# Run from a control node with kubectl access to the prole k3s cluster
# (myrddin itself is fine — it's the k3s server)
cd ~/dev/prole
git pull # pick up the trust-fix commits if you don't have them locally
# 1. RESET — wipes k3s on all three nodes
./install.sh --mode k3s --reset
# 2. DEPLOY — re-provisions services including the KDC.
# Watch the init_kdc.sh output for:
# "Applying Knoe KDC '...' in namespace '...' (realm KNOE.LOCAL)"
# "Creating outbound trust principal krbtgt/PROLE.ORG@KNOE.LOCAL"
# "Creating inbound trust principal krbtgt/KNOE.LOCAL@PROLE.ORG"
# "Note: Samba-side trust account is provisioned out-of-band by ..."
./install.sh --mode k3s
# 3. Sync the Samba-side trust account to the (possibly new) cluster
# Secret value.
ANSIBLE_VAULT_PASSWORD_FILE=$PWD/.vault_pass \
ansible-playbook infrastructure/playbooks/kerberos_trust_setup.yml
TDD acceptance criteria
All four must pass. Capture the output of each and paste into your MR description.
AC1 — KDC pod has realm KNOE.LOCAL
kubectl --context=default -n knoe-system exec deploy/auth -- \
cat /etc/krb5.conf | grep default_realm
# EXPECT: default_realm = KNOE.LOCAL
AC2 — Both cross-realm krbtgts exist in MIT
kubectl --context=default -n knoe-system exec deploy/auth -- \
kadmin.local -q 'listprincs' | grep -E '^krbtgt/'
# EXPECT (at minimum):
# krbtgt/KNOE.LOCAL@KNOE.LOCAL
# krbtgt/KNOE.LOCAL@PROLE.ORG
# krbtgt/PROLE.ORG@KNOE.LOCAL
AC3 — KDC data PVC is bound
kubectl --context=default -n knoe-system get pvc knoe-kdc-data
# EXPECT: STATUS=Bound
If you see Pending, look at the StorageClass:
kubectl get sc
kubectl --context=default -n knoe-system describe pvc knoe-kdc-data
The default k3s StorageClass local-path should be sufficient. If
your cluster doesn't have a default SC, re-run install with
PROLE_KDC_STORAGE_CLASS=local-path in the environment.
AC4 — End-to-end cross-realm ticket flow
# Pre-stage a test SPN in KNOE.LOCAL so we have something to kvno
# against beyond just the krbtgt pair (no real KNOE.LOCAL services
# are deployed yet that aren't already covered by AC2).
kubectl --context=default -n knoe-system exec deploy/auth -- \
kadmin.local -q 'addprinc -randkey test/myrddin.prole.org@KNOE.LOCAL'
# From myrddin (which has a PROLE.ORG client config in /etc/krb5.conf
# and a working /etc/krb5.conf KNOE.LOCAL block courtesy of the
# kerberos_trust_setup.yml playbook):
kdestroy
kinit chrisfu@PROLE.ORG # enter your AD password
kvno krbtgt/KNOE.LOCAL@PROLE.ORG # cross-realm referral
kvno test/myrddin.prole.org@KNOE.LOCAL # full service ticket
klist
# EXPECT: both kvno calls succeed with no errors. klist shows three
# entries:
# chrisfu@PROLE.ORG (initial TGT)
# krbtgt/KNOE.LOCAL@PROLE.ORG (cross-realm TGT)
# test/myrddin.prole.org@KNOE.LOCAL (service ticket)
# Cleanup
kubectl --context=default -n knoe-system exec deploy/auth -- \
kadmin.local -q 'delprinc -force test/myrddin.prole.org@KNOE.LOCAL'
If any AC fails
Surface the failure with diagnostics rather than papering over it:
| AC | If it fails, capture |
|---|---|
| AC1 | `kubectl get deploy -A |
| AC2 | The init script's stdout from the pod logs: kubectl logs deploy/auth |
| AC3 | kubectl get sc + kubectl describe pvc knoe-kdc-data; the deployment events: kubectl describe deploy authority-knoe-auth |
| AC4 | KRB5_TRACE=/dev/stdout kvno test/myrddin.prole.org@KNOE.LOCAL 2>&1; kubectl logs deploy/auth --tail=100 |
Do not manually kadmin.local addprinc to paper over a missing
principal — that's exactly the trap of the previous attempt. If a
principal is missing, the bug is in etc/init_kdc.sh (or in its
environment) and we want to know.
Also be alert for any leftover pods named auth-* or
authority-prole-auth-* from the pre-rebrand era. The --reset path
should wipe them, but if kubectl get pods -n knoe-system shows
multiple deployments still alive, delete the stale ones explicitly:
kubectl -n knoe-system delete deploy auth authority-prole-auth --ignore-not-found
Out of scope
- GKE (
--mode gke) deployment — different realm (KNOE.DEV), different cluster, separate brief if needed. - Image rebuild — every config the KDC pod consumes is generated by
etc/init_kdc.shand applied viakubectl apply. No image work required. - The OIDC issuer (
auth.knoe.dev) — that runs on GKE; this brief is about the in-cluster k3s KDC only. - HA storage for the KDC PVC —
local-path(k3s default StorageClass) is fine for now. A future task can move to longhorn / synology iSCSI.
Commit shape
One commit covering this brief plus any etc/init_kdc.sh follow-ups
needed during execution. Title:
kdc: verify install.sh --reset converges to a working cross-realm trust
Body should embed the AC1-AC4 outputs as proof of acceptance.
Reference: the cascade that led here
The hand-debugging session on 2026-05-10/11 walked these layers, in order, before reaching the root cause. Preserved so the next person who lands at "PROCESS_TGS error" doesn't repeat it:
- "Cannot find KDC for realm" → workstation
/etc/krb5.confdidn't know about KNOE.LOCAL. Fixed by editing the workstation conf viainfrastructure/playbooks/workstation_kerberos.yml. - "Server not found" → kadmin.local writes were going to the wrong
pod. Service
authrouted toauthority-knoe-auth-..., not theauth-...pod we'd been editing. - "Decrypt integrity check failed" → kvno mismatch between Samba's
re-keyed
krbtgt_KNOE.LOCAL(kvno=4) and MIT's freshly-createdkrbtgt/KNOE.LOCAL@PROLE.ORG(kvno=1). - "PROCESS_TGS" → MIT pod's realm was
PROLE.LOCAL, notKNOE.LOCAL, so even matching kvnos failed because the realm-of-decryption didn't match the realm-of-issue. default_realm = PROLE.LOCALbaked into the Deployment env → came from the proleetc/init_kdc.sh:327shell default never being updated post-rebrand. THIS IS THE ROOT.
All of (1)–(4) above are downstream symptoms of (5).