mirror of
https://github.com/dredx/prole.git
synced 2026-09-28 05:44:31 +00:00
df6de9138e
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4a8d9cc90d |
feat: full GKE/prod deployment pipeline from UI to Artifact Registry
## GCP / Cluster Environment Screen - Auto-populate Cloud tab from conf/prod/gcp.cfg on screen open (org_id, billing_account, billing_project, project_id) - gcloud auth validity checked on screen startup; friendly modal dialog streams gcloud auth login output live so user never leaves the app - Live GKE cluster browser: fetches clusters via gcloud container clusters list, displays with checkmark selector, auto-selects saved cluster - Selecting a cluster runs get-credentials, sets KUBECONFIG/KUBECONTEXT, and syncs the region dropdown to the selected cluster's location - Region dropdown populated live from gcloud compute regions list with checkmark on currently selected region; graceful fallback when offline - New 'GCP Storage' tab with workload->StorageClass mapping (CNPG->premium-rwo, Redis/Monitoring->standard-rwo, Garage->garage-hdd) and Fetch from Cluster - Provider readonly field styled correctly (no solid-black on macOS) - Stale prole.cfg/conf/prole.cfg symlinks removed; all config I/O now resolves env-specific paths via prole_conf.entrypoint_path() ## GKE Autopilot Compatibility (Common Services) - Synology iSCSI StorageClass and static PVs guarded behind PROLE_MODE!=k8s in init_openbao.sh (GKE Autopilot forbids hostPath/iSCSI volumes) - In-cluster Docker registry (hostPath) skipped in k8s mode; GCP Artifact Registry used instead - Kong renamed knoe-svc-kong in k8s mode; all health-check kubectl calls in init_common_services.sh and status_common_services.sh updated accordingly - DNS endpoints switched from *.prole.org to *.knoe.dev in k8s mode (api.knoe.dev, git.knoe.dev, svc.knoe.dev); ingress uses gce class - New GKE-clean Kong manifests under deploy/opentofu/k8s/manifests/prole/: no k3s node affinity, explicit Autopilot resource requests/limits ## Garage S3 Store (GKE) - New garage-statefulset-gcp.yaml targeting garage-hdd StorageClass (pd-standard, avoids SSD_TOTAL_GB quota exhaustion in us-west3) - New storageclass-gcp-hdd.yaml (pd-standard, Retain, WaitForFirstConsumer) - GCP StorageClass manifests skipped on re-runs (Autopilot built-ins are immutable; skip-if-exists guard added) - PVC deletion guard extended to cover any storageClass (not just synology) so stale claims are cleaned before StatefulSet recreation ## Topology (GKE Autopilot) - DaemonSet collector skipped in prod mode (forbidden in kube-system by GKE Warden); Kubernetes-only node facts path used instead - All ready GKE nodes assumed cnpg-eligible and monitoring-eligible without taint/synology-mount checks (skip_collector + assume_nodes_eligible flags) ## KUBECONFIG / kubectl (k8s mode) - actions.py: new elif mode==k8s branch sets KUBECONFIG=~/.kube/config and injects KUBECONTEXT from prole_cfg_data into script env - _build_kubectl_cmd falls back to Global.KUBECONTEXT when init_cluster.selected_kubectx is empty - _activate_selected_gke_cluster persists KUBECONFIG/KUBECONTEXT to prole_cfg_data and saves prole.cfg immediately after get-credentials ## Database Build Screen (GKE) - Registry display shows correct Artifact Registry URL (<region>-docker.pkg.dev/<project>/<namespace>/knoe-db) in green - Build+push: gcloud auth configure-docker, auto-creates AR repository named after SERVICE_NAMESPACE (e.g. knoe-system) if missing, then docker tag + push; falls back to gcr.io if region unavailable - GCP config loaded from conf/prod/gcp.cfg on every screen entry; keys normalised to lowercase so project_id lookup is always consistent ## Config / Namespace persistence - prole_conf.py activate_environment: symlink creation removed; sets CLUSTER_ENV env-var so all subsequent calls resolve correct env directory - knoe/ui/screens/__init__.py: startup config load uses entrypoint_path() instead of hardcoded conf/prole.cfg; seeds SERVICE_NAMESPACE=knoe-system for managed envs so Common Services never defaults to 'default' - cfg.py _save_prole_cfg: saves to env-specific path via entrypoint_path() - etc/prole_cfg.sh: removed all ln -snf symlink creation Co-authored-by: Junie <junie@jetbrains.com> |
||
|
|
eb9430df9d |
fix(gitlab,infra): ARM64 RPi service cluster – GitLab deploy in gitlab ns on gandalf
Namespace & routing - milestones.py: GitOpsMilestone now resolves namespace from gitops.gitlab_namespace (new) → Global.GITLAB_NAMESPACE → 'gitlab' hardcoded; never falls through to gitops.namespace (was 'gitea') - prole.cfg: add gitops.gitlab_namespace=gitlab + GITLAB_NAMESPACE=gitlab - init_gitlab.sh: NAMESPACE defaults to gitlab, NODE_SELECTOR blanked so only gitaly+minio are node-pinned; GITOPS_NAMESPACE fallback removed GitLab on ARM64 RPi (16 KB kernel pages) - init_gitlab.sh: DaemonSet compiles jemalloc-5.3.0 with --with-lg-page=14 (glibc/Ubuntu) on every node; LD_PRELOAD injected per Ruby component - Minio: quay.io 2022 image (ARM64); configure init container replaced with ARM64 alpine that writes credential files; MINIO_ROOT_USER/PASSWORD injected directly into main container env via secretKeyRef - Minio buckets auto-created post-deploy (registry, lfs, artifacts, etc.) - webservice/sidekiq: replicaCount=1, reduced memory (1500M/800M), liveness probe initialDelaySeconds=3600 (Rails loads 25-40min on RPi) - allowedHosts set as flat string list (chart 9.x default is list-of-maps which breaks URI initializer in 7_gitlab_http.rb) - gitaly+minio always pinned to gandalf (local PV); other workloads spread Storage - Static PVs created for gitaly (50Gi) + minio (10Gi) on synology d005 - Synology dirs created before PVs; bucket creation idempotent Redis (shared for GitLab KAS) - init_redis.sh: persistence disabled (no dynamic provisioner); Redis used as pub/sub broker only Infrastructure / pi.prole.org - Removed pi.prole.org from [k3s_agents] – dedicated pihole node, OOM - host_vars: k3s_enabled=false, k3s_state=absent (storage preserved) - New playbook: infrastructure/playbooks/disable_pi_k3s.yml (drain + disable) - monitoring.py: node-exporter DaemonSet excludes pi.prole.org - init_monitoring.sh: pi.prole.org excluded from node-exporter affinity - kong-deployment.yaml: affinity rule prevents scheduling on pi (pihole owns 80/443) Co-authored-by: Junie <junie@jetbrains.com> |
||
|
|
761d80486b |
feat: add storage probing and operational service updates
- add reusable storage probing subsystem with discovery, bounded probe execution, IO classification, caching, and topology integration - render per-node storage inventory in Cluster Nodes UI and extend installer test coverage for topology/storage behavior - introduce core service operation modules and align actions, milestones, services, and supporting configs/scripts for repair/update workflows - update CNPG/Supabase/database artifacts, placement and port mapping configs, plus related integration tests Co-authored-by: Junie <junie@jetbrains.com> |