Hit in production 2026-08-13: settings crash-looped with 'password
authentication failed' after a routine --only settings redeploy. The
settings Postgres role was only ever created once, in deploy_postgres()'s
init SQL (fresh-volume-only) — so if CREDS_FILE's SETTINGS_DB_PASS ever
drifted from the role's actual password, redeploying settings had no way
to self-heal.
Mirrors the CREATE-then-fallback pattern other apps' deploy functions use,
but falls back to ALTER instead of swallowing the error, since CREATE
failing because the role already exists doesn't fix a drifted password.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Generate KDS_DB_PASS (matching KITCHEN_DB_PASS's pattern), add to --only
dispatch's _append_secret list
- Create a scoped `kds` Postgres role: SELECT/INSERT/UPDATE/DELETE on
kds_tickets/kds_course_bumps, read-only on resos_bookings/resos_opening_hours,
column-scoped SELECT+UPDATE on just the kds_* columns of kitchen_settings
(not the NewBook/ResOS/Nextcloud/Dext/SambaPOS/Azure/Anthropic credentials
that live in the same table). Runs in two passes since kds_tickets/
kds_course_bumps don't exist until KDS's own migrations create them.
- kds .env now carries DATABASE_URL (scoped `kds` role, runtime) and
MIGRATION_DATABASE_URL (privileged `kitchen` role, migrations only)
- Fix KDS theme_color seed (teal -> orange, matched kitchen's tile before)
See kitchen-port-log.md E17.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
deploy_plant() and its --only case entry were already wired in, but
the separate VALID allowlist string gating --only wasn't updated,
so --only plant failed with "Unknown service 'plant'".
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
LXC 123, no admin-VLAN NIC needed (plant only talks to the shared
MQTT broker over the internal network). Wired into --only dispatch
and the main deploy sequence after deploy_calendar.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
deploy_settings() now copies the shared deploy SSH key (same as
deploy_management()) so settings can SSH into LXC 104 to manage
dynamic-security clients, and writes MQTT_BROKER_HOST/MQTT_ADMIN_USER/
MQTT_ADMIN_PASS into its .env from the credentials file.
Moved deploy_mqtt_broker() before deploy_settings() in the full-install
sequence so a fresh install has the admin credentials already generated
by the time settings needs them — previously only worked via a later
`--only settings` redeploy after the broker existed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Creates a throwaway dynsec client+role scoped to a private topic,
publishes/subscribes to confirm round-trip delivery, checks a bogus
login is rejected, then cleans up. Syntax verified against a real
eclipse-mosquitto:2 + dynamic-security broker locally before writing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
eclipse-mosquitto:2 runs as its own uid/gid 1883 and doesn't chown
bind-mounted volumes itself, so the root-owned host dirs from mkdir -p
left it unable to write its log file — broker was stuck restart-looping.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Mosquitto 2.1.2 config format has no nested-block syntax; plugin
options must be flat plugin_opt_<name> directives, not a
plugin_opts { } block. Broker was crash-looping on every start.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Was seeded as 'Operations' here, drifting from the auth service's
seed which uses 'Hotel' — both ON CONFLICT-update the same column, so
whichever ran last silently changed the portal grouping.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
calendar was provisioned ad hoc via add-app.sh and never got the
standard installer wiring other apps have (deploy_<slug>() + --only
dispatch), so it couldn't be redeployed/rebuilt the normal way.
LXC 128, hvac_db, hvac app — plus generic dual-NIC support in create_lxc()
(admin VLAN bridge/tag/IP asked lazily via whiptail, persisted per-hotel in
the credentials file, since these vary per site). Used now for hvac's future
Modbus/Midea/Daikin direct-LAN drivers; the shared MQTT broker LXC will reuse
the same ensure_admin_vlan_config()/ensure_admin_vlan_ip() helpers.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Runs before deploy_reports so UTILITIES_API_KEY exists when reports'
.env is written on a fresh install.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
add-app.sh had a separate Docker-install block that never got the AppArmor
fix from install-stack.sh's install_docker() (commit 9a45e39) — every app
added individually via add-app.sh since then was exposed to Docker builds
failing with "docker-default profile could not be loaded ... while confined".
Hit this deploying the calendar app to LXC 126 on the dev stack.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Creates wages_db and deploys the wages app on LXC 124 (10.10.10.124).
Wired into --only wages dispatch, _append_secret, and full deploy sequence.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Reports no longer queries forecasting_db directly — it calls the
forecasting public API. Remove GRANT SELECT and FORECAST_DATABASE_URL;
add FORECASTING_URL and FORECASTING_API_KEY (provisioned separately in
the forecasting app's API key manager).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Reports LXC now gets SELECT grants on forecasting_db tables (bookings
stats, net revenue, forecasts, budgets) and the FORECAST_DATABASE_URL
env var so the Directors Forecast section can cross-query the forecasting
DB. Also adds 'edit' capability to the auth seeding block.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously the management .env was missing SETTINGS_SECRET (needed by
the backup service to fetch Nextcloud creds) and used stale variable
names (BACKUP_PG_PASS, BACKUP_REMOTE) that the docker-compose never read.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- deploy_kitchen: 2 GB RAM, 10 GB disk; creates kitchen_db (shared with kds);
FastAPI backend with MSSQL ODBC drivers (~5 min build); seeds 10 caps into auth DB
- deploy_kds: 1 GB RAM; connects to kitchen_db (no separate DB — KDS shares schema);
slim FastAPI build (httpx only, no MSSQL ODBC); seeds 3 caps + Staff role grants
- KITCHEN_DB_PASS added to _append_secret section for --only deploys
- Both added to case dispatch and full deploy sequence
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Creates reports_db, deploys reports app to LXC 122, seeds auth DB with
app entry and capabilities, registers /reports/ location in NPM.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The docker run approach failed (image tag/context issues); use the same
direct-psql pattern as add-app.sh which is proven reliable.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds deploy_room_planner() function for LXC 120 / 10.10.10.120.
Wires it into --only room-planner, postgres init SQL (07-room-planner.sql),
gen_secrets, credentials file template, and the main deploy sequence.
NEWBOOK_LOCATION_ID written to .env with a warning if unset.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The proxy-host lookup grepped for '"id":N,"domain_names":[...]' in a
fixed field order that NPM's JSON doesn't guarantee, so the host was never
found ("proxy host not found"). Replace both the cashup and hk-planner
inline blocks with a shared npm_add_location() helper that parses the
proxy-hosts list with python3 and matches DOMAIN inside domain_names.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
npm_get_token now attempts the sourced NPM_ADMIN_* pair, then every
email/pass pair present in the credentials file, then admin@example.com/
changeme. Makes proxy-host automation resilient to duplicate, reordered,
or placeholder NPM entries regardless of how they got there.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The NPM API was reached via NPM_LAN_IP, which breaks when that value is a
placeholder or unset (and the :-10.10.10.103 fallback was wrong — the
internal IP is .3, not .103). NPM listens on all interfaces, so the host
can always reach it at 10.10.10.3:81 over vmbr1. NPM_LAN_IP now only drives
user-facing messages and the NPM LXC's LAN net0.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
In --only mode the creds file provides OFFICE_IP_CHECK, not OFFICE_IP,
so the bare ${OFFICE_IP} tripped 'set -u'. Use the same tolerant
${OFFICE_IP_CHECK:-${OFFICE_IP:-disabled}} form as cashup/hk-planner.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Save NPM_LAN_IP to creds file and reload it in --only mode
- npm_get_token and deploy_cashup NPM patch fall back to 10.10.10.103
(internal vmbr1 IP) when NPM_LAN_IP is unset
- Fix health check URL: /cashup/api/health → /cashup/health
(nginx proxies /cashup/health to backend:3001/health; /cashup/api/
proxies to /api/ which has no /health route)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add manage@hotel.com NPM credentials to creds file template and
reload block (NPM_ADMIN_EMAIL was missing from the --only path)
- Extract npm_get_token() so both configure_npm_proxy_hosts and
deploy_cashup share one auth call
- deploy_cashup now patches the live NPM proxy host to add /cashup/
→ 10.10.10.117:3083 after containers are healthy (idempotent)
- /cashup/ also added to the fresh-install locations array
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Using 'if ! pct exec ... > tmp 2>&1' avoids the bash set -e + $()
interaction where the shell exits inside the subshell before || fires.
Errors are now captured to a temp file and printed via msg_error.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Removed &>/dev/null suppression; output is now captured and shown in
msg_error with a debug command when clone or pull fails.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>