Saturn Fleet Plan
smultar-dev/saturn.os/docs/fleet-plan.md

Saturn Fleet Plan

One flake that hosts, tests and develops your apps across a laptop, a desktop, a bare-metal Incus server and its containers and VMs, an AI box, and two public VPSes. Edit a module once and every machine picks it up. Edit locally when you're offline.

proposal 2026-10-05 branch delta 18 phases ready for handoff
Status · 6 Oct

Work continues on branch delta. Milestones 1–3 are done and deployed. Milestone 4 is done except two items that wait for titan (per-group setup keys and the svc.moons.internal zone): the module merge, default-deny NetBird access as code, the moons.internal switch and the voyager/pioneer renames are live, and all 13 hosts are verified in sync with the repo. Gogs was removed early and backed up. Milestone 5 is blocked on your decisions (unmanaged Incus leftovers such as aether.link/production, the OpenTofu state backend, when to drop deploy-rs); a draft plan is in work.md. The live checklist is in docs/execution-plan.md; the host catalog is docs/hosts.md.

01 · change flow

How a change reaches every machine

Push once. Every host pulls the result itself, and offline edits still work.

PUSH

Edit a module

Commit on saturn (or anywhere) and push to main.

BUILD

CI on titan

The runner builds every host on saturn and pushes signed paths to Attic.

DEPLOY

Panix, canary first

CI runs panix deploy: one moon first, then the rest of the fleet. Hosts copy only the paths they're missing.

CATCH UP

Hourly re-run

Hosts that were offline, like the laptop, get redeployed when they're reachable again.

OFFLINE

Local switch

Git-sync keeps the repo current; run nh os switch .

json@saturn ~/repositories/saturn.os
# normal path: push; CI on titan builds and runs Panix
$ git push origin main
# manual, urgent or first install
$ panix deploy -t enceladus
# NetBird / firewall / sshd changes: test mode first (a reboot reverts)
$ panix deploy -t mimas --activation-mode test
# offline on the laptop
$ nh os switch .
Why Panix only

Panix and comin both deploy the whole system, containers included. Running both is redundant, so Panix pushes everything from saturn and comin is dropped. Builds always happen on saturn, so even pioneer with 1 GiB of RAM only receives finished results.

Risky changes

When a commit touches NetBird, the firewall, sshd or access, CI deploys the canary in test mode (a reboot reverts it) and waits for your approval before switching the rest.

GitHubapp + bot repos, releases
titanrunner · Attic · Infisical · Grafana · Loki
saturnIncus host · builder
enceladusForgejo · restic
iapetuspersonal lab
epimetheus · mimas · dione · telestodev nodes
pandoraAetherLink apps
calypsoeverything Blender
hyperionAI node
voyager · pioneerNetBird control plane · proxy
janus · rheaworkstations + git-sync
NetBird mesh · 10.66.0.0/16 · <host>.moons.internal · control plane on voyager
02 · decisions

Decisions

Settled in the Q&A.

AreaDecision
tailscaleRemove everywhere: the module, every networking.ts line, tailscale0 rules, the ublockdns and netbird-server hooks, and janus + hermes.
netbirdOne directory, networking/netbird/{client,server,proxy}.nix, with options saturn.netbird.{client,server,proxy}.
deploysOne inventory file generates panix.yml and .sops.yaml. Panix only: CI runs it after every merge, plus an hourly catch-up. deploy-rs and comin are dropped.
buildssaturn is the builder, Attic on titan is the cache, and require-sigs goes back on across the fleet.
testingCI builds every host on every push. Apps get Incus preview containers.
appsDocker stays, declared with oci-containers. Every compose stack gets converted now. No k3s.
app cdApp and bot repos export nixosModules and become flake inputs: release → lock-bump PR → merge → CI runs Panix.
secretsInfisical for anything you edit; each service's .env is written to tmpfs when it starts. sops-nix only holds bootstrap secrets. Vault is retired.
gitForgejo replaces Gogs on enceladus (was hestia).
direnvnix-direnv on workstations and dev servers, plus an .envrc in saturn.os.
opsMonitoring, Discord alerts, restic backups to B2, weekly auto-update PRs and central logs, all on titan.
vmsIncus VMs join the fleet with a VM image and a vm-guest role.
03 · hosts

Hosts

Grouped by work group, which matches the Incus projects on saturn. Hostnames checked against the repo. Full catalog: docs/hosts.md.

Naming scheme

Moons run workloads; every new host or sandbox takes the next free moon name, and old names are never reused. Spacecraft (voyager, pioneer; cassini, huygens, dragonfly in reserve) are network infrastructure. saturn is the planet everything orbits. Mesh names become <host>.moons.internal; the public names (netbird, relay, proxy.smultar.com) don't change.

Rollout

Each host is renamed in the phase that first touches it: hostname, NetBird peer, Incus instance (a short stop), sops alias, inventory. The DNS domain switch is fleet-wide and happens once, in phase 2.

Platform

bare metal

The planet: the Incus host everything orbits, and the fleet builder.

HostStatusTodayTarget
saturnThe planet itself, with hundreds of moons in orbit.changeIncus host, deployer, ublockdns, zeronIncus host and remote builder, Panix (also run from titan's CI). ublockdns forwards moons.internal to NetBird DNS.

Ops & data

incus: core

Shared services every other group depends on.

HostStatusTodayTarget
titanSaturn's largest moon, and the only one wrapped in a thick atmosphere.new—LXC ops guest: Attic, GitHub runner, Infisical, Prometheus, Grafana, Loki, Alertmanager → Discord.
enceladus was hestiaAn icy shell over a hidden global ocean.renameGogs, Vault, DockerForgejo and restic. Vault goes once the Infisical migration is done.

Development

incus: core

Persistent remote dev boxes on roles.dev.

HostStatusTodayTarget
epimetheus was prometheusShares an orbit with Janus; the two swap places every four years.renameEx-NetBird head, idleDevelopment node. The rename stops it clashing with the Prometheus monitoring tool.
mimas · dioneMimas wears the Death Star crater; Dione streaks bright ice cliffs.changeDuplicated dev configsroles.dev replaces the 8 lines each host repeats.
telesto was hyperionRides ahead of Tethys at a Lagrange point.renameDev boxSame role; frees hyperion for the AI box.

AetherLink

incus: aether.link

AetherLink distribution: databases and services behind cloudflared.

HostStatusTodayTarget
pandora was al.canary / aetherlinkA shepherd moon that keeps the F ring in line.renamecloudflared + hand-run composeroles.app-host: stacks move into the flake, creds into sops, env into Infisical.

Blender

incus: blender.org

Discord bots and GitHub runners.

HostStatusTodayTarget
calypso was blenderTrails Tethys at a Lagrange point: the opposite of Telesto.renameDocker, Mongo commented outroles.app-host with its stacks converted. The node for every Blender-related task and service, including its Discord bots (as flake-input modules). Bots for other projects live on that project's host.

Personal lab

incus: json

Private. Only the rename.

HostStatusTodayTarget
iapetus was lapetusTwo-faced: one side coal-dark, the other bright as snow.renamePrivate nodeSpelling fix (the moon is Iapetus). Otherwise untouched.

AI

bare metal

Local inference over the mesh.

HostStatusTodayTarget
hyperion was hermesA sponge-like moon that tumbles chaotically.renameOwn nixpkgs pin, Tailscale, upstream llama.cpp (low → extra-high tiers), zeron with tool pins offroles.ai-node on the main channel (tool pins back on), with the LiteLLM keys in Infisical.

Network

public VPS · spacecraft

NetBird infrastructure, named after spacecraft so it's never mistaken for a moon.

HostStatusTodayTarget
voyager was atlasStill sending data home from interstellar space.renameNetBird control plane on :latestPinned images, trustedPeers fix, backups of the store and secrets.
pioneer was daphnisThe first spacecraft ever to reach Saturn (1979).renameNetBird proxy on :latestPinned to the same version as voyager. Never builds locally (1 GiB).

Workstations

local

Pushed by Panix (tag workstation, hourly catch-up), or by hand with nh os switch . when offline.

HostStatusTodayTarget
janusShares an orbit with Epimetheus; the two swap places every four years.changeLaptop, Tailscale, hand-placed keysops recipient, git-sync, nix-direnv, Attic, joins NetBird grace-access for work. Optional btrfs reinstall later.
rheaSaturn's second-largest moon.changeWindows on the MP700 (NetBird peer rhea-pc), data on the SN850NixOS on the MP700 (btrfs), SN850 as a btrfs data disk, bulk/games disk to come. Windows moves to a portable USB. NVIDIA, Ollama on CUDA.

Images & archive

not deployed

Build outputs and retired hosts.

HostStatusTodayTarget
container-imagechangeLXC imageMoves to images/lxc, next to a new images/vm.
jimothy · ruru · tethys · homelab · al.prodarchiveDeprovisionedMove to attic/ along with templates/.
04 · netbird

NetBird

How the mesh works today, and what changes in phase 2.

PartTodayTarget
control planeatlas: server, dashboard and relay containers on :latest behind Caddy (Let's Encrypt), built-in login, SQLite store in a Docker volume, secrets generated on first bootvoyager: pinned images, trustedPeers fixed, store and secrets backed up, secrets in sops
clientEvery host: custom 0.77.0 binary, one shared setup key, migration units, Nix trust settings and a CA bundled into the same modulesaturn.netbird.client: nixpkgs package if it's ≥ 0.77, one setup key per group, Nix trust and CA moved out
proxydaphnis: reverse-proxy on :latest, token in a file placed by handpioneer: pinned to voyager's version, token in sops
mesh dns<host>.saturn.moons, so saturn.saturn.moons; peer names out of sync with hostnames<host>.moons.internal; every peer renamed to its hostname
access rulesDashboard, plus direct sqlite edits from a scriptGroups, rules and setup keys in the repo, applied by OpenTofu (NetBird provider)
Fix first · trustedPeers

Our config.yaml trusts 127.0.0.1 and has no trustedPeers. Caddy actually connects from the Docker bridge, so up to v0.79 client IPs are recorded wrong. After v0.79 the server trusts forwarded client-IP headers from any source. Fix: give the compose network a fixed subnet and trust only its gateway, in both settings and in the relay.

Upgrade procedure

Back up, read the release notes, then bump server, dashboard, relay and proxy in one commit. Deploy voyager first and pioneer second. The proxy must never run newer than management.

05 · modules

Modules

Every file under modules/ is imported on every host and does nothing until its enable flag is set.

Removed

4
  • networking/tailscale
  • services/crowdsec: dead file with placeholder keys
  • services/gogs: once Forgejo is running
  • services/vault and vault-agent: once Infisical is running

Reshaped

4
  • The 3 NetBird modules merge into networking/netbird/
  • profiles/*, incus-vm and hardware/vm become roles/
  • disko gets selectable layouts instead of mkForce false
  • deployer loses deploy-rs and Vault and gains the Infisical CLI

New

added
  • core/{git-sync, pki}
  • Platform: services/attic, services/github-runner
  • Data: services/infisical, services/forgejo, services/backup
  • Observability: services/monitoring, services/logging
  • stacks/<svc>, one module for each converted stack
  • fleet/inventory.nix, the single list of hosts

Kept

cleaned
  • base, access, hardening, nh, build-vm, tool-pins
  • ublockdns, docker, incus, ollama, llama-cpp-server, litellm-proxy
  • desktop/hyprland, home, apps/*, hardware/*
RoleBundlesUsed by
serverSSH on :7700, access, hardening, firewallevery non-workstation host
workstationHyprland, Home Manager, nix-direnv, git-syncjanus, rhea
devdocker, nh, direnv, tool pins, zeron, remote editorsepimetheus, mimas, dione, telesto
app-hostdocker, stacks, Infisical env, cloudflared, resticpandora, calypso
ai-nodellama.cpp, LiteLLM, T3, opencode confighyperion
lxc-guest · vm-guestIncus container or VM plumbingall Incus guests
public-vpsstatic IP, qemu-guestvoyager, pioneer
06 · compose stacks

Your compose stacks

The /containers/<svc>/{docker-compose.yml,.env,data} layout carries over to oci-containers.

Composeoci-containers
./data:/x"/containers/<svc>/data:/x". Your data stays where it is.
env_file: .envenvironmentFiles pointing at /run/secrets/<svc>.env, written from Infisical when the service starts
networks:compose2nix creates a docker-network-<svc> unit
depends_onsystemd after and requires
build:Not supported. CI builds the image and pushes it to GHCR.
image: x:latestPin the tag and digest so a rebuild never silently pulls a new image
07 · phases

Phases

Each phase is one PR, and every host has to build in CI before it merges.

#PhaseContentsRisk
0HygieneMove dead hosts to attic/, delete crowdsec, add .envrc, set stateVersion per hostlow
1Remove TailscaleEvery reference; janus and hermes get test mode firstmedium
2NetBirdDefault-deny policies as code, service DNS zone. Back up and fix trustedPeers first, then: one module directory, pinned images, atlas→voyager and daphnis→pioneer, the moons.internal domain switch, every peer renamed to its hostnamemedium
3Inventory + rolesGenerate panix.yml, .sops.yaml and the OpenTofu config; import existing Incus resources; roles replace duplicated config; deploy-rs droppedlow
4titan, cache, CIAttic signing, saturn as builder, the runner, CI builds; require-sigs back onmedium
5CI deploys + git-syncCI runs Panix after merge (canary, then the fleet), approval gate for risky paths, hourly catch-up, git-sync on workstationsmedium
6InfisicalMachine identities via sops, Vault KV migrated, Vault retiredmedium
7Forgejohestia→enceladus, repos migrated from Gogs, GitHub mirrors, gogs deletedlow
8App hostsal.canary→pandora and blender→calypso, every stack converted, Infisical env, app template, first Discord botmedium
9Dev nodesprometheus→epimetheus, hyperion→telesto, roles.dev, direnv, remote editorslow
10AI nodehermes→hyperion (once telesto frees the name), build on unstable, test mode, compare llama.cpp performance, drop the pinsmedium
11Observability + backupsExporters, Grafana, Loki and Discord alerts; restic to B2 with a monthly restore testlow
12Auto-updatesA weekly flake.lock PR; CI, merge and Panix take it from therelow
13Incus VMsVM image, vm-guest role, Incus VM profilelow
14rhea installBack up C: and D:, build the Windows USB, disko wipes the MP700 (btrfs), install, reformat the SN850 as btrfs and restore the data, NVIDIAmedium
15janus to btrfsOptional: back up /home, reinstall with rhea's btrfs layoutmedium
16Incus isolationGuests stop trusting eth0; OpenTofu creates a bridge per project and ACLs between guestsmedium
17AI sandboxesfleet MCP server, roles.sandbox, ephemeral keys, reaper, promotemedium
08 · review

Review findings

Problems found in the current config. Each one is fixed in a phase above.

  • The server profile forces stateVersion = "25.05" on every server, so hermes has to mkForce its own.
  • LXC guests run NetworkManager and networkd at the same time.
  • The LXC profile is misnamed and trusts a stale /home/json/rhea.os path.
  • The VM module hardcodes kvm-amd, but saturn is Intel.
  • Docker puts 127.0.0.1 first in container DNS, which breaks lookups inside containers.
  • The Incus API listens on 0.0.0.0:8443.
  • require-sigs = false is set on most hosts, in three separate places.
  • The NetBird images run :latest, and the LiteLLM master key is hardcoded.
  • Two node lists are kept in sync by hand, and secretspec.toml is still named rhea.os.
09 · blocking

Open questions

Later phases need these answered before they can start.

  1. Compose files: copy each host's /containers/*/docker-compose.yml into the repo, without the .env files or data.
  2. Alerts: which Discord webhook should alerts go to?
Settled

calypso is the node for everything Blender, bots included; other bots live with their project. iapetus is in the json project, rename only. janus and hyperion join NetBird grace-access for work. Infisical is mesh-only at infisical.moons.internal.

10 · rhea

rhea disk layout

Windows leaves the internal disks for a portable USB drive, so NixOS owns every internal disk. No encryption.

DiskTodayTargetFilesystem
MP700 1 TBWindows C:NixOS system: ESP + subvolumes @ @home @nix @log @snapshots, zram for swapbtrfs
WD SN850 1 TBWindows D: (data)Secondary data disk: @data, @games, mounted nofailbtrfs
third diskcoming soonBulk / games at /mnt/bulkbtrfs
USB drive—Portable Windows, picked from the firmware boot menuNTFS
Why btrfs (janus is ext4 today)

ext4 is simple and fast but has no snapshots or compression. btrfs adds hourly /home snapshots for one-command restores, zstd compression (often 20–40% smaller for code, the Nix store and models), subvolumes that share free space, checksums against silent corruption, and btrfs send backups to the bulk disk. Windows can't read it, which doesn't matter once Windows lives on its own USB.

Before installing

Back up everything on C: and D: you want to keep, and build and test the Windows USB first. The MP700 is wiped. Use nodatacow on any VM-image or database folders.

11 · access

NetBird access & groups

Default deny: the built-in All → All policy is deleted, and every flow is an explicit policy in fleet/netbird.nix, applied by OpenTofu after each merge.

GroupMembers
adminjanus, rhea, iris (phone), saturn: full access
g-platformsaturn
g-opstitan, enceladus
g-devepimetheus, mimas, dione, telesto
g-aetherlinkpandora
g-blendercalypso
g-labiapetus
g-aihyperion
g-networkvoyager, pioneer
ephemeralAI sandboxes
grace-accessjanus, hyperion, grace devices
FromToPortsWhy
adminallallYour devices
titanall servers7700CI deploys with Panix
all serverstitan443, LokiCache, secrets, logs
titanall servers9100Metrics
g-dev, ephemeralhyperion8080, 4000LLM API (opt-in)
ephemeraltitan443Attic cache
grace-accessgrace networkexistingWork
Service names

<svc>.svc.moons.internal comes from a DNS zone on titan, generated from the inventory's services list. Caddy on each host routes it to the local port. The same entry generates the access policy. Services that need separate access get separate ports.

Gap until phase 16

Guests also share saturn's Incus bridge and trust eth0, so they can reach each other there regardless of NetBird. Phase 16 closes this with guest firewalls, Incus ACLs and a bridge per project. Also: a lost phone in admin is a full-access device, so remove it from the dashboard right away.

12 · sandboxes

AI sandboxes

Tell an AI what machine you want; it builds a NixOS container that joins the mesh by itself.

ASK

Claude Code

The AI calls the fleet MCP server on titan with a spec: packages, services, ports, size, optional group.

BUILD

Nix on saturn

Spec + roles.sandbox + the LXC image are built into a NixOS system.

LAUNCH

Incus

An unprivileged container in the sandbox project, with its own bridge and limits.

ENROLL

Ephemeral key

A one-use, 1-hour NetBird key puts it in ephemeral plus the requested group, under the next free moon name, e.g. kiviuq.moons.internal.

EXPIRE

24 h reaper

Destroyed unless extended. NetBird drops the peer once it's offline. fleet promote turns it into a real host through a PR.

Defaults

24 h, 4 CPU, 8 GiB, 30 GiB, up to 5 at once. Reaches the internet, titan's cache and hyperion's LLM API, nothing else unless its group allows it.

Group on creation

Only groups on an allow list (e.g. g-dev, g-blender). admin, g-network and g-ops can never be requested, so an AI can't create a full-access or infrastructure peer.

13 · infrastructure

OpenTofu via terranix

Incus and NetBird are declared in Nix (terranix), generated from fleet/inventory.nix, and applied by OpenTofu from titan's CI.

LayerTool
incusOpenTofu, lxc/incus provider: projects, profiles, a bridge per project, network ACLs, storage, long-lived guests
netbirdOpenTofu, NetBird provider: groups, policies, setup keys, the svc.moons.internal DNS zone
inside hostsNixOS flake + Panix
ai sandboxesfleet MCP server → Incus API directly (create-on-request doesn't suit Terraform state); OpenTofu defines their project, bridge and limits
Flow

Edit the inventory and push. CI posts tofu plan on the PR. On merge, tofu apply creates or changes the Incus guest first, then Panix deploys its NixOS config.

State

The pg backend in Postgres on enceladus, with OpenTofu state encryption (key in sops), backed up by restic. Existing projects and guests are imported once in phase 3, then the bootstrap script and the preseed are retired.

14 · progress

Master checklist

71 items, mirrored from docs/execution-plan.md. The agent ticks items there with each commit and updates this section after every milestone.

todoin progressdoneblocked

Milestone 1: Phase 0, hygiene

  • 0.1 Stop copying the flake into /etc (server profile, container-image)
  • 0.2 scripts/eval-all.sh + record baseline drvPaths for every host
  • 0.3 Archive jimothy, ruru, tethys, homelab, al.prod, templates/ to attic/
  • 0.4 Delete modules/services/crowdsec/default_.nix
  • 0.5 Per-host system.stateVersion (drvPaths unchanged)
  • 0.6 .envrc + devShell tools (sops, age, ssh-to-age, nixfmt-tree)
  • 0.7 Fix .cursor/rules/nix.mdc and notes/core-incus.md
  • Owner: deploy 0.1 to the fleet (verified: /etc/nixos-config gone on all 8 guests; hestia first rolled back to gen 19 by hand)

Milestone 2: Phase 1, remove Tailscale

  • 1.1 Delete modules/networking/tailscale
  • 1.2 Remove every saturn.networking.ts.enable line
  • 1.3 hermes: drop services.tailscale.extraUpFlags
  • 1.4 janus: firewall, ports and comments
  • 1.5 Hooks in ublockdns, netbird-server, netbird-proxy, incus
  • 1.6 Notes cleaned
  • Verify: grep -ri tailscale hosts modules is empty; nix-diff shows only Tailscale changes (0 refs; nix-diff: firewall only)
  • Owner: tailscale logout on janus + hermes; hermes test-mode deploy; janus nh os switch; rest of fleet

Milestone 3: Phase 2a, NetBird safety on atlas

  • Ask owner: running image versions (atlas, daphnis) (read from the live hosts by the agent)
  • 2a.1 Pin server, dashboard, relay and proxy images
  • 2a.2 Fixed compose subnet; trustedHTTPProxies + trustedPeers + relay trusted proxies
  • 2a.3 NetBird store and secrets backup timer
  • Owner: manual backup (agent ran it, copies in ~/backups/atlas on saturn), test-mode deploy on atlas, check logs, real deploy

Milestone 4: Phases 2b–2e, NetBird (plan mode first)

  • Ask owner: NetBird API token, live peer and group list (token in ~/.config/saturn/netbird-token; live lists read from the API)
  • Merge into modules/networking/netbird/{client,server,proxy}.nix, options saturn.netbird.*
  • Move Nix trust and the hestia CA out of the client module; drop old migration units
  • Owner: deploy stage 4a (all 14 hosts verified live on the new build; hestia needed two reboots, see 854a61d)
  • fleet/netbird.nix groups, policies, services, setup keys (OpenTofu NetBird provider). Groups and policies applied and verified; services and per-group setup keys still to do
  • Delete the All → All policy (default deny), add the required fleet flows (Default was already disabled; saturn-moons-mesh removed, admin and dev-to-ai flows live)
  • svc.moons.internal zone (CoreDNS on titan, later) + Caddy routing
  • Mesh domain saturn.moons → moons.internal + every hard-coded reference (incl. hermes /etc/hosts pin) (f34a18d; live switch applied via tofu, verified)
  • Rename atlas → voyager, daphnis → pioneer; rename every peer to its hostname (d6c68d7, 9809909; the other peers already carried their hostnames)
  • Remove the unused saturn-moons group
  • saturn resolves mesh names: static pins from fleet/netbird.nix, since the uBlockDNS client cannot forward a domain (b6461be, deployed; all 13 hosts verified in sync)
  • Per-group setup keys (janus is already a sops recipient; rhea waits for its NixOS install in Milestone 11)

Milestone 5: Phase 3, inventory, roles, OpenTofu

  • Ask owner: unmanaged Incus leftovers stay out of OpenTofu, aether.link/production stays unmanaged, local state until Milestone 8, deploy-rs kept until Milestone 7 (plan approved)
  • fleet/inventory.nix (hosts, groups, tags, roles, descriptions, services); Incus placement and sops keys included, roles filled in step 5.3
  • fleet/moon-names.nix + fleet name + CI uniqueness check (nix run .#fleet -- name; check fleet-names)
  • Roles: server, lxc-guest, dev, public-vps, ai-node, workstation (all 15 drvPaths identical except prometheus: a duplicate trusted-users entry removed). app-host, vm-guest and sandbox wait: nothing uses them yet (5f9c531..66fa91f)
  • Generate panix.yml, .sops.yaml and deploy.nodes from the inventory (nix run .#fleet-config, check fleet-config); deploy-rs kept until Milestone 7 by decision
  • terranix + OpenTofu (Incus + NetBird providers): fleet/incus.nix, scripts/incus-tofu; local state by decision, moves to Postgres in Milestone 8
  • tofu import existing Incus projects, profiles and guests; bootstrap script and preseed retired (ff46bad, bbba186; saturn deployed, 11 guests untouched)
  • Rename lapetus → iapetus

Milestone 6: Phase 4, titan, cache, CI

  • Create titan (via OpenTofu)
  • Attic + signing key; saturn as remote builder
  • GitHub runner; CI builds every host (checks)
  • require-sigs = true fleet-wide; remove unsigned trust

Milestone 7: Phase 5, CI deploys

  • CI runs Panix after merge: canary → fleet
  • Approval gate for networking, access and hardening paths
  • Deploy-guard watchdog (confirm or revert)
  • Hourly catch-up via system.configurationRevision
  • git-sync on workstations

Milestone 8: Phases 6–7, Infisical + Forgejo

  • Infisical on titan (mesh-only), machine identities via sops
  • Migrate Vault KV → Infisical; switch LiteLLM, cloudflared, etc.; remove Vault and vault-agent
  • Rename hestia → enceladus
  • Gogs removed early (stale data; module, hestia config, secrets contract and notes deleted; 120 MB backup in ~/backups/hestia-gogs on saturn)
  • Forgejo on enceladus (nothing to migrate: the Gogs repos were stale); GitHub mirrors if wanted

Milestone 9: Phases 8–10, apps, dev, AI

  • Ask owner: /containers/*/docker-compose.yml from each app host
  • Rename al.canary → pandora, blender → calypso
  • Convert every compose stack to modules/stacks/<svc> (oci-containers, pinned digests, Infisical env)
  • App template (nix flake init -t …#app) + flake-input CD; first Discord bot on calypso
  • Rename prometheus → epimetheus, hyperion → telesto; roles.dev, nix-direnv, remote editors
  • Rename hermes → hyperion; move to unstable; tool pins back on; drop the hermes / stable pins. Also fixes litellm-proxy, which crash-loops on the pinned nixpkgs (litellm 1.97.0 vs fastapi 0.141.1); unstable has 1.102.1, verified to start

Milestone 10: Phases 11–13, 16–17

  • Ask owner: Discord webhook
  • Prometheus, Grafana, Loki, Alertmanager → Discord; status page (docs/status-mock.html design)
  • restic → B2 + monthly restore test
  • Weekly flake.lock update PR
  • Incus VM image + roles.vm-guest
  • Incus isolation: guests stop trusting eth0, bridge per project, ACLs (OpenTofu)
  • AI sandboxes: fleet MCP server, roles.sandbox, ephemeral keys, reaper, promote

Milestone 11: Phases 14–15, workstations

  • Ask owner: backups of C: and D: done, Windows USB tested
  • rhea: disko btrfs on MP700, SN850 btrfs data, NVIDIA, install
  • janus btrfs reinstall (optional)