Saturn Fleet Plan
One flake that hosts, tests and develops your apps across a laptop, a desktop, a bare-metal Incus server and its containers and VMs, an AI box, and two public VPSes. Edit a module once and every machine picks it up. Edit locally when you're offline.
Work continues on branch delta. Milestones 1–3 are done and deployed. Milestone 4 is done except two items that wait for titan (per-group setup keys and the svc.moons.internal zone): the module merge, default-deny NetBird access as code, the moons.internal switch and the voyager/pioneer renames are live, and all 13 hosts are verified in sync with the repo. Gogs was removed early and backed up. Milestone 5 is blocked on your decisions (unmanaged Incus leftovers such as aether.link/production, the OpenTofu state backend, when to drop deploy-rs); a draft plan is in work.md. The live checklist is in docs/execution-plan.md; the host catalog is docs/hosts.md.
How a change reaches every machine
Push once. Every host pulls the result itself, and offline edits still work.
Edit a module
Commit on saturn (or anywhere) and push to main.
CI on titan
The runner builds every host on saturn and pushes signed paths to Attic.
Panix, canary first
CI runs panix deploy: one moon first, then the rest of the fleet. Hosts copy only the paths they're missing.
Hourly re-run
Hosts that were offline, like the laptop, get redeployed when they're reachable again.
Local switch
Git-sync keeps the repo current; run nh os switch .
# normal path: push; CI on titan builds and runs Panix $ git push origin main # manual, urgent or first install $ panix deploy -t enceladus # NetBird / firewall / sshd changes: test mode first (a reboot reverts) $ panix deploy -t mimas --activation-mode test # offline on the laptop $ nh os switch .
Panix and comin both deploy the whole system, containers included. Running both is redundant, so Panix pushes everything from saturn and comin is dropped. Builds always happen on saturn, so even pioneer with 1 GiB of RAM only receives finished results.
When a commit touches NetBird, the firewall, sshd or access, CI deploys the canary in test mode (a reboot reverts it) and waits for your approval before switching the rest.
Decisions
Settled in the Q&A.
| Area | Decision |
|---|---|
| tailscale | Remove everywhere: the module, every networking.ts line, tailscale0 rules, the ublockdns and netbird-server hooks, and janus + hermes. |
| netbird | One directory, networking/netbird/{client,server,proxy}.nix, with options saturn.netbird.{client,server,proxy}. |
| deploys | One inventory file generates panix.yml and .sops.yaml. Panix only: CI runs it after every merge, plus an hourly catch-up. deploy-rs and comin are dropped. |
| builds | saturn is the builder, Attic on titan is the cache, and require-sigs goes back on across the fleet. |
| testing | CI builds every host on every push. Apps get Incus preview containers. |
| apps | Docker stays, declared with oci-containers. Every compose stack gets converted now. No k3s. |
| app cd | App and bot repos export nixosModules and become flake inputs: release → lock-bump PR → merge → CI runs Panix. |
| secrets | Infisical for anything you edit; each service's .env is written to tmpfs when it starts. sops-nix only holds bootstrap secrets. Vault is retired. |
| git | Forgejo replaces Gogs on enceladus (was hestia). |
| direnv | nix-direnv on workstations and dev servers, plus an .envrc in saturn.os. |
| ops | Monitoring, Discord alerts, restic backups to B2, weekly auto-update PRs and central logs, all on titan. |
| vms | Incus VMs join the fleet with a VM image and a vm-guest role. |
Hosts
Grouped by work group, which matches the Incus projects on saturn. Hostnames checked against the repo. Full catalog: docs/hosts.md.
Moons run workloads; every new host or sandbox takes the next free moon name, and old names are never reused. Spacecraft (voyager, pioneer; cassini, huygens, dragonfly in reserve) are network infrastructure. saturn is the planet everything orbits. Mesh names become <host>.moons.internal; the public names (netbird, relay, proxy.smultar.com) don't change.
Each host is renamed in the phase that first touches it: hostname, NetBird peer, Incus instance (a short stop), sops alias, inventory. The DNS domain switch is fleet-wide and happens once, in phase 2.
Platform
bare metalThe planet: the Incus host everything orbits, and the fleet builder.
| Host | Status | Today | Target |
|---|---|---|---|
| saturnThe planet itself, with hundreds of moons in orbit. | change | Incus host, deployer, ublockdns, zeron | Incus host and remote builder, Panix (also run from titan's CI). ublockdns forwards moons.internal to NetBird DNS. |
Ops & data
incus: coreShared services every other group depends on.
| Host | Status | Today | Target |
|---|---|---|---|
| titanSaturn's largest moon, and the only one wrapped in a thick atmosphere. | new | — | LXC ops guest: Attic, GitHub runner, Infisical, Prometheus, Grafana, Loki, Alertmanager → Discord. |
| enceladus was hestiaAn icy shell over a hidden global ocean. | rename | Gogs, Vault, Docker | Forgejo and restic. Vault goes once the Infisical migration is done. |
Development
incus: corePersistent remote dev boxes on roles.dev.
| Host | Status | Today | Target |
|---|---|---|---|
| epimetheus was prometheusShares an orbit with Janus; the two swap places every four years. | rename | Ex-NetBird head, idle | Development node. The rename stops it clashing with the Prometheus monitoring tool. |
| mimas · dioneMimas wears the Death Star crater; Dione streaks bright ice cliffs. | change | Duplicated dev configs | roles.dev replaces the 8 lines each host repeats. |
| telesto was hyperionRides ahead of Tethys at a Lagrange point. | rename | Dev box | Same role; frees hyperion for the AI box. |
AetherLink
incus: aether.linkAetherLink distribution: databases and services behind cloudflared.
| Host | Status | Today | Target |
|---|---|---|---|
| pandora was al.canary / aetherlinkA shepherd moon that keeps the F ring in line. | rename | cloudflared + hand-run compose | roles.app-host: stacks move into the flake, creds into sops, env into Infisical. |
Blender
incus: blender.orgDiscord bots and GitHub runners.
| Host | Status | Today | Target |
|---|---|---|---|
| calypso was blenderTrails Tethys at a Lagrange point: the opposite of Telesto. | rename | Docker, Mongo commented out | roles.app-host with its stacks converted. The node for every Blender-related task and service, including its Discord bots (as flake-input modules). Bots for other projects live on that project's host. |
Personal lab
incus: jsonPrivate. Only the rename.
| Host | Status | Today | Target |
|---|---|---|---|
| iapetus was lapetusTwo-faced: one side coal-dark, the other bright as snow. | rename | Private node | Spelling fix (the moon is Iapetus). Otherwise untouched. |
AI
bare metalLocal inference over the mesh.
| Host | Status | Today | Target |
|---|---|---|---|
| hyperion was hermesA sponge-like moon that tumbles chaotically. | rename | Own nixpkgs pin, Tailscale, upstream llama.cpp (low → extra-high tiers), zeron with tool pins off | roles.ai-node on the main channel (tool pins back on), with the LiteLLM keys in Infisical. |
Network
public VPS · spacecraftNetBird infrastructure, named after spacecraft so it's never mistaken for a moon.
| Host | Status | Today | Target |
|---|---|---|---|
| voyager was atlasStill sending data home from interstellar space. | rename | NetBird control plane on :latest | Pinned images, trustedPeers fix, backups of the store and secrets. |
| pioneer was daphnisThe first spacecraft ever to reach Saturn (1979). | rename | NetBird proxy on :latest | Pinned to the same version as voyager. Never builds locally (1 GiB). |
Workstations
localPushed by Panix (tag workstation, hourly catch-up), or by hand with nh os switch . when offline.
| Host | Status | Today | Target |
|---|---|---|---|
| janusShares an orbit with Epimetheus; the two swap places every four years. | change | Laptop, Tailscale, hand-placed key | sops recipient, git-sync, nix-direnv, Attic, joins NetBird grace-access for work. Optional btrfs reinstall later. |
| rheaSaturn's second-largest moon. | change | Windows on the MP700 (NetBird peer rhea-pc), data on the SN850 | NixOS on the MP700 (btrfs), SN850 as a btrfs data disk, bulk/games disk to come. Windows moves to a portable USB. NVIDIA, Ollama on CUDA. |
Images & archive
not deployedBuild outputs and retired hosts.
| Host | Status | Today | Target |
|---|---|---|---|
| container-image | change | LXC image | Moves to images/lxc, next to a new images/vm. |
| jimothy · ruru · tethys · homelab · al.prod | archive | Deprovisioned | Move to attic/ along with templates/. |
NetBird
How the mesh works today, and what changes in phase 2.
| Part | Today | Target |
|---|---|---|
| control plane | atlas: server, dashboard and relay containers on :latest behind Caddy (Let's Encrypt), built-in login, SQLite store in a Docker volume, secrets generated on first boot | voyager: pinned images, trustedPeers fixed, store and secrets backed up, secrets in sops |
| client | Every host: custom 0.77.0 binary, one shared setup key, migration units, Nix trust settings and a CA bundled into the same module | saturn.netbird.client: nixpkgs package if it's ≥ 0.77, one setup key per group, Nix trust and CA moved out |
| proxy | daphnis: reverse-proxy on :latest, token in a file placed by hand | pioneer: pinned to voyager's version, token in sops |
| mesh dns | <host>.saturn.moons, so saturn.saturn.moons; peer names out of sync with hostnames | <host>.moons.internal; every peer renamed to its hostname |
| access rules | Dashboard, plus direct sqlite edits from a script | Groups, rules and setup keys in the repo, applied by OpenTofu (NetBird provider) |
Our config.yaml trusts 127.0.0.1 and has no trustedPeers. Caddy actually connects from the Docker bridge, so up to v0.79 client IPs are recorded wrong. After v0.79 the server trusts forwarded client-IP headers from any source. Fix: give the compose network a fixed subnet and trust only its gateway, in both settings and in the relay.
Back up, read the release notes, then bump server, dashboard, relay and proxy in one commit. Deploy voyager first and pioneer second. The proxy must never run newer than management.
Modules
Every file under modules/ is imported on every host and does nothing until its enable flag is set.
Removed
4networking/tailscaleservices/crowdsec: dead file with placeholder keysservices/gogs: once Forgejo is runningservices/vaultandvault-agent: once Infisical is running
Reshaped
4- The 3 NetBird modules merge into
networking/netbird/ profiles/*,incus-vmandhardware/vmbecomeroles/diskogets selectable layouts instead ofmkForce falsedeployerloses deploy-rs and Vault and gains the Infisical CLI
New
addedcore/{git-sync, pki}- Platform:
services/attic,services/github-runner - Data:
services/infisical,services/forgejo,services/backup - Observability:
services/monitoring,services/logging stacks/<svc>, one module for each converted stackfleet/inventory.nix, the single list of hosts
Kept
cleaned- base, access, hardening, nh, build-vm, tool-pins
- ublockdns, docker, incus, ollama, llama-cpp-server, litellm-proxy
- desktop/hyprland, home, apps/*, hardware/*
| Role | Bundles | Used by |
|---|---|---|
| server | SSH on :7700, access, hardening, firewall | every non-workstation host |
| workstation | Hyprland, Home Manager, nix-direnv, git-sync | janus, rhea |
| dev | docker, nh, direnv, tool pins, zeron, remote editors | epimetheus, mimas, dione, telesto |
| app-host | docker, stacks, Infisical env, cloudflared, restic | pandora, calypso |
| ai-node | llama.cpp, LiteLLM, T3, opencode config | hyperion |
| lxc-guest · vm-guest | Incus container or VM plumbing | all Incus guests |
| public-vps | static IP, qemu-guest | voyager, pioneer |
Your compose stacks
The /containers/<svc>/{docker-compose.yml,.env,data} layout carries over to oci-containers.
| Compose | oci-containers |
|---|---|
| ./data:/x | "/containers/<svc>/data:/x". Your data stays where it is. |
| env_file: .env | environmentFiles pointing at /run/secrets/<svc>.env, written from Infisical when the service starts |
| networks: | compose2nix creates a docker-network-<svc> unit |
| depends_on | systemd after and requires |
| build: | Not supported. CI builds the image and pushes it to GHCR. |
| image: x:latest | Pin the tag and digest so a rebuild never silently pulls a new image |
Phases
Each phase is one PR, and every host has to build in CI before it merges.
| # | Phase | Contents | Risk |
|---|---|---|---|
| 0 | Hygiene | Move dead hosts to attic/, delete crowdsec, add .envrc, set stateVersion per host | low |
| 1 | Remove Tailscale | Every reference; janus and hermes get test mode first | medium |
| 2 | NetBird | Default-deny policies as code, service DNS zone. Back up and fix trustedPeers first, then: one module directory, pinned images, atlas→voyager and daphnis→pioneer, the moons.internal domain switch, every peer renamed to its hostname | medium |
| 3 | Inventory + roles | Generate panix.yml, .sops.yaml and the OpenTofu config; import existing Incus resources; roles replace duplicated config; deploy-rs dropped | low |
| 4 | titan, cache, CI | Attic signing, saturn as builder, the runner, CI builds; require-sigs back on | medium |
| 5 | CI deploys + git-sync | CI runs Panix after merge (canary, then the fleet), approval gate for risky paths, hourly catch-up, git-sync on workstations | medium |
| 6 | Infisical | Machine identities via sops, Vault KV migrated, Vault retired | medium |
| 7 | Forgejo | hestia→enceladus, repos migrated from Gogs, GitHub mirrors, gogs deleted | low |
| 8 | App hosts | al.canary→pandora and blender→calypso, every stack converted, Infisical env, app template, first Discord bot | medium |
| 9 | Dev nodes | prometheus→epimetheus, hyperion→telesto, roles.dev, direnv, remote editors | low |
| 10 | AI node | hermes→hyperion (once telesto frees the name), build on unstable, test mode, compare llama.cpp performance, drop the pins | medium |
| 11 | Observability + backups | Exporters, Grafana, Loki and Discord alerts; restic to B2 with a monthly restore test | low |
| 12 | Auto-updates | A weekly flake.lock PR; CI, merge and Panix take it from there | low |
| 13 | Incus VMs | VM image, vm-guest role, Incus VM profile | low |
| 14 | rhea install | Back up C: and D:, build the Windows USB, disko wipes the MP700 (btrfs), install, reformat the SN850 as btrfs and restore the data, NVIDIA | medium |
| 15 | janus to btrfs | Optional: back up /home, reinstall with rhea's btrfs layout | medium |
| 16 | Incus isolation | Guests stop trusting eth0; OpenTofu creates a bridge per project and ACLs between guests | medium |
| 17 | AI sandboxes | fleet MCP server, roles.sandbox, ephemeral keys, reaper, promote | medium |
Review findings
Problems found in the current config. Each one is fixed in a phase above.
- The server profile forces
stateVersion = "25.05"on every server, so hermes has tomkForceits own. - LXC guests run NetworkManager and networkd at the same time.
- The LXC profile is misnamed and trusts a stale
/home/json/rhea.ospath. - The VM module hardcodes
kvm-amd, but saturn is Intel. - Docker puts
127.0.0.1first in container DNS, which breaks lookups inside containers. - The Incus API listens on
0.0.0.0:8443. require-sigs = falseis set on most hosts, in three separate places.- The NetBird images run
:latest, and the LiteLLM master key is hardcoded. - Two node lists are kept in sync by hand, and
secretspec.tomlis still namedrhea.os.
Open questions
Later phases need these answered before they can start.
- Compose files: copy each host's
/containers/*/docker-compose.ymlinto the repo, without the .env files or data. - Alerts: which Discord webhook should alerts go to?
calypso is the node for everything Blender, bots included; other bots live with their project. iapetus is in the json project, rename only. janus and hyperion join NetBird grace-access for work. Infisical is mesh-only at infisical.moons.internal.
rhea disk layout
Windows leaves the internal disks for a portable USB drive, so NixOS owns every internal disk. No encryption.
| Disk | Today | Target | Filesystem |
|---|---|---|---|
| MP700 1 TB | Windows C: | NixOS system: ESP + subvolumes @ @home @nix @log @snapshots, zram for swap | btrfs |
| WD SN850 1 TB | Windows D: (data) | Secondary data disk: @data, @games, mounted nofail | btrfs |
| third disk | coming soon | Bulk / games at /mnt/bulk | btrfs |
| USB drive | — | Portable Windows, picked from the firmware boot menu | NTFS |
ext4 is simple and fast but has no snapshots or compression. btrfs adds hourly /home snapshots for one-command restores, zstd compression (often 20–40% smaller for code, the Nix store and models), subvolumes that share free space, checksums against silent corruption, and btrfs send backups to the bulk disk. Windows can't read it, which doesn't matter once Windows lives on its own USB.
Back up everything on C: and D: you want to keep, and build and test the Windows USB first. The MP700 is wiped. Use nodatacow on any VM-image or database folders.
NetBird access & groups
Default deny: the built-in All → All policy is deleted, and every flow is an explicit policy in fleet/netbird.nix, applied by OpenTofu after each merge.
| Group | Members |
|---|---|
| admin | janus, rhea, iris (phone), saturn: full access |
| g-platform | saturn |
| g-ops | titan, enceladus |
| g-dev | epimetheus, mimas, dione, telesto |
| g-aetherlink | pandora |
| g-blender | calypso |
| g-lab | iapetus |
| g-ai | hyperion |
| g-network | voyager, pioneer |
| ephemeral | AI sandboxes |
| grace-access | janus, hyperion, grace devices |
| From | To | Ports | Why |
|---|---|---|---|
| admin | all | all | Your devices |
| titan | all servers | 7700 | CI deploys with Panix |
| all servers | titan | 443, Loki | Cache, secrets, logs |
| titan | all servers | 9100 | Metrics |
| g-dev, ephemeral | hyperion | 8080, 4000 | LLM API (opt-in) |
| ephemeral | titan | 443 | Attic cache |
| grace-access | grace network | existing | Work |
<svc>.svc.moons.internal comes from a DNS zone on titan, generated from the inventory's services list. Caddy on each host routes it to the local port. The same entry generates the access policy. Services that need separate access get separate ports.
Guests also share saturn's Incus bridge and trust eth0, so they can reach each other there regardless of NetBird. Phase 16 closes this with guest firewalls, Incus ACLs and a bridge per project. Also: a lost phone in admin is a full-access device, so remove it from the dashboard right away.
AI sandboxes
Tell an AI what machine you want; it builds a NixOS container that joins the mesh by itself.
Claude Code
The AI calls the fleet MCP server on titan with a spec: packages, services, ports, size, optional group.
Nix on saturn
Spec + roles.sandbox + the LXC image are built into a NixOS system.
Incus
An unprivileged container in the sandbox project, with its own bridge and limits.
Ephemeral key
A one-use, 1-hour NetBird key puts it in ephemeral plus the requested group, under the next free moon name, e.g. kiviuq.moons.internal.
24 h reaper
Destroyed unless extended. NetBird drops the peer once it's offline. fleet promote turns it into a real host through a PR.
24 h, 4 CPU, 8 GiB, 30 GiB, up to 5 at once. Reaches the internet, titan's cache and hyperion's LLM API, nothing else unless its group allows it.
Only groups on an allow list (e.g. g-dev, g-blender). admin, g-network and g-ops can never be requested, so an AI can't create a full-access or infrastructure peer.
OpenTofu via terranix
Incus and NetBird are declared in Nix (terranix), generated from fleet/inventory.nix, and applied by OpenTofu from titan's CI.
| Layer | Tool |
|---|---|
| incus | OpenTofu, lxc/incus provider: projects, profiles, a bridge per project, network ACLs, storage, long-lived guests |
| netbird | OpenTofu, NetBird provider: groups, policies, setup keys, the svc.moons.internal DNS zone |
| inside hosts | NixOS flake + Panix |
| ai sandboxes | fleet MCP server → Incus API directly (create-on-request doesn't suit Terraform state); OpenTofu defines their project, bridge and limits |
Edit the inventory and push. CI posts tofu plan on the PR. On merge, tofu apply creates or changes the Incus guest first, then Panix deploys its NixOS config.
The pg backend in Postgres on enceladus, with OpenTofu state encryption (key in sops), backed up by restic. Existing projects and guests are imported once in phase 3, then the bootstrap script and the preseed are retired.
Master checklist
71 items, mirrored from docs/execution-plan.md. The agent ticks items there with each commit and updates this section after every milestone.
Milestone 1: Phase 0, hygiene
- 0.1 Stop copying the flake into
/etc(server profile, container-image) - 0.2
scripts/eval-all.sh+ record baseline drvPaths for every host - 0.3 Archive jimothy, ruru, tethys, homelab, al.prod,
templates/toattic/ - 0.4 Delete
modules/services/crowdsec/default_.nix - 0.5 Per-host
system.stateVersion(drvPaths unchanged) - 0.6
.envrc+ devShell tools (sops, age, ssh-to-age, nixfmt-tree) - 0.7 Fix
.cursor/rules/nix.mdcandnotes/core-incus.md - Owner: deploy 0.1 to the fleet (verified: /etc/nixos-config gone on all 8 guests; hestia first rolled back to gen 19 by hand)
Milestone 2: Phase 1, remove Tailscale
- 1.1 Delete
modules/networking/tailscale - 1.2 Remove every
saturn.networking.ts.enableline - 1.3 hermes: drop
services.tailscale.extraUpFlags - 1.4 janus: firewall, ports and comments
- 1.5 Hooks in ublockdns, netbird-server, netbird-proxy, incus
- 1.6 Notes cleaned
- Verify:
grep -ri tailscale hosts modulesis empty; nix-diff shows only Tailscale changes (0 refs; nix-diff: firewall only) - Owner:
tailscale logouton janus + hermes; hermes test-mode deploy; janusnh os switch; rest of fleet
Milestone 3: Phase 2a, NetBird safety on atlas
- Ask owner: running image versions (atlas, daphnis) (read from the live hosts by the agent)
- 2a.1 Pin server, dashboard, relay and proxy images
- 2a.2 Fixed compose subnet;
trustedHTTPProxies+trustedPeers+ relay trusted proxies - 2a.3 NetBird store and secrets backup timer
- Owner: manual backup (agent ran it, copies in ~/backups/atlas on saturn), test-mode deploy on atlas, check logs, real deploy
Milestone 4: Phases 2b–2e, NetBird (plan mode first)
- Ask owner: NetBird API token, live peer and group list (token in ~/.config/saturn/netbird-token; live lists read from the API)
- Merge into
modules/networking/netbird/{client,server,proxy}.nix, optionssaturn.netbird.* - Move Nix trust and the hestia CA out of the client module; drop old migration units
- Owner: deploy stage 4a (all 14 hosts verified live on the new build; hestia needed two reboots, see 854a61d)
fleet/netbird.nixgroups, policies, services, setup keys (OpenTofu NetBird provider). Groups and policies applied and verified; services and per-group setup keys still to do- Delete the All → All policy (default deny), add the required fleet flows (Default was already disabled; saturn-moons-mesh removed, admin and dev-to-ai flows live)
svc.moons.internalzone (CoreDNS on titan, later) + Caddy routing- Mesh domain
saturn.moons→moons.internal+ every hard-coded reference (incl. hermes/etc/hostspin) (f34a18d; live switch applied via tofu, verified) - Rename atlas → voyager, daphnis → pioneer; rename every peer to its hostname (d6c68d7, 9809909; the other peers already carried their hostnames)
- Remove the unused saturn-moons group
- saturn resolves mesh names: static pins from fleet/netbird.nix, since the uBlockDNS client cannot forward a domain (b6461be, deployed; all 13 hosts verified in sync)
- Per-group setup keys (janus is already a sops recipient; rhea waits for its NixOS install in Milestone 11)
Milestone 5: Phase 3, inventory, roles, OpenTofu
- Ask owner: unmanaged Incus leftovers stay out of OpenTofu, aether.link/production stays unmanaged, local state until Milestone 8, deploy-rs kept until Milestone 7 (plan approved)
fleet/inventory.nix(hosts, groups, tags, roles, descriptions, services); Incus placement and sops keys included, roles filled in step 5.3fleet/moon-names.nix+fleet name+ CI uniqueness check (nix run .#fleet -- name; checkfleet-names)- Roles: server, lxc-guest, dev, public-vps, ai-node, workstation (all 15 drvPaths identical except prometheus: a duplicate trusted-users entry removed). app-host, vm-guest and sandbox wait: nothing uses them yet (5f9c531..66fa91f)
- Generate
panix.yml,.sops.yamlanddeploy.nodesfrom the inventory (nix run .#fleet-config, checkfleet-config); deploy-rs kept until Milestone 7 by decision - terranix + OpenTofu (Incus + NetBird providers):
fleet/incus.nix,scripts/incus-tofu; local state by decision, moves to Postgres in Milestone 8 tofu importexisting Incus projects, profiles and guests; bootstrap script and preseed retired (ff46bad, bbba186; saturn deployed, 11 guests untouched)- Rename lapetus → iapetus
Milestone 6: Phase 4, titan, cache, CI
- Create titan (via OpenTofu)
- Attic + signing key; saturn as remote builder
- GitHub runner; CI builds every host (
checks) require-sigs = truefleet-wide; remove unsigned trust
Milestone 7: Phase 5, CI deploys
- CI runs Panix after merge: canary → fleet
- Approval gate for networking, access and hardening paths
- Deploy-guard watchdog (confirm or revert)
- Hourly catch-up via
system.configurationRevision - git-sync on workstations
Milestone 8: Phases 6–7, Infisical + Forgejo
- Infisical on titan (mesh-only), machine identities via sops
- Migrate Vault KV → Infisical; switch LiteLLM, cloudflared, etc.; remove Vault and vault-agent
- Rename hestia → enceladus
- Gogs removed early (stale data; module, hestia config, secrets contract and notes deleted; 120 MB backup in ~/backups/hestia-gogs on saturn)
- Forgejo on enceladus (nothing to migrate: the Gogs repos were stale); GitHub mirrors if wanted
Milestone 9: Phases 8–10, apps, dev, AI
- Ask owner:
/containers/*/docker-compose.ymlfrom each app host - Rename al.canary → pandora, blender → calypso
- Convert every compose stack to
modules/stacks/<svc>(oci-containers, pinned digests, Infisical env) - App template (
nix flake init -t …#app) + flake-input CD; first Discord bot on calypso - Rename prometheus → epimetheus, hyperion → telesto;
roles.dev, nix-direnv, remote editors - Rename hermes → hyperion; move to unstable; tool pins back on; drop the hermes / stable pins. Also fixes litellm-proxy, which crash-loops on the pinned nixpkgs (litellm 1.97.0 vs fastapi 0.141.1); unstable has 1.102.1, verified to start
Milestone 10: Phases 11–13, 16–17
- Ask owner: Discord webhook
- Prometheus, Grafana, Loki, Alertmanager → Discord; status page (
docs/status-mock.htmldesign) - restic → B2 + monthly restore test
- Weekly flake.lock update PR
- Incus VM image +
roles.vm-guest - Incus isolation: guests stop trusting eth0, bridge per project, ACLs (OpenTofu)
- AI sandboxes:
fleetMCP server,roles.sandbox, ephemeral keys, reaper,promote
Milestone 11: Phases 14–15, workstations
- Ask owner: backups of C: and D: done, Windows USB tested
- rhea: disko btrfs on MP700, SN850 btrfs data, NVIDIA, install
- janus btrfs reinstall (optional)