Deployment
Status: deployed 2026-08-18, deploy path corrected 2026-08-25. The flagship node, Fly app patches-social (org "personal", region iad), is live at patches-social.fly.dev and has been exercised end to end with two real accounts (register, login, post, follow, like, reply, thread, notifications, home feed — see "First deploy" below). infra/docker/Dockerfile and infra/fly/fly.toml are what shipped it. Media uploads use the production R2 bucket and verification email is sent through Resend from the verified noreply@updates.allisons.dev sender; federation is off by design. As of 2026-08-18 (A-041), production DATABASE_URL points at Neon, not the original Fly Postgres cluster — see "Production database" below. Sections describing genuinely not-yet-exercised steps (custom domain, autoscaling, log drain) still say Status: planned.
Deploy paths (read this before touching production)
What actually deployed every release up to 2026-08-25, verified via gh api repos/alliecatowo/patches/deployments (creator.login: "fly-io[bot]", description: "Deploying patches-social to Fly.io"): the Fly.io GitHub App, a dashboard-configured continuous-deployment integration installed on this repo — not any file in this repo, and not gated on CI. This is a real defect, not a documentation gap: gh api showed a production deployment (6084350830, 2026-08-25T13:38:02Z) created from the same push whose CI run (32854508360) had failed. There was never a .github/workflows/deploy.yml in this repo before this change — earlier versions of this doc described one; that was aspirational, not shipped (confirmed by ls .github/workflows/ on the live branch).
What deploys now: .github/workflows/deploy.yml, added 2026-08-25 — triggers via workflow_run on the CI workflow's completion (conclusion == success, head_branch == main), the same idiom .github/workflows/web.yml already used in this repo, plus a manual workflow_dispatch. It deploys both the server and worker process groups from one flyctl deploy --config infra/fly/fly.toml invocation, asserts the fly.toml ingress shape matches the postmortem fix (see docs/operations/postmortem-2026-08-24-web-outage.md), and verifies /healthz returns 200 afterward. It is gated behind repository variables FLY_DEPLOY_ENABLED and AUTH_CODE_ENVELOPE_ROLLOUT_COMPLETE (both true as of 2026-08-25 — see "One-time auth-envelope rollout" below for why the second one exists).
Known dual-path risk — owner action still required. The Fly GitHub App is NOT disconnected as of this doc's last edit. Until an owner disconnects it (Fly dashboard → patches-social → Settings → GitHub → Disconnect, or GitHub → Settings → Applications → Installed GitHub Apps → Fly → Configure → remove this repo from its repository access), pushes to main race TWO independent deploys: the new CI-gated workflow, and the old ungated Fly GitHub App. Do not treat deploy.yml's existence as having solved the "no CI gate" problem until that disconnect happens — until then the ungated path can still ship a red-CI commit regardless of what deploy.yml does. This agent's GitHub token could not inspect or disable the App installation itself (gh api .../installations needs admin:org/app-installation scope this token lacks; flyctl has no CLI surface for it — checked flyctl apps --help, flyctl deploy --help).
The demo node (patches-demo, ADR 0044) deploys from the same workflow in a separate deploy-demo job after production; see demo-node.md.
Deploying by hand in an incident (either path failed, or both need to be bypassed): the exact manual sequence is still the "First deploy" transcript below — flyctl deploy --config infra/fly/fly.toml --remote-only, run from a shell with FLY_API_TOKEN/flyctl auth login set up. Nothing about the new workflow removes that manual escape hatch.
Per INITIAL_VISION.md §§84–91, §122, §141.
Target architecture
Internet
|
Fly Proxy
|
HTTP/2 (TLS) -> h2c
|
+-----v------+
| server | <- Fly process group "server" (public)
| NestJS gRPC |
+-----+------+
|
+----------+-----------+
| |
v v
Neon Postgres Cloudflare R2
|
+----v-------+
| worker | <- Fly process group "worker" (private, no public service)
| NestJS |
| standalone |
+------------+One Docker image (infra/docker/Dockerfile), two Fly.io process groups (infra/fly/fly.toml's [processes]). No Kubernetes, no self-managed VM, no serverless decomposition.
Container: infra/docker/Dockerfile
Multi-stage build, four stages: base (Node 24 + pnpm via corepack + the buf CLI, needed because packages/proto's gen script has no npm-installable equivalent), fetch (warms the pnpm content-addressable store from pnpm-lock.yaml alone, cached across builds via a BuildKit --mount=type=cache), build (full source, offline install, pnpm turbo run build --filter=@patches/server --filter=@patches/worker, then pnpm --filter <app> deploy --prod --legacy /deploy/<app> to produce two self-contained app directories with workspace deps copied in rather than symlinked), and runtime (node:24-slim, non-root node user, no build toolchain, NODE_ENV=production).
Build context is the repo root (the pnpm workspace needs the whole monorepo), so the Dockerfile lives at infra/docker/Dockerfile but is always invoked from the repo root:
podman build -t patches:local -f infra/docker/Dockerfile .
# or: mise run docker:build (tries podman first, falls back to docker).dockerignore lives at the repo root (not next to the Dockerfile) — ignore files are resolved against the build context, not the Dockerfile's own directory.
Why --legacy on pnpm deploy
pnpm 10+ defaults pnpm deploy to "injected" (hard-linked) workspace dependencies, which assumes the deployed package stays inside the workspace's node_modules layout. That's wrong for a Docker runtime stage, which needs each app's node_modules to be genuinely self-contained so it can be COPY'd out on its own. --legacy switches to the old behavior — copying each workspace dependency's built files into the deployed package's own node_modules — which is what actually works here. Verified locally: pnpm --filter @patches/server deploy --prod --legacy /tmp/deploy-server produces a working, portable /tmp/deploy-server (confirmed node_modules/@patches/{config,database,media,proto} present as real directories, not symlinks); without --legacy it fails outright with ERR_PNPM_DEPLOY_NONINJECTED_WORKSPACE.
Neither apps/server/package.json nor apps/worker/package.json declares a "files" field, so pnpm deploy copies the whole package directory (src/, test/, tsconfig*.json, not just dist/) into the image — functionally harmless (dead weight, not a correctness issue) but slightly bloats the runtime layer (~100-120 MB per app directory before the shared base image). Adding "files": ["dist"] to those two package.jsons would trim this; left as a follow-up since P7-001's owned file set doesn't include apps/server/apps/worker source or manifests.
Running database migrations (release_command)
packages/database's own CLI (pnpm --filter @patches/database migration:run) runs TypeScript source through tsx, a devDependency that a pnpm deploy --prod image deliberately excludes. Rather than adding a production-only entrypoint inside packages/database (out of scope for P7-001), infra/docker/migrate.mjs calls the already-compiled createDataSource export that @patches/server already ships in its own node_modules after pnpm deploy — it's copied to server/migrate.mjs in the runtime image specifically so Node's ESM resolution finds @patches/database there. It reads DATABASE_URL/DATABASE_SSL/DATABASE_SSL_CA/DATABASE_LOGGING from the environment (same variables packages/database's own CLI reads), calls dataSource.initialize() + runMigrations({ transaction: 'each' }), and exits. See that file's header comment for the full rationale, and docs/operations/database.md for migration policy generally.
If packages/database ever grows its own compiled migration entrypoint, prefer that and delete infra/docker/migrate.mjs.
What's been verified locally (no Fly account needed)
podman build -t patches:local -f infra/docker/Dockerfile .— proto codegen (buf generate), and thetsupbuilds for@patches/config/@patches/media/@patches/proto/@patches/database/@patches/testkit, all succeed inside the container. Blocked at the finalapps/server/apps/workertscstep as of 2026-08-18 by an unrelated, concurrently in-progress change toapps/server/src/modules/auth/auth.guard.ts(a different agent's uncommitted work in this shared checkout, confirmed viagit status/git diff— not caused by anything in this Dockerfile). Re-run once that lands; nothing about the Docker mechanics themselves is in question past that point given how far the build got.pnpm --filter @patches/server deploy --prod --legacy /tmp/deploy-serverand the@patches/workerequivalent — both verified directly (see above).podman run --rm patches:local node server/dist/main.js --help— not yet run; blocked on the sameapps/serverbuild failure above (the image never finished building). Run this (or... node server/dist/main.jswithDATABASE_URLpointed atpostgres://patches:patches@host.containers.internal:5432/patchesto prove full boot) once the image builds.infra/fly/fly.toml— valid TOML (parsed with Python'stomllib), and its shape (oneapp,[processes]withserver/worker,[[services]]scoped toserveronly viaprocesses = ["server"],[deploy].release_command) matchesdocs/research/fly-io.md's verified Fly config syntax. Never deployed — no Fly account in this environment..github/workflows/deploy.yml(2026-08-18 note, since superseded) — this file did not actually exist in the committed repo as of 2026-08-25 (confirmed vials .github/workflows/); production deploys ran through the Fly.io GitHub App instead, with no CI gate. See "Deploy paths" above for the corrected picture and the file that exists now.
First deploy (2026-08-18)
The blocking apps/server/apps/worker tsc failure noted above (a different agent's concurrent, since-landed change) cleared, and the image built and deployed successfully. Exact commands run, in order:
flyctl apps create patches-social --org personal
flyctl postgres create ... # cluster patches-social-db
flyctl postgres attach -a patches-social patches-social-db # sets DATABASE_URL secret
flyctl secrets set --config infra/fly/fly.toml -a patches-social \
JWT_PRIVATE_KEY=... JWT_PUBLIC_KEY=... FEDERATION_KEY_ENCRYPTION_KEY=...
# (DATABASE_URL was already set by `postgres attach`, above)
flyctl deploy --config infra/fly/fly.toml --remote-only
# builds registry.fly.io/patches-social:build-<sha>, runs the release_command
# (node server/migrate.mjs) before the new server/worker Machines take traffic
flyctl proxy 15432:5432 -a patches-social-db
# separate terminal, used to run patches-admin invite create against prod Postgres:
# DATABASE_URL=postgres://...@127.0.0.1:15432/patches \
# pnpm --filter @patches/admin start invite create --by allie --max-uses 5Secrets set (names only — see docs/operations/local-development.md for what each is): DATABASE_URL, JWT_PRIVATE_KEY, JWT_PUBLIC_KEY, FEDERATION_KEY_ENCRYPTION_KEY. Not set (media/email disabled until dashboard-only credentials are fetched — tasks.md B-031): R2_ACCOUNT_ID, R2_ACCESS_KEY_ID, R2_SECRET_ACCESS_KEY, R2_BUCKET, R2_ENDPOINT, RESEND_API_KEY.
Non-secret env, from infra/fly/fly.toml's [env]: NODE_ENV=production LOG_LEVEL=log GRPC_HOST=0.0.0.0 GRPC_PORT=50051 HTTP_PORT=8080 PUBLIC_ORIGIN=https://patches-social.fly.dev NODE_DOMAIN=patches-social.fly.dev INVITE_ONLY=true FEDERATION_ENABLED=false EMAIL_PROVIDER=console EMAIL_FROM=noreply@patches-social.fly.dev.
E2EE is always-on (ADR 0036 Amendments, 2026-08-26 owner override) — there is no dev-mode flag, approval list, or narrowing env var. The v1 franking profile is the shipped construction, so GetE2eeCapability reports ENABLED as soon as the node has a franking signing key for its current era (e2ee_node_franking_keys).
A-052 (spec §197.6) operator-transparency env — also set in infra/fly/fly.toml's [env], published unauthenticated via NodeService.GetNodePolicy: PRIVACY_NOTICE_SUMMARY (what's stored, what's public, that DMs are end-to-end encrypted but visible to the node as metadata, retention, export/deletion, contact — summarizing docs/product/privacy.md), TERMS_URL (points at the full docs/product/privacy.md on GitHub — there is no separate ToS page), APPEAL_INSTRUCTIONS (the in-client Appeals screen, or email the operator), OPERATOR_CONTACT (who runs this node), and DATA_LOCATION (Fly for compute, Neon Postgres in aws-us-east-2 for the database, Cloudflare R2 for media). See the file itself for the exact published text.
These two values are published copy, not just config. PRIVACY_NOTICE_SUMMARY in both infra/fly/fly.toml and infra/preview/fly-preview.toml is served unauthenticated through NodeService.GetNodePolicy, so a stale value is a false public statement rather than a private misconfiguration. Both asserted the retired ADR 0017 wording ("Direct messages are NOT end-to-end encrypted — this node's operator can read them") for as long as it took B-097 to notice; that became false the moment B-095 removed the server-visible mode. They now mirror requiredConversationDisclosure('E2EE_V1') from @patches/domain — including its second clause, that the node still sees who you message and when. Whenever that domain constant changes, change these too, in the same commit, and redeploy: an accurate constant behind a stale env var still lies to everyone who asks the node what it does.
Gotcha: LOG_LEVEL is log, not info. The server's logger factory (apps/server/src/common/logging/logger.factory.ts) uses Nest's own LogLevel union, whose "normal operation" level is literally named log — setting LOG_LEVEL=info (a very natural guess coming from most other Node logging libraries) is accepted by nothing here and either falls through to a default or produces confusing output; use log.
Production bug found and fixed during this deploy: the Nest hybrid app (HTTP health listener + gRPC microservice in the same process) needed connectMicroservice(options, { inheritAppConfig: true }) — without that second argument, every handled application error surfaced to gRPC clients as bare INTERNAL instead of the mapped x-patches-error-code/status. See docs/agents/LEARNINGS.md for the full entry.
Live smoke check:
node apps/tui/dist/cli.js ping
# {"ok":true,"target":"patches-social.fly.dev:443",...}Verified end to end (tmux, two real accounts — first via register bootstrap since users was empty, second via an invite from patches-admin invite create run through the flyctl proxy above): register, login, whoami, compose + post, / search → profile, f follow, l like, g n notifications (LIKE + FOLLOW entries appear), r reply, Enter thread, g h home feed shows the followed account's post.
Production integrations verified: media uploads use R2 S3 credentials and the patches-media bucket; verification messages use Resend and the verified updates.allisons.dev sending domain. Federation remains disabled by design for v0.0 (FEDERATION_ENABLED=false).
"First deploy" checklist
- [x] Fly app created (
patches-social, org personal, regioniad). - [x] Postgres provisioned and attached (Fly Postgres cluster
patches-social-db; not yet Fly Managed Postgres — see "Production database" above). - [x] Required secrets set (
JWT_PRIVATE_KEY,JWT_PUBLIC_KEY,FEDERATION_KEY_ENCRYPTION_KEY;DATABASE_URLset automatically bypostgres attach). - [x] Image built and deployed (
flyctl deploy --config infra/fly/fly.toml --remote-only). - [x]
release_command(node server/migrate.mjs) ran migrations successfully before traffic cutover. - [x]
server(gRPC 50051, Fly TLS on 443,h2_backend) andworkerprocess groups both running. - [x] Smoke ping passes (
patches pingagainst the live host). - [x] End-to-end social loop verified with two real accounts.
- [x] R2 media credentials set and a live upload/derivative flow verified.
- [x] Resend API key and verified
updates.allisons.devsending domain configured. - [ ] Deploy workflow exercised through CI (planned — configuration is present, but the latest
mainCI run failed before the deploy gate; prior deploys were manual). - [ ] Custom domain
patches.social(planned — node currently only reachable atpatches-social.fly.dev). - [x] Neon switch (production
DATABASE_URLmigrated off Fly Postgres 2026-08-18, A-041 — see "Production database" above). - [ ] Autoscaling /
[[vm]]sizing tuned for real traffic (planned — default single Machine per process group so far, confirmed viaflyctl scale show --app patches-social2026-08-27:serverandworkereach at count 1, no scaling configured. An earlierinfra/fly/fly.tomlhad an invalid[services.scaling]block thatflyctl config validatesilently accepted but never applied — removed rather than fixed, since the correct per-service knobs (auto_stop_machines/auto_start_machines/min_machines_running,docs/research/fly-io.md) haven't been tuned against real traffic yet either). - [ ] Log drain wired up (planned —
fly logs/dashboard live-tail only today).
Process groups (infra/fly/fly.toml)
[processes]
server = "node server/dist/main.js"
worker = "node worker/dist/main.js"One image, two Fly Machines-per-group. Note the paths: pnpm deploy (see above) flattens each app into its own top-level directory in the image (server/, worker/), notapps/server/dist/main.js — a common mistake if copy-pasting an example that assumes the monorepo's own layout persists into the image.
If gRPC and a later HTTP federation listener end up awkward on the same public ports within one Fly app, the plan (per INITIAL_VISION.md §87) is to deploy separate Fly apps from the same image rather than build a bespoke reverse-proxy hack.
gRPC ingress
TLS terminates at the Fly edge; the server process group speaks plain h2c behind it. Per docs/research/fly-io.md §2 (verified 2026-08-18 against fly.io/docs), there is no dedicated "grpc" handler — gRPC rides HTTP/2 via [[services.ports]] handlers = ["tls", "http"] plus [services.ports.http_options] h2_backend = true, which is what infra/fly/fly.toml sets. Production client connections use TLS end to end from the client's perspective (grpc.patches.social:443, once DNS exists).
Health checks
Status: implemented; verify with flyctl checks list after deploy.
No native gRPC check type exists on Fly (documented — docs/research/fly-io.md §3; only http/tcp). infra/fly/fly.toml keeps the [[services.tcp_checks]] against the gRPC internal_port (50051) — it only confirms the socket accepts connections, not that the app is actually healthy — alongside a real application-level check (A-043):
[checks]
[checks.healthz]
type = "http"
port = 8080
method = "get"
path = "/healthz"
interval = "15s"
timeout = "3s"
grace_period = "10s"
processes = ["server"]GET /healthz (apps/server/src/modules/system/health.controller.ts / health.service.ts) reports 200 only when the database answers SELECT 1 and the grpc.health.v1.Health status is SERVING, 503 otherwise — a top-level [checks] entry rather than [[services.http_checks]] because port 8080 isn't a routed/public [[services]] block; [checks] targets a port directly, independent of request routing.
/healthz is served differently depending on FEDERATION_ENABLED (main.ts):
FEDERATION_ENABLED=false(production default, spec §176): a standalone, single-route HTTP listener (apps/server/src/modules/system/healthz-server.ts) bindsHTTP_PORTand answers only/healthz— deliberately not Nest's own HTTP adapter.AppModuleimportsFederationModule(webfinger/actor/inbox/outbox controllers) unconditionally, and those controllers don't re-checkFEDERATION_ENABLEDthemselves (unlikeFederationMetricsController) — they stay unreachable only because nothing binds a port for Nest's adapter. Binding that adapter to serve/healthzwould have also opened the federation HTTP surface on every node, contradicting the "zero new network surface when federation is off" invariant documented onFEDERATION_ENABLEDinenv.schema.ts. The standalone listener sidesteps that entirely.FEDERATION_ENABLED=true:main.tscallsapp.listen(HTTP_PORT)on Nest's full HTTP adapter as before, andHealthControlleranswers/healthzfrom the same port alongside the federation routes.
Both paths call the same HealthService.check(), so the response is identical either way. apps/server also implements the standard grpc.health.v1.Health service internally (grpc-health-check, see apps/server/src/grpc-options.ts) for gRPC-aware clients/probes — a different, additional thing from either Fly check above.
Production database
Status: implemented 2026-08-18 (A-041). Production DATABASE_URL on patches-social now points at Neon — project patches (id shy-recipe-96135980, org org-plain-leaf-04797948, region aws-us-east-2), default branch production (br-twilight-dew-axkmolfo), database neondb, role neondb_owner, sslmode=require. Get the current connection string (never print it in a log or commit it):
neonctl connection-string --project-id shy-recipe-96135980 --api-key "$NEON_API_KEY" \
| sed 's/&channel_binding=require//'(--api-key reads NEON_API_KEY, kept in the repo-root .env, gitignored — not committed. The sed strips channel_binding=require, which TypeORM's pg driver in this codebase doesn't need and which has caused connection issues with some pg client stacks; verified working without it.)
History: the first deploy (2026-08-18, see "First deploy" above) ran on a self-managed Fly Postgres cluster (patches-social-db, flyctl postgres attach) because neonctl wasn't authenticated in this environment at the time. Once it was, the data was migrated to Neon and DATABASE_URL/DATABASE_SSL secrets were repointed. The Fly Postgres cluster is now stopped and kept as a cold fallback (not actively serving traffic); its volume (vol_r1j3on1n5m85wpwr) has scheduled daily snapshots with 14-day retention (flyctl volumes update vol_r1j3on1n5m85wpwr --snapshot-retention 14) so it isn't itself an unbacked-up liability while it's kept around.
Cutting production over to a restored/branched database (e.g. after a Neon branch restore — see docs/operations/backups.md):
flyctl secrets set -a patches-social DATABASE_URL=<connection string> DATABASE_SSL=true
# rolls patches-social's Fly Machines onto the new DATABASE_URLNeon provides both point-in-time recovery and instant branching (neonctl branches create --project-id shy-recipe-96135980 --parent production) as backup/ restore primitives — see docs/operations/backups.md for the full backup/restore runbook, including a restore drill actually run against this project on 2026-08-18.
Not used: Fly Managed Postgres (fly mpg) was considered per docs/research/fly-io.md §6 and docs/decisions/0003-typeorm-postgres.md, but Neon was chosen instead for the actual migration — that ADR predates this decision and is not updated by this change (ADRs are architect's territory; see tasks.md A-041 for the open question of whether it needs a formal update or a new ADR).
Secrets
Set via fly secrets set, never committed. Full variable list: docs/operations/local-development.md's "Environment variables" section. Exact commands (fill in real values before running):
fly secrets set --config infra/fly/fly.toml \
DATABASE_URL="postgres://..." \
JWT_PRIVATE_KEY="$(base64 -w0 < jwt-private.pem)" \
JWT_PUBLIC_KEY="$(base64 -w0 < jwt-public.pem)" \
AUTH_CODE_DELIVERY_ACTIVE_KEY_ID="prod-YYYY-MM" \
AUTH_CODE_DELIVERY_KEYS='{"prod-YYYY-MM":"<32-byte-base64-key>"}' \
R2_ACCOUNT_ID="..." \
R2_ACCESS_KEY_ID="..." \
R2_SECRET_ACCESS_KEY="..." \
R2_BUCKET="patches-media" \
R2_ENDPOINT="https://<account-id>.r2.cloudflarestorage.com" \
RESEND_API_KEY="..." \
EMAIL_FROM="Patches <noreply@updates.allisons.dev>" \
GITHUB_CLIENT_ID="<client-id>" \
GITHUB_CLIENT_SECRET="<client-secret>" \
INVITE_ONLY="true"(fly mpg attach sets DATABASE_URL automatically — the line above is only needed if connecting an externally-provisioned Postgres instead.) JWT_PRIVATE_KEY/JWT_PUBLIC_KEY generation: pnpm keys:generate (see root package.json).
AUTH_CODE_DELIVERY_KEYS is a JSON keyring shared by the server (encrypts verification and password-reset jobs) and worker (decrypts them); AUTH_CODE_DELIVERY_ACTIVE_KEY_ID selects the write key. It must be independent of the JWT keypair. Rotate additively: add the new key to both processes first, deploy, switch the active id, wait until every pre-rotation auth job is terminal, then remove the old key.
One-time auth-envelope rollout
Status: planned — not exercised against Fly. The commands below are the reviewed operator sequence for the first rollout; do not treat them as a record of a completed deployment.
Migration 1787420562003-AuthCodeDeliveryEnvelopes adds a constraint that rejects the old plaintext auth-email job shape. Therefore the first rollout must not use the automatic GitHub deploy workflow: its normal rolling strategy could leave an old server producing plaintext jobs after the migration is applied. Use this one-time quiesced rollout, accepting a short sign-in/registration outage:
- Set the GitHub production environment variable
FLY_DEPLOY_ENABLED=falseand confirm no deploy run is active. Record current counts withfly scale show --app patches-social --config infra/fly/fly.toml. - Put only the two
AUTH_CODE_DELIVERY_*assignments above in a mode-0600 temporary file, then stage them without restarting Machines:fly secrets import --stage --app patches-social < auth-envelope-secrets.env. Remove that temporary file immediately after the command succeeds. - Build and push the reviewed commit while the old version is still serving:
fly deploy --app patches-social --config infra/fly/fly.toml --build-only --push --build-arg PATCHES_BUILD_SHA=<full-reviewed-commit-sha>. Record the exact registry image reference printed by Fly. - Quiesce every old producer and consumer:
fly scale count 0 --app patches-social --process-group server --yes, thenfly scale count 0 --app patches-social --process-group worker --yes. Confirm both are zero withfly scale showbefore continuing. - Deploy the already-built image (the release command applies the migration while no old process can enqueue):
fly deploy --app patches-social --config infra/fly/fly.toml --image <recorded-registry-image> --strategy immediate. - Restore the exact server and worker counts recorded in step 1 with separate
fly scale count <count> --process-group <group> --yescommands. Run the deployment smoke checks below and exercise verification and password reset. Finally set bothFLY_DEPLOY_ENABLED=trueandAUTH_CODE_ENVELOPE_ROLLOUT_COMPLETE=truein the GitHub production environment to enable routine later releases.
If build or review fails, stop before step 4. Once step 5 applies the constraint, do not roll back to the plaintext-producing image; fix forward with the reviewed envelope-aware image.
R2 bucket + CORS (Status: deployed): the dedicated patches-media bucket and scoped S3 credentials back the presigned-URL upload/derivative flow. Browser-origin CORS remains a client-specific follow-up; the current TUI upload path does not require it.
Resend (Status: deployed): updates.allisons.dev is verified and production sends as Patches <noreply@updates.allisons.dev> using the scoped RESEND_API_KEY Fly secret.
PUBLIC_READ (Status: implemented, default unchanged on the live deploy): owner decision, 2026-08-19 — INVITE_ONLY gates posting, not reading; this node's public content stays readable logged-out by default (PUBLIC_READ=true, the default, so nothing needs to change on patches-social.fly.dev to keep today's behavior). An operator who wants a fully closed node sets fly secrets set PUBLIC_READ=false (or the env var directly for a self-hosted node): every RPC outside SystemService.*, NodeService.GetNodeInfo/GetNodePolicy, and AuthService.* then requires a session (UNAUTHENTICATED/SIGN_IN_REQUIRED) — see apps/server/src/common/guards/public-read.guard.ts and docs/architecture/api.md §7.
PASSWORD_AUTH (Status: implemented, default unchanged on the live deploy): P15-002 — off | optional | required, default optional (unchanged behavior). An operator who wants to force SSH/GitHub-only sign-in sets fly secrets set PASSWORD_AUTH=off: Login, a password-carrying Register, and AddCredential(PASSWORD) then reject with FAILED_PRECONDITION/PASSWORD_AUTH_DISABLED, and AuthService.GetAuthPolicy tells clients to hide password UI. required is accepted but not yet enforced — see docs/architecture/auth.md §10.
Passkeys on the web client (Status: configured for the reference deploy): passkey credentials are bound to the configured WebAuthn RP ID. The reference node's web client is served from https://patches-web.pages.dev, so Fly sets PASSKEY_RP_ID=patches-web.pages.dev and PASSKEY_ORIGINS=https://patches-web.pages.dev. If the RP ID changes, any passkey enrolled under the previous RP ID is no longer usable: the RP ID is baked into the credential and users must re-enroll it.
GITHUB_CLIENT_ID / GITHUB_CLIENT_SECRET (Status: deployed 2026-08-30, #147): create an OAuth App at github.com/settings/applications/new ("Enable Device Flow" must be checked; homepage and callback set to https://patches-web.pages.dev; OAuth Apps have no create-API so the web form is the only path), then set both values on the app:
fly secrets set GITHUB_CLIENT_ID=<client-id> GITHUB_CLIENT_SECRET=<client-secret> -a patches-socialFly rolls/restarts the server machine automatically — no separate deploy step needed. The web client requires no separate secret because the device flow (BeginGitHubLogin / PollGitHubLogin in patches.v1.auth) runs entirely server-side. Verification confirmed via:
curl -s -X POST https://patches-social.fly.dev/patches.v1.AuthService/GetAuthPolicy \
-H 'content-type: application/json' -d '{}'
# → {"passwordAuth":"PASSWORD_AUTH_MODE_OPTIONAL","githubAuth":true}User-visible behavior: the web login page shows a GitHub button (anonymous = login or first-time register); a logged-in user links GitHub at Settings → Credentials ("Link another account"); in the TUI, the Accounts screen's add-account flow offers GitHub and runs the same device flow. Starting the device flow while authenticated links the credential to the current account; anonymously it logs in or registers. See also docs/architecture/auth.md §10 for the policy endpoint shape.
Error monitoring
Status: decided; log drain setup planned (needs live environment).
Decision (per INITIAL_VISION.md §100/§159 "error monitoring works or a documented alternative exists"): Patches v0 does not wire in Sentry (or any third-party error tracker). The documented alternative is structured JSON logs plus Fly's own log infrastructure:
apps/serverandapps/workeralready emit structured JSON logs in production (NODE_ENV=production) through the shared logger factory (apps/server/src/common/logging/logger.factory.ts) — every request/RPC is logged with a request-context (actor/session where available), and unhandled errors are surfaced as structuredrpc.errorevents byapps/server/src/common/errors/rpc-exception.filter.tsrather than bare stack traces.apps/worker's job runner (apps/worker/src/jobs/) uses the same logger factory for its own job lifecycle/error logging; dedicated federation delivery/inbox counters are tracked as a separate follow-up (seetasks.mdA-036) and are not yet emitted.- Once deployed,
fly logs/the Fly dashboard give a live tail of this structured output for free. For retention/search beyond Fly's short live-tail window, Fly supports shipping logs to an external log drain (destination — Loki, Datadog, or similar — is an operator choice, not prescribed here); the exactflyctl/fly.tomlincantation for wiring up a drain is not verified in this environment (no live Fly app to test it against) — confirm againstfly logs --help/current Fly docs at setup time rather than assuming a specific flag or config block from this note. - Alerting is log-based, not APM-based: a rule watching for
level: "error"/rpc.errorevents (or a spike in their rate) in whatever the log drain's destination supports (most log-drain backends — Grafana Loki, Datadog, etc. — support alert rules on log content), routed to whatever the team's current communication channel is at the time (seedocs/operations/incidents.md's detection step). - Why not Sentry: it would be the first third-party SaaS dependency purely for observability, with its own DSN-as-secret to manage and a stack-trace-shipping surface that needs privacy review before it ships to a hosted service; structured logs already carry the same "what broke and where" information without that added dependency, which fits the project's minimal-infrastructure bias (spec §153's no-unnecessary-managed-service posture, applied here by extension). Revisit if log-based alerting proves insufficient once there's real production traffic.
packages/observability'sinitializeTelemetry(used for OTel tracing, a separate concern from this error-monitoring decision) has no Sentry-specific code path — it only reads the genericOTEL_EXPORTER_OTLP_ENDPOINT/OTEL_EXPORTER_OTLP_HEADERSenv vars. An operator who wants traces routed to Sentry's OTLP ingestion (currently an open-beta Sentry feature, not GA) sets those two vars to Sentry's documented values themselves — seedocs/research/sentry-otlp.md(verified 2026-08-27) for the exact endpoint shape (https://o<orgId>.ingest.sentry.io/api/<projectId>/integration/otlp/v1/traces) and auth header (x-sentry-auth: sentry sentry_key=<public-key>, not a bearer token). An earlier revision ofinstrumentation.tshardcoded a wrong endpoint/header pair behind asentryDsnoption; that branch was removed rather than fixed, since the generic OTLP path already covers this and Sentry's beta API shape may still change.- Nothing here has been exercised against a live Fly log drain — there is no deployed node in this environment to point a drain at. The application-side structured logging itself is implemented and covered by existing server/worker tests; only the drain/alerting setup remains planned.
Domains
Per INITIAL_VISION.md §91:
patches.social marketing/docs
api.patches.social future HTTP API
grpc.patches.social gRPC
social.patches.social federation origin if desiredStatus: planned — no domain purchased/configured from this environment. fly certs add grpc.patches.social --config infra/fly/fly.toml provisions a Fly-managed TLS cert once DNS points at the app; confirm the current fly certs flags against fly certs --help at setup time (not independently re-verified for this note).
CI/CD
pull request
|
CI (format, lint, typecheck, buf checks, build, unit + integration tests, migration check)
|
merge to main
|
CI runs again on main -> ci-ok
|
Deploy workflow (workflow_run, triggered by CI's completion) -> flyctl deploy --remote-only
| (release_command runs migrations first, before new Machines take traffic)
|
verify: GET https://patches-social.fly.dev/healthz returns 200 (retried up to ~100s).github/workflows/deploy.yml (added 2026-08-25 — see "Deploy paths" above) implements this, using the same workflow_run-on-CI idiom as web.yml, plus workflow_dispatch for a manual redeploy. FLY_API_TOKEN is configured. Routine deployment requires both vars.FLY_DEPLOY_ENABLED=true and vars.AUTH_CODE_ENVELOPE_ROLLOUT_COMPLETE=true (both true as of 2026-08-25 — the one-time rollout above is complete); an operator can set either back to false to pause automatic deploys without editing the workflow file. Deploy credentials are never exposed to pull requests from forks — the workflow only triggers off workflow_run (main-only, same-repo) and manual dispatch, both of which run with the repo's own secrets, never a fork's. actionlint and python3 -c "import tomllib; ..." both pass on the workflow file and infra/fly/fly.toml respectively (checked 2026-08-25); the workflow itself has not yet completed a real CI-triggered deploy from this environment (no way to push to main and wait for it here).
Depot-accelerated builds (issue #262, 2026-08-28): the deploy step passes --depot=auto --depot-scope=org explicitly to flyctl deploy — docs/research/fly-io.md §10 confirms this is flyctl's own default as of v0.4.92 (flyctl deploy --help), made explicit so a future flyctl default change can't silently drop it. Fly runs the Depot builder itself; no separate Depot account, token, or CI step is needed. Combined with the BuildKit pnpm-store cache mounts already in infra/docker/Dockerfile and the runner-side pnpm store cache in .github/actions/setup/action.yml, repeat deploys of this monorepo image should reuse both the CI-runner pnpm store and Depot's own layer cache. Status: verified via flyctl deploy --help output only — no live flyctl deploy was run from this environment (would deploy to the real patches-social production app, out of scope for this change); measure the actual before/after deploy time on the next real production deploy.
Smoke tests
.github/workflows/deploy.yml's final step curls https://patches-social.fly.dev/healthz directly (up to 20 attempts, 5s apart) rather than a gRPC client round trip — simpler than standing up a gRPC smoke client in the runner, and it's the exact same endpoint infra/fly/fly.toml's own [checks.healthz] already polls, so a workflow failure and a Fly health-check failure mean the same thing. Status: planned — never run against a real deployment from this environment (same reason as above).
The workflow passes the exact deployed commit as PATCHES_BUILD_SHA at image build time, same as before this change — GetServerInfo.server_version (queryable via patches ping or the TUI) still reports <package-version>+<short-sha> in deployed images.
Graceful shutdown
Already implemented in application code (apps/server/src/main.ts calls app.enableShutdownHooks() and flips the gRPC health status to NOT_SERVING on SIGTERM/SIGINT before Nest's own shutdown hooks drain the process; apps/worker similarly stops claiming new jobs on shutdown — see apps/worker/src/main.ts). Fly can terminate Machines on its own schedule; nothing here assumes a generous shutdown window.
npm packaging (patches-social, P9-003 / A-046)
See apps/tui/README.md's "Self-contained build" section for the full picture. Summary: apps/tui/tsup.config.ts bundles the app and its three private workspace dependencies (@patches/domain, @patches/proto, @patches/terminal-media) into a single dist/cli.js, so the published tarball no longer needs those unpublished packages resolved separately — the earlier known gap (an npm install -g 404ing on them) is closed.
The npm-facing package name is patches-social (the bare patches name is taken; checked 2026-08-18). It's set via publishConfig.name in apps/tui/package.json — a pnpm 11.18+ feature (this repo pins 11.22.0) that publishes under a different name than the workspace's own package.json name, which stays @patches/tui. That means every --filter @patches/tui reference elsewhere in the repo (mise.toml, root package.json, CI workflows) is untouched by this.
Verified locally (2026-08-18), from a shell with the repo's node_modules off PATH, against the live node:
pnpm --filter @patches/tui build
pnpm --filter @patches/tui pack --pack-destination /tmp/patches-tui-pack
# -> /tmp/patches-tui-pack/patches-social-0.1.0.tgz (~80 KB)
mkdir -p /tmp/pfx/bin
PATH="/tmp/pfx/bin:$PATH" PNPM_HOME=/tmp/pfx pnpm add -g /tmp/patches-tui-pack/patches-social-*.tgz
cd /tmp && /tmp/pfx/bin/patches --version # -> 0.1.0
cd /tmp && /tmp/pfx/bin/patches ping --server patches-social.fly.dev:443 # -> {"ok": true, ...}Inspecting the packed package.json confirmed pnpm-workspace.yaml's catalog: entries resolve to concrete semver ranges at pack time (e.g. "ink": "^7.1.1", not the literal string catalog:) — required for the tarball to be installable by a plain npm install, not just within this pnpm workspace.
GitHub prerelease tarballs are published and are the supported installation channel today; see docs/operations/try-it.md for the current URL.
npm registry status: blocked on authentication — publishing itself (npm login, then mise exec -- pnpm --filter @patches/tui publish --access public, using the workspace-local filter name; pnpm rewrites the published name to patches-social via publishConfig.name) is a manual, one-time step for the package owner and hasn't been run — npm login isn't available in this environment. Everything up to producing and installing the tarball is verified above.
Related documents
docs/research/fly-io.md— the verified Fly.io config reference this doc builds on.docs/operations/database.md— migration policy.docs/operations/backups.md— backup/restore procedure.docs/operations/incidents.md— rollback and incident response.
Scale to zero (2026-10-02)
infra/fly/fly.toml runs a single server Machine with auto_stop_machines = "stop", auto_start_machines = true, min_machines_running = 0 on the http and grpc services. The job worker no longer has its own Machine: WORKER_IN_PROCESS=true makes the server spawn worker/dist/main.js as a child process (apps/server/src/embedded-worker.ts), because a service-less worker Machine can never be woken by traffic. Jobs are claimed while the Machine is awake; delayed or scheduled jobs (retention sweeps, 30-day purge) run on the next wake. A long-lived gRPC/TUI connection keeps the Machine up. After deploying, destroy the old worker Machine if it still exists (fly machine destroy <id> -a patches-social --force).