Files
jobtrackingapp/deploy/README.md
T

330 lines
14 KiB
Markdown

# Production deployment notes
## Gitea Actions
This repo includes `.gitea/workflows/ci-deploy.yml` for:
- backend build
- backend tests
- frontend tests
- frontend production build
- deployment to Ubuntu after successful tests on `main`
### Required secrets in Gitea
- `PROD_HOST`
- `PROD_USER`
- `PROD_SSH_KEY`
## Ubuntu server setup
Recommended app path:
- `/opt/job-tracker/app`
Persistent runtime secrets path:
- `/opt/job-tracker/shared/.env`
Requirements:
- Docker Engine
- Docker Compose plugin
- reverse proxy in front (Nginx, Caddy, or Traefik)
- shared env file present on server in `/opt/job-tracker/shared/.env`
- network connectivity from the backend container to your `mariadb` container/service
The deploy script will automatically create a symlink from:
- `/opt/job-tracker/shared/.env`
to:
- `/opt/job-tracker/app/.env`
This keeps secrets outside the uploaded repo checkout so they are not wiped by CI deploys.
### Frontend API base URL
The production frontend already proxies `/api` to the backend container via Nginx.
Recommended default:
- leave `NEXT_PUBLIC_API_BASE_URL` unset/empty in production
Only set `NEXT_PUBLIC_API_BASE_URL` if the UI must call a different external API origin on purpose.
## Example production `.env`
```env
DATABASE_PROVIDER=mariadb
JOBTRACKER_CONNECTION_STRING=server=mariadb;port=3306;database=jobtracker;user=jobtracker;password=REPLACE_ME
AUTH_JWT_KEY=replace_with_long_random_secret
AUTH_ADMIN_EMAIL=you@example.com
AUTH_ADMIN_PASSWORD=replace_with_strong_password
AUTH_REQUIRE_EMAIL_VERIFICATION=true
APP_PUBLIC_BASE_URL=https://your-domain.example
WEB_PROXY_SUBNET=172.31.250.0/29
STRIPE_SECRET_KEY=sk_live_...
STRIPE_PRICE_PREMIUM=price_...
STRIPE_WEBHOOK_SECRET=whsec_...
AI_SERVICE_BASE_URL=http://ai-service:8001
OLLAMA_BASE_URL=http://ollama:11434
OLLAMA_MODEL=qwen2.5:7b
EMAIL_FOLLOWUPREMINDERS_ENABLED=true
EMAIL_FOLLOWUPREMINDERS_UPCOMINGDAYS=2
# Optional backward-compatible alias if older config still references the previous name:
SUMMARIZER_BASE_URL=http://ai-service:8001
```
## Database recommendation
For production, yes — use a real database.
### Recommended direction
Short term:
- SQLite is acceptable for a single-user or very small deployment
- keep backups and volume persistence
Better production choice:
- MariaDB or PostgreSQL
### My recommendation
- **PostgreSQL** if you want the best long-term maintainability and fewer edge cases
- **MariaDB** is also fine if that is what you already know or host elsewhere
If you stay on SQLite:
- okay for small personal use
- not ideal for concurrent writes, larger scale, or operational robustness
## Practical recommendation for this project
If this app is going to be a real production service on Ubuntu:
- move to PostgreSQL first if possible
- MariaDB is still a reasonable option if preferred
## Deployment flow
Production automation always selects `docker-compose.yml` explicitly. Local development must add
`docker-compose.dev.yml`; never add that file to a production command. The base configuration has no
host bindings for frontend, backend, ai-service, or bundled Ollama.
The external Traefik configuration is operator-owned and is not stored here. Before deployment it
must route only the exact host from `APP_PUBLIC_BASE_URL` to frontend port 80 on
`jobtracker_shared`, terminate TLS, replace `X-Forwarded-For` and `X-Forwarded-Proto=https`, and
expose no direct application or Ollama host ports.
Nginx independently rejects non-canonical Hosts except `/health`, forwards Traefik's sanitized
single-hop values, and reaches the backend only over `WEB_PROXY_SUBNET`. The backend fails startup
if forwarded-header trust is enabled without a valid known CIDR.
1. push to `main`
2. Gitea Actions runs tests
3. if green, workflow uploads repo to server
4. `deploy/deploy.sh` links `/opt/job-tracker/shared/.env` into the repo checkout, then explicitly runs `docker compose -f docker-compose.yml build` and `up -d`
5. if `OLLAMA_MODEL` is set, the deploy script waits for Ollama, pulls the configured model if missing, then restarts `ai-service` so hybrid CV classification can use it
6. workflow checks service status after deployment
## Post-deploy verification you should also do manually the first time
- confirm reverse proxy routes to the frontend correctly
- confirm API auth/login works with production config
- confirm backend can connect to MariaDB
- confirm AI service container is reachable from backend
- confirm reminder and admin/system pages load
- verify follow-up reminder emails are enabled only when intended and that links open the correct job/tab
hat links open the correct job/tab
---
# Backups, restore and rollback
**Database restore and application rollback are two different operations.** A bad deploy usually needs
only the rollback. Restore the database only if the data itself is wrong or lost — it discards
everything written since the dump.
## Environment loading
`deploy/deploy.sh` **loads `/opt/job-tracker/shared/.env` into its own shell** before it decides
anything. The symlink it creates in the checkout is for docker compose, which reads `.env` itself;
the script needs the values too, to pick the right backup and to check its own configuration.
- Parsed line by line, not `source`d — a compose `.env` is not a shell script, so an unquoted value
containing spaces would execute as a command.
- **Variables already set in the environment win**, so CI-provided `APP_VERSION`, `APP_COMMIT_SHA`
and `APP_BUILD_STAMP` still override the file.
- No value is ever echoed. Error messages name variables, never their contents.
This was added on 2026-07-19. Before it, the script read an empty environment: `DATABASE_PROVIDER`
fell back to `sqlite` on a MariaDB host, so the deploy tarred the data volume, printed
`Backup verified`, and continued with no database dump at all. See
`docs/release-candidate-review.md` (B1).
## Required production variables
`validate_deploy_config` runs **before** the backup, and therefore before anything is built, stopped
or replaced. A missing variable aborts the deploy while the running stack is still untouched.
| Variable | Required | Why it is checked here |
|---|---|---|
| `DATABASE_PROVIDER` | **Always** — no default | Selects the backup. Guessing it wrong backs up the wrong database and reports success |
| `JOBTRACKER_CONNECTION_STRING` | When provider is `mariadb`/`mysql` | Without it there is no way to dump the database |
| `AI_SERVICE_TOKEN` | Always | `docker-compose.yml` declares it with `:?`; missing it kills the stack *after* the images are built |
| `AUTH_JWT_KEY` | Always | Compose sets `Auth__Require=true`, and the backend throws at startup on a blank key — after the containers have been replaced |
| `APP_PUBLIC_BASE_URL` | Always | Canonical HTTPS origin for links, OAuth callbacks, billing redirects, secure cookies, Host validation, and the public smoke check |
| `AUTH_MICROSOFT_TENANT` | When `AUTH_MICROSOFT_CLIENT_ID` is set | Exact Microsoft application sign-in account mode; distinct from the Graph mailbox tenant |
| `WEB_PROXY_SUBNET` | Always | Dedicated nginx-to-backend CIDR trusted for exactly one forwarded hop; must not overlap another Docker network |
### Microsoft sign-in migration gate
Before enabling `AUTH_MICROSOFT_CLIENT_ID` with the canonical identity release, take the normal
backup and record counts only. Do not print subjects or email addresses:
```sql
SELECT COUNT(*) AS legacy_links
FROM AspNetUsers
WHERE MicrosoftSubject IS NOT NULL OR MicrosoftEmail IS NOT NULL;
SELECT COUNT(*) AS legacy_without_alternate_credential
FROM AspNetUsers
WHERE (MicrosoftSubject IS NOT NULL OR MicrosoftEmail IS NOT NULL)
AND PasswordHash IS NULL
AND GoogleSubject IS NULL;
SELECT COUNT(*) AS duplicate_legacy_subject_groups
FROM (
SELECT MicrosoftSubject
FROM AspNetUsers
WHERE MicrosoftSubject IS NOT NULL
GROUP BY MicrosoftSubject HAVING COUNT(*) > 1
) duplicate_subjects;
SELECT COUNT(*) AS duplicate_legacy_email_groups
FROM (
SELECT MicrosoftEmail
FROM AspNetUsers
WHERE MicrosoftEmail IS NOT NULL
GROUP BY MicrosoftEmail HAVING COUNT(*) > 1
) duplicate_emails;
```
Apply `20260802212509_AddCanonicalMicrosoftIdentity` before deploying code that queries the two new
columns. The migration does not backfill legacy rows and adds a unique nullable composite index.
Keep Microsoft sign-in disabled if the inventory or migration fails. Roll back the application
binary while leaving the additive columns in place; never roll back to email auto-linking.
`DATABASE_PROVIDER` deliberately has **no default**. An unset value used to mean "sqlite"; it now
means "stop and tell me".
## Backup creation
`deploy/deploy.sh` takes a backup **before** it builds, stops or replaces anything, and **aborts the
deploy if the backup fails**. Nothing else in the deploy runs without a restore point.
- **Location:** `/opt/job-tracker/backups` — override with `BACKUP_DIR`. If an earlier root-run deploy owns the directory, the deploy script repairs its ownership through Docker before writing.
- **Selection:** driven solely by `DATABASE_PROVIDER`, which must be set.
- `mariadb` / `mysql` → SQL dump, `jobtracker-<database>-<UTC timestamp>.sql.gz`, e.g.
`jobtracker-jobtracker-20260719T153759Z.sql.gz`
- `sqlite` → data volume archive, `jobtracker-sqlite-<UTC timestamp>.tar.gz`
- anything else → the deploy stops
- **The filename tells you which path ran.** If you expect a MariaDB deploy and find a
`jobtracker-sqlite-*.tar.gz`, the environment is wrong — that is the exact failure this check exists
to make visible.
- **Naming:** the UTC timestamp makes every file unique, so a deploy never overwrites an earlier backup.
- **Credentials** come from `JOBTRACKER_CONNECTION_STRING` and are passed via `MYSQL_PWD`, never on the
command line, so they cannot appear in the process list or the deploy log.
- **Compression:** gzip. A small database compresses to a few KB.
- **SQLite volume resolution:** compose prefixes volume names with the project name, so the script
resolves `<project>_jobtracker_data` and **fails if that volume does not exist**. Naming the bare
volume would silently create an empty one and back *that* up.
### Verification — each provider gets the check that proves its own format
A backup that exists but is empty, truncated, or the wrong *kind* is worse than none, because it looks
like a restore point.
| Provider | Checks |
|---|---|
| MariaDB | non-empty; valid gzip; contains `CREATE TABLE`; contains the `Dump completed` trailer that `mariadb-dump` writes last, so a dump that died partway through is rejected |
| SQLite | non-empty; valid gzip; the archive actually contains `jobtracker.db` |
A failed check deletes the file rather than leaving something that looks like a backup.
### Taking one by hand
```bash
BACKUP_DIR=/opt/job-tracker/backups
mkdir -p "$BACKUP_DIR"
MYSQL_PWD='<password>' mariadb-dump \
--host=127.0.0.1 --port=3306 --user=<user> \
--single-transaction --routines --events --quick \
jobtracker | gzip -c > "$BACKUP_DIR/jobtracker-manual-$(date -u +%Y%m%dT%H%M%SZ).sql.gz"
```
### Retention
**Nothing is deleted automatically.** Backups accumulate in `BACKUP_DIR` until you remove them. Watch
disk usage and prune deliberately — a suggested policy is to keep every backup for 30 days and one per
month after that, but the script does not enforce it and will not delete your files.
## Restoring the database
Tested end-to-end against a clean MariaDB 11 container: dump taken from a seeded database, restored
into an empty one, rows verified identical.
```bash
# 1. Stop the application so nothing writes during the restore.
docker compose stop backend
# 2. Restore. This REPLACES the current contents of the named database.
gzip -dc /opt/job-tracker/backups/jobtracker-jobtracker-20260719T153759Z.sql.gz \
| MYSQL_PWD='<password>' mariadb --host=127.0.0.1 --port=3306 --user=<user> jobtracker
# 3. Verify before starting anything.
MYSQL_PWD='<password>' mariadb --host=127.0.0.1 --port=3306 --user=<user> jobtracker \
-e "SELECT COUNT(*) AS applications FROM JobApplications;"
# 4. Start again.
docker compose start backend
```
Restoring a **SQLite** deployment instead:
```bash
docker compose stop backend
docker run --rm -v jobtracker_data:/data -v /opt/job-tracker/backups:/backup \
-e ARCHIVE_NAME=jobtracker-sqlite-20260719T153941Z.tar.gz \
alpine:3 sh -c 'rm -rf /data/* && tar xzf "/backup/$ARCHIVE_NAME" -C /data'
docker compose start backend
```
## Restoring application containers (rollback)
This is the usual fix for a bad deploy, and it **does not touch the database**.
```bash
cd /opt/job-tracker/app # the deployment checkout
git log --oneline -5 # find the last good commit
git checkout <previous-commit>
deploy/deploy.sh
```
`deploy.sh` takes a fresh backup first, so rolling back is itself protected.
**Why a code rollback is safe here:** every Phase 4/5 migration is a no-op — the startup reconciler
owns those tables — so reverting the code never leaves migration history ahead of the schema. The
reconciler is additive and never drops a table holding rows, so the older code simply ignores the
newer tables.
**What a code rollback does not undo:** rows users created in the newer tables stay. That is usually
what you want. If you additionally restore the database, those rows are lost — so restore only when
the data is the problem.
## Choosing between them
| Symptom | Action |
|---|---|
| New version starts but behaves wrong | Rollback the code. Leave the database. |
| Backend will not start; schema looks wrong | Rollback the code, then restore only if it still fails. |
| Data is missing or corrupted | Restore the database from the most recent good dump. |
| Deploy aborted before starting | Nothing to undo — the backup ran before any change. |
## Health checks
`backend` and `frontend` both report container health, so `docker compose ps` shows real state rather
than merely "running".
- **Backend:** `curl -fsS http://127.0.0.1:8080/health`. Anonymous, and deliberately **does not touch
the database** — a health check that queried MariaDB would restart a healthy backend whenever the
database blipped. `start_period` is 90s to cover first-boot schema reconciliation.
- **Frontend:** `wget` against nginx on port 80.
- `frontend` waits for `backend` to be *healthy*, not merely started, because nginx proxies `/api` to
it and refuses to start if the upstream cannot be resolved.
A backend that cannot reach its database exits and is reported `unhealthy`, so a broken deploy does not
present as a running stack.