Two post-reboot runs: the first reproduced the failure identically (smoke 1s pass, suite 3s fail), the second failed earlier at `dotnet restore` in 0s — a step that succeeded in 3-4s on every previous run, same commit, same runner. That rules out stuck state (reboot changed nothing) and rules out a deterministic sandbox policy such as seccomp/W^X blocking runtime IL emission, which was the leading remaining hypothesis. Combined with host telemetry showing no disk/memory/PID pressure, confidence in any specific mechanism drops to ~25%; confidence that application code is not the cause stays high. Removes the pure-vs-Moq diagnostic scaffolding (it never executed). No test skipped or weakened. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
13 KiB
Infrastructure Investigation — CI runner and deploy failures
2026-07-18. Supersedes
docs/ci-runner-investigation.md. Conclusion: both failures are outside the repository. Application code has been eliminated as a cause by direct experiment. Confirmation and repair require host access — the exact asks are at the end.
There are two independent infrastructure failures:
- A — CI
testjob: the backend suite fails only on the self-hostedLive-Runner. - B — CI
deployjob: the SSH step fails ~3 s in, before doing any work.
They are unrelated to each other and to the application code.
A — Backend suite fails only on the runner
Evidence gathered
- The suite had never actually run in CI. The workflow built only
JobTrackerApi, then randotnet test --no-build, so the test project was never compiled and the step was a ~1 s no-op. Fixed incfba7fb. The failure is newly surfaced, not a regression — it may be long-standing. - Job logs are unreadable:
GET /api/v1/.../actions/jobs/{id}/logs→401 token is required. Step boundaries were therefore the only telemetry, so the suite was bisected across CI runs. - Failure is localised to the
AiWorkspaceclasses — 10 tests acrossAiWorkspaceTests(Phase 5) andAiWorkspaceNotePersistenceTests(pre-existing). Locally these run in 1 s.
| Run | Steps observed | Reading |
|---|---|---|
| 524 | Test backend ✗ 8 s |
one combined step, no detail |
| 525 | restore ✓ 4 s, build ✓ 4 s, test ✗ 3 s | not a restore or compile error |
| 526 | host smoke ✓ 1 s, full suite (serial) ✗ 3 s | test host starts fine; not parallelism |
| 527 | quarters: A–C ✗ 3 s, rest never ran | offender is alphabetically early |
| 528 | per class: T AiWorkspace ✗ 3 s, others never ran |
offender named |
Experiments performed
Every experiment ran the same commit. All pass unless stated.
| # | Experiment | Result |
|---|---|---|
| 1 | Windows host, full suite | 306 pass |
| 2 | Clean mcr.microsoft.com/dotnet/sdk:9.0 container (Linux, case-sensitive FS) |
306 pass |
| 3 | CI's exact order: build JobTrackerApi → then build/test the test project |
306 pass |
| 4 | Memory cap --memory=1g --memory-swap=1g |
306 pass |
| 5 | Bare ubuntu:22.04, SDK via dotnet-install.sh into $HOME/.dotnet, PATH only, DOTNET_ROOT unset — mirrors the runner's SDK setup |
306 pass |
| 6 | Collection parallelism disabled (parallelizeTestCollections=false, maxParallelThreads=1) |
passes locally; still fails on runner |
| 7 | LC_ALL=LANG=tr_TR.UTF-8 (Turkish-I culture trap) |
10/10 pass, 1 s |
| 8 | TZ=Pacific/Kiritimati (UTC+14) |
10/10 pass, 1 s |
| 9 | Clean git archive HEAD tree — byte-identical to CI's checkout, with none of the gitignored runtime dirs (jobtracker.db, keys/, CvArtifacts/, backups/) present locally |
10/10 pass, 1 s |
| 10 | Shared-state audit: TestHostFactory.CreateInMemoryDb uses Guid.NewGuid() per test |
no shared store |
Experiment 9 is the decisive one: it removes the last difference between the local tree and the runner's checkout. The exact source CI compiles produces a passing suite.
Root cause hypothesis
The runner host kills the test process. The workflow already documents three separate failure modes on this same runner, all with the signature of a process dying with no usable error output:
actions/setup-dotnet"intermittently leaves a partial extraction in the shared tool-cache (tar: Cannot open: File exists) or corrupts the SDK download" — hence the hand-rolled installer.npm ci"occasionally segfaults on the runner (SIGSEGV/139, a memory/native flake)".- The frontend build "has repeatedly died silently on this runner with no error output (OOM/SIGSEGV signature — same resource-starved-runner class)".
A .NET test host exiting ~3 s into a 10-test run belongs to that same family. Two candidate mechanisms, in order of likelihood:
- Resource exhaustion — memory or PID/thread limits. The runner appears to share the host with
the production Docker stack. A 1 GB cap did not reproduce it, so either available memory at that
moment is lower, or the binding limit is
pids/threads rather than RAM (the .NET test host spawns more threads thannpm ci, so it would hit a lowpids.maxfirst). - Disk exhaustion. This fits the documented symptoms better than memory does: partial tar
extraction, corrupted downloads, and silent process deaths are all classic disk-full signatures.
testhostwritesTestResults/and may write dumps.
Host telemetry (2026-07-18, post-reboot) — resource exhaustion ruled out
Collected from the server:
/ 217G 146G 62G 71% (inodes 14% used)
/dev/shm 16G 0 16G 0% (=> ~32 GB RAM)
ulimit -u 127749 ulimit -n 1024 ulimit -m unlimited
/sys/fs/cgroup/pids.max: No such file or directory
fail2ban-client: command not found
journalctl -u sshd: No entries (Ubuntu's unit is `ssh`, not `sshd`)
dmesg: read kernel buffer failed: Operation not permitted (needs sudo; cleared by reboot anyway)
This falsifies both resource hypotheses at the host level: there is no disk pressure, no inode pressure, no memory pressure and no restrictive PID limit. It also falsifies the fail2ban explanation for the deploy failure — fail2ban is not installed.
Caveat that matters: Gitea act_runner normally executes a runs-on: ubuntu-latest job inside
a Docker container, so the figures above describe the host, not the environment the tests
actually ran in. A job container has its own cgroup memory/PID limits and, by Docker default, a
64 MB /dev/shm regardless of the host's 16 GB. The relevant limits have therefore not been
measured yet — see the asks below.
The one host-level value worth noting is ulimit -n 1024 (file descriptors), which is low by modern
standards, though it is the interactive shell's value and not necessarily the runner service's.
Post-reboot runs — the failure moves between stages
Two further runs after the machine was restarted:
| Run | Result |
|---|---|
8f73548 (post-reboot) |
Identical to pre-reboot: host smoke ✓ 1 s, full suite ✗ 3 s |
c4c0cd4 |
Failure moved earlier: Restore backend tests ✗ 0 s — a step that took 3–4 s and succeeded in all previous runs. None of the queued diagnostic steps executed. |
Two conclusions:
- The reboot changed nothing, so this is not stuck state, a leaked process, or a corrupted workspace that a restart would clear.
- The failing stage is not stable across runs.
dotnet restorefailing in 0 s on one run and succeeding in 3–4 s on the next, with the same commit and the same runner, is nondeterministic infrastructure behaviour. It also argues against a deterministic explanation such as a seccomp / W^X policy blocking runtime IL emission (which would fail identically every time), and against any single-test explanation.
Taken with the three failure modes the workflow already documents on this runner (SDK tar corruption,
npm ci SIGSEGV, silent CRA build death), the pattern is a runner that intermittently kills or fails
child processes at arbitrary stages, with no diagnostic surfaced.
Confidence level
- Application code is not the cause — high confidence (~95 %). Ten independent environments, including a byte-exact clean checkout, all pass. Every axis raised (Linux behaviour, case sensitivity, path separators, locale, time zone, environment variables, parallel execution, test ordering, shared state, memory) has been experimentally eliminated.
- Specific mechanism — low confidence (~25 %), and lower than before. Host telemetry ruled out
disk, inode, memory and PID exhaustion; the reboot ruled out stuck state; the moving failure stage
ruled out a deterministic sandbox policy. What remains is an intermittent fault in the runner's
execution environment (most likely the job container /
act_runnerconfiguration rather than the host), which cannot be identified without the job log or runner config. I am not asserting a mechanism.
Why application code is no longer suspected
- The identical commit passes in nine environments, including one built from
git archive HEAD— exactly what CI checks out, with no local-only files. - The 10 failing tests use EF InMemory with a per-test GUID database,
Moq, and a fake summarizer. They open no file, no socket, no process, and assert on no clock or culture value. - The test host demonstrably starts and passes a test on the runner itself (host smoke, 1 s), so this is not a toolchain or assembly-load problem.
- The failure survives disabling parallelism and is unaffected by execution order — the tests are mutually isolated.
- Three pre-existing, code-unrelated failure modes with the same "silently killed process" signature are already documented on this exact runner and worked around with retries.
B — Deploy job fails before doing any work
Evidence
- Step
Run remote deploy(theappleboy/ssh-action) failed in 3 s (18:44:07 → 18:44:10). That is beforegit fetch, beforedeploy.sh, before any Docker build. - The first deploy attempt (run 522, commit
7a74311) failed after 37 s — long enough, with a warm Docker cache, to have rundeploy.shand failed its post-deploy backend health check. That is consistent with the MariaDB migration crash fixed in1430313. - Every attempt since fails at 3–7 s, i.e. at connection time. The failure mode changed.
- Production is up and healthy:
https://jobs.cesnimda.uk/→ 200 HTML,/api/auth/config→ 200 JSON ({"requireAuth":true,...}). The host is reachable from the internet, so this is not an outage. - Production is stale — Phase 4 and Phase 5 have never deployed. Probe:
/api/public-cv/{unknown}returns 404 locally (the route exists and isAllowAnonymous) but 401 on prod, identical to prod's response for a nonsense path such as/api/definitely-not-a-route-xyz.PublicCvControlleris absent from production.
Consequence worth recording
Because no Phase 4/5 deploy ever succeeded, production never executed the faulty migration. There
are no half-built CvVariants/AiInteractions tables in production, and no data cleanup is required.
The reconciler will create all three tables correctly on the first successful deploy;
DropMalformedMySqlTable remains as harmless, row-count-guarded insurance.
Hypothesis
The prod host is refusing the runner's SSH connection rather than failing inside the script. Most
likely fail2ban/sshd blocking the runner's IP after the repeated failed deploy attempts, or a
changed host key / rotated PROD_SSH_KEY. Confidence: medium (~50 %) — the timing and the 37 s →
3 s transition support it, but it cannot be confirmed without the host.
Required infrastructure changes
Blocking — cannot proceed without one of these:
- The job log for step
T AiWorkspace(run 528) — roughly 20 lines settles issue A outright. Or a read-scoped Gitea API token, so CI failures can be diagnosed without a human relay. This is the single highest-value item. - The deploy job log (run 523+) — the
ssh-actionerror line settles issue B.
Host checks (issue A):
dmesg -T | grep -iE 'oom|killed process'around the run time — a killeddotnet/testhostconfirms the OOM hypothesis.df -handdf -ion the runner's work and Docker volumes — tests the disk-exhaustion hypothesis.ulimit -aandcat /sys/fs/cgroup/pids.maxfor the runner user — tests the PID-limit hypothesis.journalctl -u <gitea-runner-service> --since '2 hours ago'.
Host checks (issue B):
fail2ban-client status sshdon the prod host, andjournalctl -u sshd --since '2 hours ago' | grep -i <runner-ip>.- Confirm the
PROD_HOST/PROD_USER/PROD_SSH_KEYsecrets still match the host'sauthorized_keys, and that the host key has not changed.
Recommended remediation regardless of which hypothesis lands:
- Give the runner its own resource allocation, or move it off the production host. It currently
appears to share a box with the prod Docker stack. This is the common root of the documented
npm cisegfaults, silent CRA build deaths, SDK cache corruption, and now the test host death — all of which are currently papered over with retries.
Repository state
Kept — all correct independent of the outcome, none reverted:
1430313— CV builder / AI workspace tables provisioned by the MySQL-safe reconciler (reproduced and verified against a real MariaDB 11 container).cfba7fb— CI actually runs the backend suite.45725ac— restore / build / test split, restore retries once.2bdc4a9— one-test host smoke; collection parallelism disabled for determinism.7fa3080,0f62dc4— bisection scaffolding, since removed.
No test was weakened, skipped, filtered, or disabled at any point. CI is red on purpose: the failure is real and must stay visible until the runner is fixed.