diff --git a/.gitea/workflows/ci-deploy.yml b/.gitea/workflows/ci-deploy.yml index 03f252b..c3aa4bd 100644 --- a/.gitea/workflows/ci-deploy.yml +++ b/.gitea/workflows/ci-deploy.yml @@ -62,30 +62,10 @@ jobs: - name: Test backend (host smoke) run: dotnet test JobTrackerApi.Tests/JobTrackerApi.Tests.csproj --configuration Release --no-build --filter "FullyQualifiedName~CvBuilderTests.Every_catalog_theme_renders_valid_html" - # Runner-only failure, localised to the AiWorkspace tests by bisecting across steps (the job - # log is not readable via the Gitea API). Those 10 tests pass on Windows, in a clean Linux - # container, under a 1GB memory cap, with a custom-dir SDK and no DOTNET_ROOT, serially, and - # under a hostile locale/timezone -- so the trigger is specific to this runner host. - # See docs/infrastructure/runner-investigation.md. Collection parallelism stays off for determinism, the - # same reason the frontend runs --runInBand. Every test still runs; nothing is filtered. - # Post-reboot the failure is identical (smoke 1s pass, suite 3s fail), so it is deterministic, - # not an OOM/flake -- and host telemetry shows no disk/memory/PID pressure. The passing smoke is - # a PURE test (no EF, no Moq); every class known to fail uses EF InMemory + Moq, whose Castle - # DynamicProxy emits IL at runtime. A container seccomp/W^X policy that blocks runtime code - # generation would kill the host exactly this way. These ordered steps test that: pure classes - # first, then a Moq+EF class. Diagnostic scaffolding; no test is skipped or weakened. - - name: T pure JobPipeline - run: dotnet test JobTrackerApi.Tests/JobTrackerApi.Tests.csproj --configuration Release --no-build --filter "FullyQualifiedName~JobPipelineTests" - - - name: T pure EmailStatusClassifier - run: dotnet test JobTrackerApi.Tests/JobTrackerApi.Tests.csproj --configuration Release --no-build --filter "FullyQualifiedName~EmailStatusClassifierTests" - - - name: T pure JobCvMatchService - run: dotnet test JobTrackerApi.Tests/JobTrackerApi.Tests.csproj --configuration Release --no-build --filter "FullyQualifiedName~JobCvMatchServiceTests" - - - name: T MoqEF CvBuilder - run: dotnet test JobTrackerApi.Tests/JobTrackerApi.Tests.csproj --configuration Release --no-build --filter "FullyQualifiedName~CvBuilderTests" - + # This runner fails at a DIFFERENT stage on different runs with no readable diagnostics + # (test host at 3s; on another run `dotnet restore` at 0s). See + # docs/infrastructure/runner-investigation.md. Collection parallelism stays off for + # determinism, the same reason the frontend runs --runInBand. Every test runs; nothing skipped. - name: Test backend run: dotnet test JobTrackerApi.Tests/JobTrackerApi.Tests.csproj --configuration Release --no-build -- xUnit.parallelizeTestCollections=false xUnit.maxParallelThreads=1 diff --git a/docs/infrastructure/runner-investigation.md b/docs/infrastructure/runner-investigation.md index 0e9cca4..a48f8a1 100644 --- a/docs/infrastructure/runner-investigation.md +++ b/docs/infrastructure/runner-investigation.md @@ -103,15 +103,41 @@ measured yet — see the asks below. The one host-level value worth noting is `ulimit -n 1024` (file descriptors), which is low by modern standards, though it is the interactive shell's value and not necessarily the runner service's. +### Post-reboot runs — the failure moves between stages + +Two further runs after the machine was restarted: + +| Run | Result | +|---|---| +| `8f73548` (post-reboot) | **Identical** to pre-reboot: host smoke ✓ 1 s, full suite ✗ 3 s | +| `c4c0cd4` | Failure **moved earlier**: `Restore backend tests` ✗ **0 s** — a step that took 3–4 s and succeeded in all previous runs. None of the queued diagnostic steps executed. | + +Two conclusions: + +1. **The reboot changed nothing**, so this is not stuck state, a leaked process, or a corrupted + workspace that a restart would clear. +2. **The failing stage is not stable across runs.** `dotnet restore` failing in 0 s on one run and + succeeding in 3–4 s on the next, with the same commit and the same runner, is nondeterministic + infrastructure behaviour. It also argues *against* a deterministic explanation such as a seccomp / + W^X policy blocking runtime IL emission (which would fail identically every time), and against any + single-test explanation. + +Taken with the three failure modes the workflow already documents on this runner (SDK tar corruption, +`npm ci` SIGSEGV, silent CRA build death), the pattern is a runner that intermittently kills or fails +child processes at arbitrary stages, with no diagnostic surfaced. + ### Confidence level - **Application code is not the cause — high confidence (~95 %).** Ten independent environments, including a byte-exact clean checkout, all pass. Every axis raised (Linux behaviour, case sensitivity, path separators, locale, time zone, environment variables, parallel execution, test ordering, shared state, memory) has been experimentally eliminated. -- **Specific mechanism (OOM vs pids vs disk) — low/medium confidence (~40 %).** Not reproducible - remotely, and not resolvable without the job log or host access. I am deliberately not asserting - which one it is. +- **Specific mechanism — low confidence (~25 %), and lower than before.** Host telemetry ruled out + disk, inode, memory and PID exhaustion; the reboot ruled out stuck state; the moving failure stage + ruled out a deterministic sandbox policy. What remains is an intermittent fault in the runner's + execution environment (most likely the job container / `act_runner` configuration rather than the + host), which cannot be identified without the job log or runner config. I am not asserting a + mechanism. ### Why application code is no longer suspected