Files
jobtrackingapp/tools/summarizer
cesnimda e508fb9541
CI and Deploy / test (push) Failing after 52s
CI and Deploy / deploy (push) Has been skipped
test(parser): treat reaped container zombies as terminated
2026-08-31 22:14:37 +02:00
..
2026-08-30 11:12:25 +02:00
2026-08-30 11:12:25 +02:00
2026-08-30 11:12:25 +02:00
2026-03-21 11:55:27 +01:00
2026-03-21 11:55:27 +01:00

Local AI Service

This service runs a local Hugging Face summarization model and also exposes document text extraction with OCR for supported PDFs and images.

Capabilities

  • job/role summarization
  • PDF text extraction
  • OCR fallback for scanned PDFs
  • OCR for image uploads (png, jpg, jpeg, webp)
  • DOCX / TXT / MD extraction
  • optional Ollama-backed CV block classification for harder sectioning

Install

Windows:

python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements-dev.txt
python -m uvicorn app:app --host 127.0.0.1 --port 8001 --workers 1

Linux / macOS:

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-dev.txt
python -m uvicorn app:app --host 127.0.0.1 --port 8001 --workers 1

If the host is missing python3-venv or pip, use the bootstrap script instead:

./scripts/bootstrap-and-test.sh bootstrap

Docker

The Dockerfile installs Tesseract OCR so scanned PDFs and supported images can be processed inside the container.

Document parser isolation

/extract-text performs only bounded upload and signature checks in the web process. Actual TXT, Markdown, PDF, DOCX and image decoding runs in a single-capacity child process with a minimal environment. Linux children receive CPU, address-space, output-size and file-descriptor limits; every platform enforces a parent deadline and terminates the parser process tree on timeout.

Production Compose runs the service as a non-root user with a read-only root filesystem, no Linux capabilities, no-new-privileges, PID/CPU/memory limits and a bounded /tmp tmpfs. The Hugging Face cache is the only named writable volume. Parser work directories are removed after each request and stale work-* directories are reconciled on the next extraction.

Parser limits can be made more conservative with PARSER_TIMEOUT_SECONDS, PARSER_CPU_SECONDS, PARSER_MEMORY_MIB and PARSER_STALE_WORK_SECONDS. Increasing them should follow measured benign-CV evidence; disabling import is the safe fallback when the budget is insufficient.

Tests

Run the summarizer unit tests with:

./scripts/bootstrap-and-test.sh test

The script:

  • creates .venv with stdlib venv when available
  • falls back to user-space virtualenv when host venv support is missing
  • installs requirements-dev.txt
  • writes pytest cache under tmp/pytest-cache to avoid stale root-owned .pytest_cache directories

API

  • GET /health — health check and runtime capabilities, including lazy model state (model_loaded, model_disabled, summarize_available, model_load_error) plus Ollama version/model metadata when configured
  • POST /summarize — JSON body { "text": "...", "max_length": 150, "min_length": 30 }
  • POST /extract-text — multipart file upload, returns extracted text and OCR metadata
  • POST /cv/classify-block — JSON body { "block": "..." }, uses Ollama when OLLAMA_MODEL is configured

Ollama

Set these before starting the service if you want the hybrid CV classifier enabled:

export OLLAMA_BASE_URL=http://ollama:11434
export OLLAMA_MODEL=qwen2.5:7b

Choose the model by setting OLLAMA_MODEL and then warming it with the helper script:

OLLAMA_MODEL=qwen2.5:7b ./scripts/start-ollama-cv.sh

Equivalent manual flow:

docker compose up -d ollama
docker compose exec ollama ollama pull qwen2.5:7b
docker compose up -d ai-service
  • Model weights are downloaded on first pull.
  • OCR quality depends on scan quality and language support.
  • Default OCR language is English (eng).