eac34705e3
Unblocks the documented core workflow and closes the AI-service exposure, without changing existing behaviour. Job/JobApplication split (additive; see ADR-002): - New Job entity (the opportunity) with owner-scoped query filter; nullable JobApplication.JobId FK. Nothing reads Job yet. - Migration AddJobEntityAndProspectStages, hand-edited to drop reconciler-owned tables the scaffolder re-emitted; verified against the real dev DB. Pipeline: 10 internal stages across three concerns kept separate — PipelineStage (workflow) / PipelineGroup (UI: NotApplied/Active/Closed) / PipelineCategory (analytics). Adds Saved/Interested/Preparing/Withdrawn; keeps Waiting and Ghosted. Kanban shows 3 grouped columns; cards keep a stage chip and full transitions; drag applies only safe transitions (never infers Ghosted/Withdrawn). DateApplied nullable + SavedAt. Cleared when leaving Applied so analytics stay accurate; the discarded date is preserved as an AppliedDateCleared JobEvent. AI service lockdown: no host port; private ai_internal network (backend is the only other member); X-Ai-Service-Token required on all non-/health endpoints; AI_SERVICE_TOKEN mandatory via compose. Verified backend-only against the live stack. Also carries two pre-existing working-tree files (views/ProfilePage.tsx, views/CareerWorkspacePage.tsx) so the tree is clean for the branch integration. Tests: +40 backend (247 total), +5 sidecar (16), +15 frontend. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Local AI Service
This service runs a local Hugging Face summarization model and also exposes document text extraction with OCR for supported PDFs and images.
Capabilities
- job/role summarization
- PDF text extraction
- OCR fallback for scanned PDFs
- OCR for image uploads (
png,jpg,jpeg,webp) - DOCX / TXT / MD extraction
- optional Ollama-backed CV block classification for harder sectioning
Install
Windows:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
python -m uvicorn app:app --host 127.0.0.1 --port 8001 --workers 1
Linux / macOS:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python -m uvicorn app:app --host 127.0.0.1 --port 8001 --workers 1
If the host is missing python3-venv or pip, use the bootstrap script instead:
./scripts/bootstrap-and-test.sh bootstrap
Docker
The Dockerfile installs Tesseract OCR so scanned PDFs and supported images can be processed inside the container.
Tests
Run the summarizer unit tests with:
./scripts/bootstrap-and-test.sh test
The script:
- creates
.venvwith stdlibvenvwhen available - falls back to user-space
virtualenvwhen hostvenvsupport is missing - installs
requirements-dev.txt - writes pytest cache under
tmp/pytest-cacheto avoid stale root-owned.pytest_cachedirectories
API
GET /health— health check and runtime capabilities, including lazy model state (model_loaded,model_disabled,summarize_available,model_load_error) plus Ollama version/model metadata when configuredPOST /summarize— JSON body{ "text": "...", "max_length": 150, "min_length": 30 }POST /extract-text— multipart file upload, returns extracted text and OCR metadataPOST /cv/classify-block— JSON body{ "block": "..." }, uses Ollama whenOLLAMA_MODELis configured
Ollama
Set these before starting the service if you want the hybrid CV classifier enabled:
export OLLAMA_BASE_URL=http://ollama:11434
export OLLAMA_MODEL=qwen2.5:7b
Choose the model by setting OLLAMA_MODEL and then warming it with the helper script:
OLLAMA_MODEL=qwen2.5:7b ./scripts/start-ollama-cv.sh
Equivalent manual flow:
docker compose up -d ollama
docker compose exec ollama ollama pull qwen2.5:7b
docker compose up -d ai-service
- Model weights are downloaded on first pull.
- OCR quality depends on scan quality and language support.
- Default OCR language is English (
eng).