824251d328
The structured /cv/* calls funnel through a provider router so production can offload a weak local GPU (GTX 1060) to a cloud provider without any .NET change. Default stays "ollama" (keyless/local) and /summarize remains local distilbart. - AI_PROVIDER=ollama|gemini|groq dispatch inside _ollama_generate_json/_text (entry-point names kept, so no call sites change; Ollama path is byte-identical). - Gemini (x-goog-api-key header, not URL query) and Groq (OpenAI-compatible chat/completions) added via stdlib urllib — zero new dependencies. - /health reports ai_provider + ai_provider_configured. - Keys read from env only; never logged/committed. - Compose + .env.example pass AI_PROVIDER/GEMINI_*/GROQ_* through. Tests: 11 passed (default Ollama unchanged, Gemini/Groq dispatch, missing-key 503, health reports provider). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Local AI Service
This service runs a local Hugging Face summarization model and also exposes document text extraction with OCR for supported PDFs and images.
Capabilities
- job/role summarization
- PDF text extraction
- OCR fallback for scanned PDFs
- OCR for image uploads (
png,jpg,jpeg,webp) - DOCX / TXT / MD extraction
- optional Ollama-backed CV block classification for harder sectioning
Install
Windows:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
python -m uvicorn app:app --host 127.0.0.1 --port 8001 --workers 1
Linux / macOS:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python -m uvicorn app:app --host 127.0.0.1 --port 8001 --workers 1
If the host is missing python3-venv or pip, use the bootstrap script instead:
./scripts/bootstrap-and-test.sh bootstrap
Docker
The Dockerfile installs Tesseract OCR so scanned PDFs and supported images can be processed inside the container.
Tests
Run the summarizer unit tests with:
./scripts/bootstrap-and-test.sh test
The script:
- creates
.venvwith stdlibvenvwhen available - falls back to user-space
virtualenvwhen hostvenvsupport is missing - installs
requirements-dev.txt - writes pytest cache under
tmp/pytest-cacheto avoid stale root-owned.pytest_cachedirectories
API
GET /health— health check and runtime capabilities, including lazy model state (model_loaded,model_disabled,summarize_available,model_load_error) plus Ollama version/model metadata when configuredPOST /summarize— JSON body{ "text": "...", "max_length": 150, "min_length": 30 }POST /extract-text— multipart file upload, returns extracted text and OCR metadataPOST /cv/classify-block— JSON body{ "block": "..." }, uses Ollama whenOLLAMA_MODELis configured
Ollama
Set these before starting the service if you want the hybrid CV classifier enabled:
export OLLAMA_BASE_URL=http://ollama:11434
export OLLAMA_MODEL=qwen2.5:7b
Choose the model by setting OLLAMA_MODEL and then warming it with the helper script:
OLLAMA_MODEL=qwen2.5:7b ./scripts/start-ollama-cv.sh
Equivalent manual flow:
docker compose up -d ollama
docker compose exec ollama ollama pull qwen2.5:7b
docker compose up -d ai-service
- Model weights are downloaded on first pull.
- OCR quality depends on scan quality and language support.
- Default OCR language is English (
eng).