feat(ai): provider router (ollama|gemini|groq) for heavy CV calls
The structured /cv/* calls funnel through a provider router so production can offload a weak local GPU (GTX 1060) to a cloud provider without any .NET change. Default stays "ollama" (keyless/local) and /summarize remains local distilbart. - AI_PROVIDER=ollama|gemini|groq dispatch inside _ollama_generate_json/_text (entry-point names kept, so no call sites change; Ollama path is byte-identical). - Gemini (x-goog-api-key header, not URL query) and Groq (OpenAI-compatible chat/completions) added via stdlib urllib — zero new dependencies. - /health reports ai_provider + ai_provider_configured. - Keys read from env only; never logged/committed. - Compose + .env.example pass AI_PROVIDER/GEMINI_*/GROQ_* through. Tests: 11 passed (default Ollama unchanged, Gemini/Groq dispatch, missing-key 503, health reports provider). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -13,6 +13,16 @@ AI_SERVICE_BASE_URL=http://ai-service:8001
|
||||
OLLAMA_BASE_URL=http://ollama:11434
|
||||
OLLAMA_MODEL=qwen2.5:7b
|
||||
|
||||
# AI provider for the heavy /cv/* calls: ollama (default, local) | gemini | groq.
|
||||
# /summarize always stays local (distilbart). To offload a weak production GPU,
|
||||
# set AI_PROVIDER=gemini (or groq) and provide the matching key below.
|
||||
# Keys are read from the environment only — never commit real keys.
|
||||
AI_PROVIDER=ollama
|
||||
GEMINI_API_KEY=
|
||||
GEMINI_MODEL=gemini-2.0-flash
|
||||
GROQ_API_KEY=
|
||||
GROQ_MODEL=llama-3.3-70b-versatile
|
||||
|
||||
# Optional: only needed if you want the UI to call a non-default API base URL.
|
||||
# In production the UI defaults to `/api`.
|
||||
REACT_APP_API_BASE_URL=
|
||||
|
||||
Reference in New Issue
Block a user