litellm: add local_qwen2.5-14b, the structured-output tier
paperless-ai needs a model that holds strict JSON without being asked, because it does not send response_format. llama3.1:8b breaks the shape and qwen2.5:72b cannot answer inside its client timeout. 14B clears both. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
dd8589a20b
commit
3b1de05760
@@ -94,6 +94,16 @@ litellm_models:
|
||||
backend: ollama_chat/llama3.1:8b
|
||||
api_base: "{{ anvil_ollama_base }}"
|
||||
api_key: ollama-no-auth
|
||||
# The structured-output tier, added 2026-09-12. llama3.1:8b answers fast but
|
||||
# will not hold strict JSON: paperless-ai got "Expected ',' or '}' after
|
||||
# property value in JSON at position 263" on a 1,921-token reply. The Kestra
|
||||
# flows escape this by setting response_format={"type":"json_object"};
|
||||
# paperless-ai does not, so it needs a model that keeps the shape unprompted.
|
||||
# 14B is the smallest that does, and at ~9 GB it co-exists with anything here.
|
||||
- name: local_qwen2.5-14b
|
||||
backend: ollama_chat/qwen2.5:14b
|
||||
api_base: "{{ anvil_ollama_base }}"
|
||||
api_key: ollama-no-auth
|
||||
# Shadow alias SOLELY for paperless-ai's hardcoded gpt-4o-mini validation
|
||||
# probe (see the "Narrow exception" note above). Routed to free local
|
||||
# llama3.1:8b so the probe costs nothing. Do not point real traffic here.
|
||||
|
||||
Reference in New Issue
Block a user