Files
infrastructure/docs
Kaloyan Danchev 2a34c42270
ci/woodpecker/push/woodpecker Pipeline was successful
docs(services): correct Ollama iGPU entry — remove flash-attn/q8_0, note large-prompt crash
The flash-attention + q8_0 KV cache settings were removed after they (and the
gfx1103 ROCm path generally) crash on large prompts. Document that the 780M iGPU
is direct/interactive-chat only, not an agent backend (Hermes ~15.5k-token prompt
crashed every call; reverted to cloud model).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 22:14:41 +03:00
..
2026-01-31 13:06:02 +02:00