

I’m not the OP, but I’m trying to figure out if local models are always painfully slow or if I’m missing something obvious in my tuning.
Stable Diffusion can whip up a picture in less time on the same hardware, than lama.cpp takes to decide to call an MCP function.
It seems like I must be missing something in my lama.cpp setup, but none of the guides I’ve read have clued me in to what I’ve done wrong.
Ollama performs similarly poorly on the same harsware, so I’ve probably managed to make the se mistake(s) at least twice.
Anyway, that’s the main thing I’m reading along for. Trying to increase my understanding until I catch my own mistakes.













I will study these. Thank you!