Running Ollama Right: Local Models and Your Own Fine-Tunes on One GPU
Ollama is the least painful way to serve local models on a single GPU, but the defaults are tuned for demos, not for work. What actually matters (quantization, context memory, keep-alive) and how to get a fine-tuned model of your own running behind the same API.