Ollama v0.32.3 Enhances Linux CUDA/ROCm iGPU Memory, Adds B200 Support
Ollama v0.32.3 was released today, bringing significant updates for local large language model inference. The update focuses on expanding GPU support and optimizing resource usage on Linux systems. Key improvements include lower memory usage for Linux CUDA and ROCm integrated GPUs (iGPUs). Additionally, the release adds support for NVIDIA's B200 GPUs through CUDA 12. The underlying MLX and llama.cpp engines were also updated to enhance inference capabilities. Other changes address model download stalls and improve integrations, such as restoring Claude Code Channels.
Sources
- v0.32.3 - GitHub: ollama/ollama