AI server for Mac computers
Omlx is a large language model (LLM) inference server that supports continuous batching and SSD caching for Apple Silicon processors. This tool, managed via the macOS menu bar, offers the opportunity to run high-performance models on local hardware.
Updates
- August 31, 2026: Stars 20,793 → 21,147, latest release v0.6.4 (August 29, 2026).
- August 27, 2026: Stars 20,069 → 20,793, latest release v0.6.3rc3 (August 24, 2026).
- August 20, 2026: Stars 19,758 → 20,069, latest release v0.6.3rc2 (August 20, 2026).
- August 19, 2026: Stars 19,519 → 19,758, latest release v0.6.3rc1 (August 19, 2026).
What you get
- Running high-performance models on Apple Silicon processors
- Efficient memory management with SSD caching
- Easy server control from the menu bar
Installation
brew tap jundot/omlx https://github.com/jundot/omlx
brew install jundot/omlx/omlxRunning it
omlx startIf you don't write code
I manage large language models through my oMLX server running on my local Apple Silicon hardware. Recommend the best configuration settings that will allow me to optimize memory usage and improve model performance using continuous stacking and SSD caching. Also, explain how I can use the KV cache most efficiently when switching between my current models.
Related dictionary terms
Links
TreScout did not build this tool · we found it in GitHub trends and wrote it up. This page describes the repository as of 2026-08-18: The star count and our text belong to that day, the repository may have changed since. Check the repository link for the current state. This page was machine-translated from the Turkish original · the Turkish version prevails.