High-performance OpenAI-compatible embedding server on Apple Silicon. Runs Qwen3 embedding models locally via MLX with optimized batched GPU inference — no API keys needed. Up to 5x faster than Ollama for the same models.

WWW: https://github.com/dmarkey/roocode-code-indexer-macos
