Portable builds — Download and run on Windows, macOS, or Linux without additional setup.
Multiple backends — Switch between llama.cpp, ExLlamaV3, Transformers, and TensorRT-LLM without restarting.
Tool-calling — Models can execute custom Python functions including web search and math tools.
Vision and file attachments — Send images, PDFs, and documents to understand their contents.
OpenAI-compatible API — Use as a drop-in replacement for OpenAI and Anthropic APIs with tool support.
100% offline local LLM interface with zero telemetry. Portable builds get you running in under a minute with no Python setup. Supports vision, tool-calling, fine-tuning, image generation, and works with multiple inference backends—swap between them mid-session.
Python 3.9+ for manual install; portable builds are self-contained. 10GB disk space for full feature set.