olla social preview
model-router302Apache 2.0

olla

High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local a…

Updated Sep 19, 2026
Platforms
docker
Pricing
free-open-source
Status
active
License
Apache 2.0

What it does

Core capabilities at a glance

  • AMD
  • Golang
  • Intel
  • Llama CPP
  • Llamacpp
  • LLM Inference
  • LLM Proxy
  • LLM Router

Deep dive

The full breakdown - performance, comparisons, and setup

olla

olla is a model router/gateway - High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.

Overview

Olla is a high-performance, low-overhead, low-latency proxy and load balancer for managing LLM infrastructure. It intelligently routes LLM requests across self-hosted inference nodes with a wide variety of natively supported endpoints and extensible enough to support others. Olla provides model discovery and unified model catalogues within each provider, enabling seamless routing to available models on compatible endpoints.

Olla works alongside API gateways like LiteLLM or orchestration platforms like GPUStack, focusing on making your existing LLM infrastructure reliable through intelligent routing and failover. You can choose between two proxy engines: Sherpa for simplicity and maintainability or Olla for maximum performance with advanced features like circuit breakers and connection pooling.

Single CLI application and config file is all you need to go Olla!

For large GPU deployments, enterprise and data centre use, see TensorFoundry FoundryOS. For an inference control plane, consider Alloy.

olla is open-source, written primarily in Go, with 302 GitHub stars under the Apache 2.0 license. The latest release is v0.0.29 (2026-08-10).

Key capabilities

From the project's documentation:

  • 🔁 Intelligent Retry: Automatic retry on connection failures with immediate transparent endpoint failover
  • 🔧 Self-Healing: Automatic model discovery refresh when endpoints recover
  • ⚡ High Performance: Sub-millisecond endpoint selection with lock-free atomic stats
  • 🎯 LLM-Optimised: Streaming-first design with optimised timeouts for long inference
  • Home Lab: Olla → Multiple Ollama (or OpenAI Compatible - eg. vLLM) instances across your machines
  • Hybrid Cloud: Olla → Local endpoints + LiteLLM → Cloud APIs (OpenAI, Anthropic, Bedrock, etc.)

Install

A quick way to get started (always check the official docs for the latest):

go install github.com/thushan/olla@latest

How it fits a local-AI stack

olla runs on your own hardware, so pair it with a model and a GPU sized to your needs. Use the VRAM calculator to pick a model that fits your card, and see what you can run for hardware guidance. Related model router/gateways in the directory:

Sources

Stats from GitHub, 2026-09-19.

Frequently asked

Quick answers to common questions

What is olla?

olla is a model-router tool for local AI workloads. High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local a…

Is olla free and open source?

Yes, olla has 302 GitHub stars and is licensed under Apache 2.0. You can self-host it for free on docker.

What platforms does olla support?

olla runs on docker.

What hardware do I need for olla?

The hardware requirements depend on which models you run. Check our hardware directory for compatible GPUs and systems. olla has 302 GitHub stars and an active community.

Does olla support GPU acceleration?

olla's GPU support depends on your specific setup. Check the documentation for details. For the best performance, pair it with an NVIDIA RTX 4090 or 5090.

What are the best alternatives to olla?

Popular alternatives include other model-router tools in our directory. Browse our full collection at /tool for comparisons, community reviews, and benchmark data to find the right fit for your workflow.

How much does olla cost?

olla is free-open-source. It is completely free and open source to self-host.

Pairs well with

Complementary tools, models, and hardware

Comments coming soon

Configure NEXT_PUBLIC_GISCUS_REPO_ID and NEXT_PUBLIC_GISCUS_CATEGORY_ID at giscus.app to enable.