OlliteRT social preview
inference-server309Apache 2.0

OlliteRT

Turn your Android phone into an OpenAI-compatible LLM inference server - Fully local, private and Open Source

Updated Sep 10, 2026
Platforms
Pricing
free-open-source
Status
active
License
Apache 2.0

What it does

Core capabilities at a glance

  • Android
  • Anthropic API
  • Gemma
  • Home Assistant
  • Kotlin Android
  • Litert
  • Litert LM
  • LLM Inference

Deep dive

The full breakdown - performance, comparisons, and setup

OlliteRT

OlliteRT is a local inference server - Turn your Android phone into an OpenAI-compatible LLM inference server - Fully local, private and Open Source.

Overview

Think of it as Ollama for Android. Pick a model, tap Start, and your phone becomes an LLM server — runs LLMs on your mobile GPU/CPU via Google's LiteRT-LM runtime and serves them as a standard OpenAI-compatible HTTP API on your local network.

No cloud. No API keys. No subscriptions. Just your phone.

OlliteRT is open-source, written primarily in Kotlin, with 309 GitHub stars under the Apache 2.0 license. It was last updated on 2026-09-05.

Key capabilities

From the project's documentation:

  • Benchmark Built-in — Test and compare models on your device to find the best fit for your hardware
  • Activity Logs — Detailed request/response logs with search, filtering, and JSON highlighting
  • Broad Compatibility — Home Assistant, Open WebUI, OpenClaw, Python, curl — if it talks to OpenAI, it works
  • FAQ — Model support, privacy, battery, architecture, tool calling
  • Troubleshooting — Connection issues, performance, crashes, auto-start, storage
  • Privacy Policy — no data is collected, no telemetry, no analytics

How it fits a local-AI stack

OlliteRT runs on your own hardware, so pair it with a model and a GPU sized to your needs. Use the VRAM calculator to pick a model that fits your card, and see what you can run for hardware guidance. Related local inference servers in the directory:

Sources

Stats from GitHub, 2026-09-10.

Frequently asked

Quick answers to common questions

What is OlliteRT?

OlliteRT is a inference-server tool for local AI workloads. Turn your Android phone into an OpenAI-compatible LLM inference server - Fully local, private and Open Source

Is OlliteRT free and open source?

Yes, OlliteRT has 309 GitHub stars and is licensed under Apache 2.0. You can self-host it for free on .

What hardware do I need for OlliteRT?

The hardware requirements depend on which models you run. Check our hardware directory for compatible GPUs and systems. OlliteRT has 309 GitHub stars and an active community.

Does OlliteRT support GPU acceleration?

OlliteRT's GPU support depends on your specific setup. Check the documentation for details. For the best performance, pair it with an NVIDIA RTX 4090 or 5090.

What are the best alternatives to OlliteRT?

Popular alternatives include other inference-server tools in our directory. Browse our full collection at /tool for comparisons, community reviews, and benchmark data to find the right fit for your workflow.

How much does OlliteRT cost?

OlliteRT is free-open-source. It is completely free and open source to self-host.

Pairs well with

Complementary tools, models, and hardware

Comments coming soon

Configure NEXT_PUBLIC_GISCUS_REPO_ID and NEXT_PUBLIC_GISCUS_CATEGORY_ID at giscus.app to enable.