port 11434 · bearer auth · LAN only

An OpenAI-compatible LLM server that runs on your iPhone.

On Device: LAS turns a phone into a local inference endpoint. Point aider, Continue, or anything that speaks the OpenAI API at it and the tokens come off the device in your pocket — no account, no cloud round trip, no telemetry.

bash — streaming from 192.168.1.8
LOCAL
ON DEVICE : LAS Local-first LLM runtime for iPhone

Architecture

SwiftUI interface
ModelResidency
MemoryAdvisor
MLX + llama.cpp
LocalAPIServer
RuntimeLogCenter

Capabilities

On-device inference
Model import + export
Resumable downloads
OpenAI / Anthropic / Ollama
Tool calling + parsers
Debugger + live terminal

Request flow

Client
Local API
Runtime gate
Resident model
Bearer-authenticated · LAN only
>_ iOS 18+· No telemetry· No silent cloud fallback· MIT source
What it is

A server build, not a chat app

On Device: LAS is the server-only build of the local runtime. It keeps model loading, memory and thermal safety, authentication, and the API compatibility layer — and drops the assistant, lens, voice, and share-extension surfaces entirely. The app's job is to hold a model resident and answer HTTP.

🔌

Three dialects, one server

OpenAI, Anthropic Messages, and Ollama routes are served from the same port. Clients connect with the SDK they already use.

⚡

Token-level streaming

SSE frames are written and flushed per token, so a client paints text as the model decodes it rather than waiting for the generation to finish.

🧠

Reasoning kept separate

Chain-of-thought streams as reasoning_content while content carries only the final answer. Split markers across token boundaries are handled.

🛠

Native tool calls

Provider-native shapes with preserved call IDs, structured arguments, and parallel calls. On MLX, decoding is grammar-constrained rather than prompt-hoped.

📦

MLX and GGUF

Load an MLX model or import a GGUF for the llama.cpp runtime. The repository ships no weights — you bring the model.

🔒

Closed by default

Bearer-authenticated HTTP intended for a trusted LAN. The listener stops when iOS backgrounds the app. Nothing leaves the device.

Previews

On the device

The server build is a control surface, not a chat app — a homepage that holds a model resident and answers HTTP, with runtime state, memory admission, and logs in view.

On Device: LAS server homepage showing device health, memory headroom, and Local API server status
Server homepageDevice health, memory headroom, and the Local API server toggle.
Model library screen listing installed on-device models with capacity readouts
Model libraryImport, download, and make an MLX or GGUF model resident.
On-device debugger showing the memory admission snapshot
On-device debuggerAdmission snapshot — process budget, entitlements, thermal state.
Verbose terminal streaming live server log events
Verbose terminalLive server log — every request and lifecycle event.
Compatibility

Endpoints

Compatibility is intentionally scoped. Options the local runtime does not implement are ignored where that is safe and rejected explicitly where it is not — the server does not pretend to support what it cannot do.

OpenAI

aider · Continue · openai-python

  • GET/v1/models
  • GET/v1/models/{model}
  • POST/v1/chat/completions
  • POST/v1/responses

Anthropic

anthropic-sdk · Messages API

  • POST/v1/messages

Ollama

Ollama-native clients

  • GET/api/tags
  • GET/api/ps
  • GET/api/version
  • POST/api/show
  • POST/api/chat
  • POST/api/generate
Requirements

What a coding agent needs

Agent clients like aider drive long tool-calling conversations, and that is the demanding case — not chat. A model has to carry a real 65,536-token context and a KV cache to match before the server will treat it as agent-ready; below that the app shows a Hermes-compatibility warning rather than failing mid-session.

Model familyContextTool parser Parallel callsMax outputKV cache
Qwen3
needs YaRN ×2
65,536hermes2 4,0961.1–4.3 GB
Qwen3.5 65,536qwen3_coder2 4,096~1.1 GB
Ornith 65,536qwen3_xml2 4,096~1.1 GB
Bonsai
not agent-ready
≤ 32,768qwen3_coder1 2,048≤ 0.5 GB
Other / imported
no tool support
catalognone1 ≤ 4,096varies

Why 64K, not 32K

Qwen3 is natively 32K. Reaching a real 65,536-token runtime needs a YaRN factor of 2, which the app applies to the model configuration before MLX decodes it — your original config is backed up, not overwritten.

Why the KV cache matters

A 64K cache is 1.1 GB on a small Qwen3 and about 4.3 GB on an 8B, on top of the weights. This is the reason the memory entitlements below are not optional.

Parallel tool calls

Capped at 2 concurrent calls on the families that handle them reliably, and forced sequential elsewhere. The limit is also settable per install.

Quickstart

Talking to your phone

The app shows its LAN URL and bearer key on the server homepage. Rotate the key after installing, then point a client at it.

01 Check the model is resident
$ curl -s http://192.168.1.8:11434/v1/models \
    -H "Authorization: Bearer $KEY" | jq .data[0].id
02 Stream a completion
$ curl -N http://192.168.1.8:11434/v1/chat/completions \
    -H "Authorization: Bearer $KEY" \
    -H 'Content-Type: application/json' \
    -d '{"model":"your-model-id","stream":true,
         "messages":[{"role":"user","content":"Explain SSE in one line."}]}'
03 Point aider at it
$ export OPENAI_API_BASE=http://192.168.1.8:11434/v1
$ export OPENAI_API_KEY=<key from the app>
$ aider --model openai/your-model-id
Availability

Coming soon to the App Store

Public sideload downloads have been retired. Official App Store links will be published at ondevice.fun when the apps are available.

Source code and historical release notes remain available on GitHub. Models are downloaded separately after installation.

Boundaries

What it will and will not do

Trusted LAN only

Bearer-authenticated plain HTTP. It is designed for a network you control, not for exposure to the internet. Do not port-forward it.

Foreground only

The listener stops when iOS backgrounds the app. A request that arrives before a model is resident gets a service-unavailable response, not a hang.

No silent execution

The server returns tool calls; it does not run them. Executing a call is always the client's decision.

Independent project. On Device: LAS is community-maintained and is not affiliated with, endorsed by, or supported by Apple Inc. Source code remains public; App Store releases are coming soon.