I built an OpenAI-compatible server that runs Gemma 4 E4B on a Pixel 10 Pro XL (Tensor G5) — fully offline, ~11 tok/s decode, Tailscale-encrypted option, Apache 2.0

Wait 5 sec.

What it is: an Android app (PixelUnlockGPU) that turns a Pixel into an OpenAI-compatible HTTP server. Standard `/v1/chat/completions` with streaming, so it talks to TypingMind or any OpenAI client directly — no cloud, no subscription, model runs entirely on-device via LiteRT-LM. Device: Pixel 10 Pro XL (Tensor G5, 16 GB shared LPDDR). Model: Gemma 4 E4B instruct, GPU bundle, 2.97 GB, SHA-256 verified on download. Context window capped at 32k. Measured numbers (not benchmarks — real on-device measurements): - ~11 tok/s steady-state decode (first-token-to-last over a ~300-word generation) - Follow-up turns in ~1.7 s: the server auto-reuses the KV prefix across turns, so stateless clients like TypingMind don't re-prefill history - Engine warm build ~12 s once per model change (visible in-app, split out of the metrics on purpose) - Short replies read slower than 11 tok/s because warm + prefill dominate the window — the UI separates decode tok/s from prefill ms so nobody has to guess Security/access: three independent modes — loopback only, raw LAN (unencrypted), or Tailscale (binds the CGNAT tailnet IP; WireGuard end-to-end from e.g. a laptop on the same tailnet; degrades to loopback if the VPN drops). Verified with a real chat completion over the tailnet from a MacBook. Honest limitations: - NPU path aborts on stock Tensor G5 firmware — GPU is the shipping backend (documented with the full investigation) - Android blocks named GPU temp sensors for normal apps, so the in-app gauge shows OS thermal headroom instead of °C - No real token counts anywhere — LiteRT-LM exposes none, so usage is estimated at ~4 chars/token and labeled as such - `stop` sequences rejected explicitly (engine has no per-request stop API); client-side emulation is a filed issue Why not llama.cpp/ollama on the phone: wanted the official LiteRT-LM GPU path on Tensor specifically, an always-on Android service (Ktor/Netty), and the OpenAI wire so existing clients just work. Happy to add a GGML backend if people want it — the engine layer is abstracted. Built on, with credit: server/inference foundation derived from mlnomadpy/localllm (Apache 2.0); Tensor G5 runtime knowledge and the prebuilt dispatch lib from jegly/Box (Apache 2.0). Full attribution in NOTICE + per-file headers. Apache 2.0, contributions welcome — there are labeled good-first-issues (usage block, stop-sequence emulation, docs). Repo + APKs (v0.1.0/v0.1.1 on releases): https://github.com/cannitellinicholas-spec/PixelUnlockGPU Happy to answer anything about Tensor G5 quirks — I've done more Gate-2 debugging than I planned to.   submitted by   /u/VerityAISolutions [link]   [comments]