Hey everyone, Most mobile LLM setups for iOS suffer from two issues: Web UIs in mobile Safari tend to drop streaming the second your screen locks or you switch apps, with zero access to native iOS APIs. Most App Store clients push aggressive $15/month subscriptions and route your private prompts through their own cloud proxies. I built Eron as a clean, native iOS companion specifically for people running their own local hardware (Ollama, vLLM, LM Studio) or using their own API keys (BYOK). Technical details & v1.4 architecture: Direct Socket / Zero Proxy: Direct HTTP/WebSocket connection straight to your local IP or Tailscale/WireGuard node. No intermediate servers, no telemetry, no account required. Zero-Buffer Streaming: Rewrote the streaming pipeline from scratch. Instead of waiting for sentence buffers, tokens render as raw chunks as fast as your GPU outputs them. Reasoning Stream: Native streaming and collapsible rendering for reasoning blocks (DeepSeek R1, Qwen reasoning, etc.). Local iOS Tool Calling: If your local model supports function calling, Eron provides native bridges to Apple Reminders, Calendar events, and HomeKit smart home control directly from your prompt. Workspaces: Isolated project workspaces with persistent custom system prompts to keep coding contexts separate from daily chats. v1.4.1: Native dual-screen layout ready for the upcoming iPhone Duo form factor. Pricing & Community Codes: It’s a $2.99 one-time purchase on the App Store To get feedback from this community, I have 20 App Store promo codes to give away to anyone running a local setup who wants to test it for free. Just drop a comment with your setup (what models/hardware you’re running) and I’ll DM you a code! App Store: https://apps.apple.com/app/eron/id6760043923 Setup docs: https://henningwinter.com/app/eron Self-promotion disclosure: I am the sole developer.   submitted by   /u/RA2B_DIN [link]   [comments]