Chapper: Native iOS Client for LM Studio, Ollama, and OpenAI-Compatible Local Models

Chapper is a native SwiftUI iOS client for connecting to local AI models running on LM Studio, Ollama, and any OpenAI-compatible server. The app runs entirely on-device with no cloud requirements, web views, or mandatory accounts.
Core Features
- Real-time token streaming with live inference speed display
- Full sampling controls: temperature, top-p, top-k, min-p, TFS-Z, repeat/presence/frequency penalty
- Structured output/JSON schema mode
- Markdown rendering with syntax-highlighted code blocks
Reasoning Model Support
- Collapsible thought process panel inline above each response
- Works with Qwen3, DeepSeek-R1, and any model using <think> tags
- Custom <think> tag parser for reasoning model output
Model Management
- In-app model management: browse, load, configure context length
- Flash attention support
- GPU KV-cache offload
Conversation Features
- Personas with persistent system prompts per chat
- Full-text search across all conversations + pinned chats
- Memory system that injects long-term context automatically
- Scratchpad for working notes while chatting
Output Options
- Export in 7 formats: PDF, HTML, Markdown, JSON, CSV, XML, TXT
- TTS in three modes: native iOS voices, local on-device Kokoro model (experimental), or custom TTS server
- Background playback support
Technical Implementation
- Native async streaming over SSE
- MCP tool integration for web search, file access, URL fetching
- iCloud sync (optional)
- On-device analytics dashboard
- 12 language support
- Custom haptics with toggle option
Pricing & Availability
Free + Pro model with one-time purchase, no subscription. Core chat is free. Pro unlocks advanced sampling, unlimited history, all export formats, custom icons, and unlimited personas. Works on iPhone and iPad.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Claude Code's Tool API Details Revealed
A Reddit user extracted details about Claude Code's tool API, including file system operations, bash execution, web search, and how tool calls are structured using XML-like blocks.

Bio-Inspired Memory System for Local LLMs: LTP and Selective Oblivion Implementation
A developer built a local MCP server implementing bio-inspired memory mechanics including Long-Term Potentiation reinforcement, selective oblivion decay, and weekly consolidation cycles. The system uses hybrid search with sqlite-vec and text fallbacks, non-blocking architecture with asyncio executors, and maintains state via a persistent 'Soul' file.

Browser Harness: Giving LLMs raw CDP access to self-correct browser tasks
Browser Harness strips away browser frameworks, giving LLMs direct CDP websocket access and letting them write missing tools mid-task. Demonstrated by self-inventing an upload_file() function.

Self-Evolving Skill pattern validation: 5-round experiment results
A developer tested the Self-Evolving Skill design pattern for Claude Code with a 5-round experiment on a MySQL database with 29 tables and 590MB of smart building management data. Key results include a 63.6% Five-Gate rejection rate, incremental convergence, and 100% accuracy with no incorrect knowledge surviving.