Fine-tuning llama3.2 3B for personalized health coaching using Apple Watch data and MLX

A developer created a personalized health coach LLM by fine-tuning llama3.2 3B on a Mac using Apple Health and Whoop data. The entire fine-tuning process took approximately 15 minutes using MLX.
Technical pipeline
The implementation follows this workflow:
- Apple Health and Whoop data stored in local SQLite database
- SQL RAG layer converts natural language queries to SQL
- Claude API used once to generate ~270 gold-standard training examples (anonymized question/SQL/result pairs, no personal health data sent)
- LoRA fine-tuning on llama3.2 3B via MLX
- Fused model served locally at 127.0.0.1:8080
Before vs. after fine-tuning
The source provides concrete examples of the improvement:
Before fine-tuning: "Your HRV is an important measure of autonomic nervous system function..." [500 words of generic advice]
After fine-tuning: "Your HRV averaged 68ms this week, down 12% from last week's 77ms. Coincides with 3 nights under 7 hours sleep. Consider reducing training intensity for 48 hours."
Memory footprint and hardware
- Model (4-bit): ~2 GB
- LoRA adapter: ~50 MB
- Training memory: ~4-5 GB total
- Runs on M-series Mac, no GPU needed
The developer mentions including technical details on SQL hallucination guardrails, cross-metric context enrichment, and the training pipeline in their full writeup. They also offer to answer questions about the MLX setup or RAG layer implementation.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Hybrid Local+API Approach Cuts AI Costs by 79% in Month-Long Test
A developer running a 24/7 AI assistant on a Hetzner VPS reduced monthly costs from $288 to $60 by strategically combining local models with API calls. The setup uses nomic-embed-text for embeddings and Qwen2.5 7B for background tasks, routing more complex work to Claude models.

Project Slayer: Halo-inspired browser shooter built with Claude Code
A developer built Project Slayer, a Halo-inspired arena shooter playable in browser, using Claude Code (Opus 4.6) over two weeks with approximately 200 working hours and over 400 git commits. The game runs on FP Engine, a custom game engine built on Babylon.js.

OpenClaw on AWS Lightsail: Cost Breakdown and Configuration Lessons
A developer spent $100 in a week running OpenClaw on AWS Lightsail with Claude Sonnet 4.6 via Bedrock, discovering that sandbox settings, token management, and prompt size significantly impact functionality and costs.

User discovers hypoxic-ischemic encephalopathy diagnosis through Claude conversation
A 22-year-old from São Paulo used Claude to identify hypoxic-ischemic encephalopathy after 22 years of misdiagnosis. The AI helped connect birth complications with persistent cognitive symptoms that didn't match autism.