Running a 6-agent behavioral coaching pipeline on self-hosted Qwen3 235B with vLLM

Multi-agent behavioral coaching system
A developer has implemented a 6-agent cognitive pipeline for behavioral coaching that runs entirely on self-hosted Qwen3 models via vLLM. The system uses Claude Code instances as agents calling a vLLM endpoint, with four specialist agents firing simultaneously on each user message.
Hardware and setup
- Development: Qwen3 30B on 2x RTX 4090s
- Production: Qwen3 235B on RunPod A40 pods
- All 6 agents are Claude Code instances calling the vLLM endpoint
Pipeline architecture
Each user message triggers 6 agents in sequence:
- Shadow - Runs first, writes cross-session behavioral patterns to a shared blackboard (stated goals vs revealed priorities, follow-through prediction, pattern classification)
- Persona - OCEAN scoring, recurring goal detection, follow-through prediction percentages, growth edge identification
- Plasticity - Personality-informed coaching strategy, maps OCEAN scores to communication preferences
- Stability - Risk framework with severity/detectability/reversibility ratings, identifies blocked moves the coach should not suggest
- Coach - Fires early for an immediate response while the other agents process (~seconds)
- Synth (Pineal) - Merges all worker outputs, applies voice calibration, delivers the full response
Performance characteristics
The user sees an immediate Coach response, then the full synthesis appends approximately 40 seconds later on 2x RTX 4090s. On the A40 configuration, this takes about 108 seconds - counterintuitively slower due to different memory architecture.
Key implementation insights
What worked:
- Parallel dispatch is the key unlock for performance
- Shadow must write first because synthesis needs the blackboard content to aggregate correctly
- The sequencing logic to guarantee Shadow completes before Synth picks up adds meaningful complexity but is non-negotiable
- Context management at 235B scale is expensive - each agent gets a full context brief plus session history
- Aggressive compaction between sessions and tight per-agent context budgets have been the main reliability levers
What is hard:
- Getting agents to write structured output reliably enough for synthesis to aggregate without hallucinating merge artifacts
- Main failure mode: Synth seeing conflicting signals from Persona and Stability on the same session
The developer is seeking input from others running multi-agent systems on self-hosted inference, particularly regarding parallelism strategies at 235B scale.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Developer Creates 3D GitHub City Visualization Using Claude Code in One Day
A developer built Git City, a 3D visualization where GitHub users appear as pixel art buildings with height based on commits and width on repositories, using Claude Code exclusively in one day. The project uses Next.js, Three.js, Supabase, and Vercel.

Corporate Developer's Claude Workflow for Backend Development
A backend developer at a large US finance company shares their Claude workflow: providing detailed task descriptions with specs and internal documents, using Claude to create a working markdown document, then employing a codeReviewing agent with organizational style guidelines.

OpenClaw Agent Automates AI News Pipeline with LLM Curation
An OpenClaw agent runs a fully automated AI news pipeline that scans 25 RSS feeds, 13 Reddit subreddits, Twitter, GitHub, and web searches, then uses Gemini Flash for editorial curation and Claude Sonnet for writing. The system costs about $5/month and publishes to a Telegram channel.

Running an AI-Operated Store: Lessons from Ultrathink.art
The team behind ultrathink.art, an e-commerce store where every function is handled by AI agents, shares insights on treating agents like contractors rather than fancy autocomplete. Key differences include how you scope their work, what information you provide, and how you verify completion.