Developer Considers Switching from DeepSeek to Grok for Finance AI Agent

Finance AI Agent Performance Issues and Potential Switch
A developer has built a finance AI web app in FastAPI/Python that functions similarly to Perplexity but for stocks. The application runs a parallel pipeline before the LLM processes queries, including live stock quotes from several finance APIs, live web search from finance search APIs, and earnings calendar data. All this structured context gets injected into the system prompt, with the model handling only reasoning and formatting while facts come from APIs, making hallucination rates less relevant for this use case.
Current Model Performance Problems
The developer is currently using DeepSeek V3.2 Reasoning and reports significant performance issues:
- TTFT (Time to First Token): ~70 seconds
- Output speed: ~25 tokens per second
- Streaming experience described as "terrible"
- Stream start timeout set to 75 seconds to avoid constant timeouts
Application Requirements
The finance AI agent has two main features:
- Chat stream: Perplexity-style finance analysis with inline source citations
- Trade check stream: Trade coach that outputs GO/NO-GO/WAIT with entry, stop-loss, target, and R:R ratio
Model requirements include:
- Fast performance with low TTFT and high tokens per second for streaming UX
- Low cost for a small project
- Smart enough for multi-step trade reasoning
- Good instruction following for strict output formats in trade checks
Considering Grok 4.1 Fast Reasoning
The developer is considering switching to Grok 4.1 Fast Reasoning based on these comparisons:
- TTFT: ~15 seconds (vs DeepSeek's ~70s)
- Output speed: ~75 tokens per second (vs DeepSeek's ~25 t/s)
- AA intelligence score: 64 vs DeepSeek's 57
- Input cost: $0.20 vs $0.28 per million tokens
Other Models Considered
The developer has also looked at Minimax 2.5, Kimi K2.5, new Qwen 3.5 models, and Gemini 3 Flash, but notes most are relatively expensive and not better for their specific use case.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Kepler builds verifiable AI for financial services with Claude: 26M+ filings indexed, audit-ready answers
Kepler's platform indexes 26M+ SEC filings across 14,000+ companies, using Claude for multi-step reasoning and a deterministic verification layer to ensure every output traces back to source documents.

UPSC StatsBuddy Bot: Telegram Interface for Indian Government Data via Claude AI
A developer built a Telegram bot called UPSC StatsBuddy that connects to India's MoSPI MCP server, using Claude AI to transform complex government datasets into clear, citeable answers for UPSC aspirants in under 30 hours.

Non-coder builds cryptographically secure AI microservice using Claude, Gemini, and ChatGPT
A 60-year-old with zero coding experience built an AI microservice called AgentGate in one week using Claude Code for writing, with cross-auditing by Gemini and ChatGPT. The system includes SQLite database, progressive rate-limiting, Ed25519 cryptographic signing, and 50+ passing tests.

Porting Quake to Three.js with Claude Code: Workflow and Limitations
A developer used Claude Code to port Quake's source code to JavaScript and Three.js, creating a web-based version. The project involved significant prompting work and revealed Claude's difficulty with porting multiplayer server code to Deno+WebTransport.