Cloudflare's AI Platform: Unified Inference Layer for AI Agents

What Cloudflare's AI Platform Offers
Cloudflare has expanded its AI capabilities into a unified inference layer designed specifically for AI agents. The platform addresses the challenge of AI models changing rapidly and the need to use multiple models for different tasks within agentic workflows.
Key Features and Implementation
The core offering is one API to access any AI model from any provider. For Workers users, you can call third-party models using the same AI.run() binding already used for Workers AI. Switching between providers requires only a one-line code change.
const response = await env.AI.run('@cf/moonshotai/kimi-k2.5', {
prompt: 'What is AI Gateway?'
}, {
metadata: {
"teamId": "AI",
"userId": 12345
}
});The platform provides access to 70+ models across 12+ providers including Alibaba Cloud, AssemblyAI, Bytedance, Google, InWorld, MiniMax, OpenAI, Pixverse, Recraft, Runway, and Vidu. Model offerings now include image, video, and speech models for building multimodal applications.
Cost Management and BYOM Support
All AI spend can be managed in one place through AI Gateway. By including custom metadata with requests, you can get cost breakdowns by attributes like free vs. paid users, individual customers, or specific workflows.
For custom model needs, Cloudflare is working on letting users bring their own models to Workers AI using Replicate's Cog technology. This involves containerizing machine learning models with a cog.yaml file and Python inference code, abstracting away CUDA dependencies, Python versions, and weight loading.
Recent Updates and Availability
Recent additions include zero-setup default gateways, automatic retries on upstream failures, and more granular logging controls. REST API support for non-Workers users is coming in the coming weeks.
📖 Read the full source: HN AI Agents
👀 See Also

ClawControl iOS client released for OpenClaw self-hosted servers
ClawControl v1.50 is now available on iOS as a privacy-focused mobile client for self-hosted OpenClaw/Claw servers. The open-source app enables real-time chat with streaming responses, agent management, and session control from mobile devices.

FlowBoard v5: Event-Sourced Project Workspace for Multi-Agent Teams
FlowBoard v5 rebuilds the project context layer on React with an event-sourced task store, enables multi-agent coordination across OpenClaw, Claude Code, and Cursor, and introduces an Ideas-to-Specs pipeline.

SkyClaw v2.2 Rust AI Agent Runtime Adds OpenAI OAuth and Custom Tool Authoring
SkyClaw v2.2 introduces OpenAI OAuth authentication using ChatGPT Plus/Pro subscriptions, custom tool authoring where agents write their own bash/python/node tools at runtime, and daemon mode for background operation. The Rust-based runtime benchmarks at 31ms cold start, 15MB idle RAM, and 9.3MB binary size.

OpenClaw Model Performance Review: Codex 5.3 Leads, GLM Models Disappoint
A developer tested multiple AI models with OpenClaw, finding Codex 5.3 performs best with 9/10 rating, while GLM 4.7 and GLM 5 scored 5/10 due to high token usage, slow responses, and inconsistent output.