HF Viewer: Visualize Any Hugging Face Model Graph Instantly

HF Viewer is a new web tool that lets you visualize the architecture of any Hugging Face model directly in the browser. No installation, no export step, no config hunting. Just paste a model URL or repo name — e.g., gpt2 — and get an interactive graph showing the high-level structure, from encoder-decoder transformers to sparse MoE reasoning models.
Key Features
- Direct URL magic: Replace
huggingface.cowithhfviewer.comin any model URL to view it instantly. - Granularity levels: Zoom from overview down to specific sub-structures, such as attention blocks, vision encoders, or expert routing.
- Model family comparison: Compare related models side-by-side with synchronized pan/zoom — currently showcased for the Gemma 4 family.
- Embed in model cards: Press the Embed button to get an iframe snippet for your own model card.
How to Use
Navigate to hfviewer.com, paste a Hugging Face model URL or repo name in the input box, and click "Visualize Model". Alternatively, manually replace huggingface.co with hfviewer.com in the URL bar.
For example, to visualize GPT-2: open https://hfviewer.com/gpt2.
Use Cases
The tool is designed for developers and ML engineers who need to quickly understand a model's architecture without reading through config files or source code. It supports a range of popular models including:
- Qwen/Qwen3.5-0.8B — small instruction-tuned LLM
- google/vit-base-patch16-224 — vision backbone
- openai/clip-vit-base-patch32 — dual encoder
- t5-small — encoder-decoder
- nvidia/parakeet-tdt-0.6b-v3 — streaming Conformer-TDT speech recognizer
Interactive Blog Format
On the Gemma 4 family page, the blog text and graph are linked. You can read about an architectural decision and jump into the corresponding part of the graph, then return to the article with surrounding context intact. This graph-to-text loop offers a new way to communicate ML architecture.
HF Viewer is released as a free community tool by the Embedl team.
📖 Read the full source: HN AI Agents
👀 See Also

Control Real iPhones from OpenClaw Agents — Hands-On with AvaBerry43's Setup
A developer created an API-controlled iPhone setup for agents, testing features like iMessage drafting, Shortcuts, and mobile QA. 70 phones are available for others to experiment with.

Local AI Agent Achieves Sub-Second STT and TTS Latency with Open-Source Servers
A developer achieved ~0.2s STT latency using Whisper large-v3-turbo with hybrid thread-managed GPU architecture and ~250ms TTS latency with Coqui-TTS optimized for low-latency synthesis. Both implementations are fully self-hosted and open-sourced.

Efficient Token Management with Open-Source MCP Servers: Pare
Pare MCP servers reduce token waste and enhance efficiency when AI coding agents use developer tools by providing structured output.

Context Mode MCP Server Cuts Claude Code Context Usage by 98%
Context Mode is an MCP server that reduces Claude Code context consumption from 315 KB to 5.4 KB by sandboxing tool outputs. It supports 10 language runtimes and includes a knowledge base with full-text search.