JetBrains on Building a RAG Pipeline for Semantic Code Search: Parsing, Chunking, Vectorization
JetBrains published the first part of a developer diary on building Air Context, their semantic code search platform, a RAG pipeline that feeds LLM agents "precise, citable evidence from real repositories instead of whatever grep happens to surface." Written by Adam Malek and Ashot Kazaryan, Part 1 covers parsing, chunking, and vectorization.
Why semantic search, not grep
The post argues that keyword search and grep fail because they require the agent to know the exact text in advance. The concrete example: "an agent looking for where session tokens get refreshed cannot rely on the code helpfully containing the word 'refresh'." To reason over abstract domains, the agent needs to search by meaning. RAG indexes source code in a way that captures semantics, then lets the agent retrieve relevant pieces on demand via free-text search.
Parsing and chunking
Parsing/chunking is the pre-processing step, and JetBrains calls it "often overlooked." Production repos contain thousands of files spanning hundreds or thousands of lines each, and agents inflate codebases further by being "prolific writers." Files may contain many classes, fields, and methods with varying relatedness.
The wrong approaches, per the post:
- Whole-file embedding — even if you could fit the file into the embedding model, search would return the entire file, defeating the goal of finding a specific function, symbol, or snippet.
- Per-line embedding — individual lines are "semantically insignificant without the surrounding context." A generic function name or comment "does not merit embedding and will produce the wrong retrieval result," overloading the agent with insignificant micro-results.
The goal is properly scoped units: the agent's exploration workflow is mostly concerned with finding a specific function, symbol, or code snippet, so chunk boundaries should align with those units (the "fine AST of parsing and chunking," in the post's phrasing).
What comes next
This is Part 1 of a series. The authors say they'll cover each stage from pre-processing to storage and agent integration, including the "wrong turns" they took. The piece stops mid-sentence before detailing the vectorization stage specifics, so expect a follow-up on how chunks are transformed into an embedding representation.
Who it's for
Developers building retrieval for coding agents on large codebases, or anyone evaluating chunking strategies for code RAG.
📖 Read the full source: HN LLM Tools
👀 See Also

Academic Research Skills for Claude Code: A Human-in-the-Loop Pipeline for Paper Writing
Academic Research Skills (ARS) v3.7.0+ is a Claude Code plugin that automates reference hunting, citation formatting, data checking, and logical consistency review while keeping the human researcher in control. Install via /plugin marketplace add Imbad0202/academic-research-skills.

Skynet: Multi-Agent Collaboration Network for Claude Code Agents
Skynet is an open-source network that enables role-based collaboration between multiple Claude Code agents and humans. It's installed as a skill using npx and managed through natural language commands.

pop-pay MCP server adds payment guardrails for Claude Code agents
pop-pay is an MCP server that lets Claude Code agents handle purchases without exposing credit card numbers. It uses CDP injection to place virtual card credentials directly into payment iframes, with Claude only receiving masked confirmation numbers.

OpenClaw plugin adds persistent memory with Engram server
A developer built a TypeScript plugin connecting OpenClaw agents to Engram, a Go-based memory server using SQLite with FTS5 search. The plugin provides 11 tools, 4 lifecycle hooks, and automatic recall that injects relevant memories into prompts before each agent turn.