JetBrains on Building a RAG Pipeline for Semantic Code Search: Parsing, Chunking, Vectorization

✍️ OpenClawRadar📅 Published: October 5, 2026🔗 Source
Ad

JetBrains published the first part of a developer diary on building Air Context, their semantic code search platform, a RAG pipeline that feeds LLM agents "precise, citable evidence from real repositories instead of whatever grep happens to surface." Written by Adam Malek and Ashot Kazaryan, Part 1 covers parsing, chunking, and vectorization.

Why semantic search, not grep

The post argues that keyword search and grep fail because they require the agent to know the exact text in advance. The concrete example: "an agent looking for where session tokens get refreshed cannot rely on the code helpfully containing the word 'refresh'." To reason over abstract domains, the agent needs to search by meaning. RAG indexes source code in a way that captures semantics, then lets the agent retrieve relevant pieces on demand via free-text search.

Ad

Parsing and chunking

Parsing/chunking is the pre-processing step, and JetBrains calls it "often overlooked." Production repos contain thousands of files spanning hundreds or thousands of lines each, and agents inflate codebases further by being "prolific writers." Files may contain many classes, fields, and methods with varying relatedness.

The wrong approaches, per the post:

  • Whole-file embedding — even if you could fit the file into the embedding model, search would return the entire file, defeating the goal of finding a specific function, symbol, or snippet.
  • Per-line embedding — individual lines are "semantically insignificant without the surrounding context." A generic function name or comment "does not merit embedding and will produce the wrong retrieval result," overloading the agent with insignificant micro-results.

The goal is properly scoped units: the agent's exploration workflow is mostly concerned with finding a specific function, symbol, or code snippet, so chunk boundaries should align with those units (the "fine AST of parsing and chunking," in the post's phrasing).

What comes next

This is Part 1 of a series. The authors say they'll cover each stage from pre-processing to storage and agent integration, including the "wrong turns" they took. The piece stops mid-sentence before detailing the vectorization stage specifics, so expect a follow-up on how chunks are transformed into an embedding representation.

Who it's for

Developers building retrieval for coding agents on large codebases, or anyone evaluating chunking strategies for code RAG.

📖 Read the full source: HN LLM Tools

Ad

👀 See Also