AI Code Retrieval Primer
Embeddings, chunking, and re-ranking — the moving parts behind "the agent found the right file", explained from the ground up.
When an agent “just knows” which file to open, there’s a retrieval pipeline doing the work. This note unpacks it, one stage at a time.
Why keyword search isn’t enough
Grep finds the string you typed. Retrieval finds the thing you meant — even when the code uses different words than your question.
This is a budding note — the shape is here, but I’m still tending the examples. Check back as it grows toward evergreen.
Embeddings, briefly
An embedding turns a chunk of code or text into a vector — a point in space where “close together” means “similar in meaning”. Retrieval is then just: find the nearest points to your query.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
vectors = model.encode(["a branch is a pointer", "git refs explained"])
The cheapest quality win is usually better chunking, not a fancier model. Split on meaningful boundaries — functions, headings — not arbitrary token counts.