How to Build an Efficient RAG Application: Chunking Strategies and Incremental Ingestion

Efficient RAG is not a better embedding model. It is chunking that retrieves one idea at a time, plus an incremental ingest pipeline that upserts changed chunks by deterministic IDs and deletes the rest.

Abstract pipeline of a document splitting into overlapping chunks that flow into a vector-store cylinder with upsert arrows
R

Rajkumar

Software Engineer

Last Tuesday a support engineer pasted a customer question into our RAG chatbot: What's the refund window after a subscription downgrade? The answer came back confident and wrong. We had shipped the policy change on Monday. The vector store still held Friday's chunks.

I already knew why. Our "pipeline" was a weekend script: wipe the collection, re-chunk every markdown file, re-embed, reload. It took forty minutes and a few dollars in embedding calls. We ran it when someone remembered. The model was not stale. The index was.

An efficient RAG application is not a cleverer prompt on top of a vector dump. It is two engineering problems most tutorials skip: how you cut documents into chunks that retrieve well, and how you update those chunks when the source changes without rebuilding the world. The second one is what keeps the first honest.

The Naive RAG That Looks Finished

The demo architecture is three boxes: load documents, split into 512-token windows, embed into Pinecone or pgvector. At query time you embed the question, fetch top-k, stuff the prompt. It works on a static PDF. It falls apart the week legal edits a paragraph in the middle of a 40-page guide.

Three failure modes show up immediately:

  • Cuts in the wrong place. A 512-token window bisects a table, a heading, or an "except when..." clause. Retrieval hits the first half. The model never sees the exception.
  • One topic per chunk, except when it isn't. A chunk that mixes refund policy and shipping SLAs embeds as a blur. Similarity search returns it for both questions and neither question gets a clean passage.
  • The index is a snapshot. Change one sentence and you either re-embed the entire corpus or live with drift. Both are expensive.

Before you reach for graphs, you still need chunks that are addressable, replaceable, and small enough to be precise. Similarity finds text that looks alike. It does not survive a bad split or a stale index.

Chunking Is a Retrieval Decision, Not Preprocessing

People treat the splitter as a config default. It is the most important retrieval hyperparameter you will set. A chunk is the unit the model is allowed to see. If the unit is wrong, no reranker saves you.

The tradeoff is mechanical: smaller chunks hit more precisely and miss surrounding constraints; larger chunks keep context and dilute the embedding. I aim for one idea, plus the heading path that makes it true. A refund rule should travel with the product it applies to. If I cannot point at a chunk and say what question it should answer, the split is too coarse or too fine.

Five Strategies I Actually Use

Fixed-size token windows are the baseline. Pick 256–512 tokens, 10–20% overlap, ship it. Good for homogeneous prose. Bad for markdown with tables, or anything with "see section 4." I still use them as a fallback after structure-aware splits fail.

Recursive character splitting is the LangChain default for a reason. Split on \n\n, then \n, then spaces, then characters, until the piece fits. You keep paragraphs intact more often than a raw token slice. It is the right first production splitter for mixed docs.

Structure-aware splitting is what I want on product docs. Walk markdown headings, HTML sections, or PDF bookmarks. Each H2 becomes a parent. Paragraphs under it become children. The heading path goes into metadata (Refunds > Subscriptions > Downgrades). Retrieval can filter or boost on that path. This is the strategy that stops "except when" from landing in the next chunk.

Semantic splitting embeds sentences, then cuts where cosine similarity drops. It finds topic shifts that headings miss. It is slower and less deterministic, which will matter in a minute. I use it for long narrative sources (incident write-ups, research notes) and not for versioned policy docs.

Parent-document retrieval (small-to-big) is the trick that pays for itself. Embed and search small children. Return the parent section to the model. You get precise hits without starving the prompt of context.

flowchart LR
    D[Document] --> S[Structure split]
    S --> C[Child chunks]
    C --> E[Embed children]
    E --> V[Vector store]
    V --> R[Retrieve child]
    R --> P[Expand to parent]
    P --> L[LLM context]

Default I ship: markdown or recursive split into 300–400 token children, 10–20% overlap, parent expansion at query time. FAQs go smaller (100–200). Narrative reports go larger (400–800). Top-k is 4–8 children, then expand. Put source_id, heading_path, content_hash, and splitter_version on every chunk. Incremental ingestion is impossible without that metadata.

Why Full Re-Ingest Becomes the Product

The first version of our pipeline was a batch job. New docs arrived. We rebuilt. Then the wiki moved to a CMS with webhooks, and "rebuild at 2 a.m." meant the index was wrong for 22 hours after every edit.

Embedding APIs charge per token. Re-embedding 10,000 documents because three pages changed burns budget and still serves stale answers during the run. Insert without wiping and you get duplicate vectors for the same paragraph. The model sees old and new and picks whichever ranks higher.

An index you cannot patch is a cache you pretend is a database.

Deterministic Chunk IDs Are the Whole Trick

The operational core is boring: every chunk needs an ID you can compute again from the source, then upsert. No random UUIDs. UUIDs mean you can never find yesterday's vector when today's text is 90% the same.

The naive deterministic ID looks like this:

chunk_id = f"{source_id}:{chunk_index}"

That ID is stable only if the document never grows from the top. Insert a paragraph on page one and every later chunk_index shifts. You re-embed the entire file for a one-line change. I have watched this look like "the pipeline is working" in metrics because upsert count equals chunk count.

Content-addressed IDs fix the cascade:

def chunk_id(source_id: str, text: str, splitter_version: str) -> str:
    digest = hashlib.sha256(
        f"{source_id}\n{splitter_version}\n{normalize(text)}".encode()
    ).hexdigest()[:32]
    return f"{source_id}:{digest}"

Unchanged text keeps the same ID. New text gets a new ID. The vector store upserts by ID. If you also store content_hash in metadata, you can skip the embedding call when the hash already matches what is stored.

Put splitter_version in the ID so a later geometry change does not mix two chunk shapes in one collection. For structure-aware splits I use {source_id}:{heading_path}:{hash(text)}: readable in the dashboard, stable when a sibling section moves.

The Incremental Pipeline

Updating a RAG system is not a rebuild. It is a diff against the last successful ingest.

flowchart TD
    E[Edit / webhook / git] --> H{Doc content hash unchanged?}
    H -->|yes| S[Skip]
    H -->|no| C[Chunk with pinned splitter]
    C --> I[Compute deterministic chunk IDs]
    I --> D[Diff vs existing IDs for source_id]
    D --> N[Embed only new or changed]
    N --> U[Upsert vectors]
    D --> X[Delete orphan IDs]
    U --> F[Index is current]
    X --> F

The loop is five steps:

  1. Detect. File watcher, CMS webhook, S3 event, or a git diff on docs/. Prefer a source_id that survives renames (CMS id beats file path).
  2. Hash the document. Normalize newlines and ignore generated frontmatter you do not retrieve. If the hash matches documents.content_hash, stop.
  3. Chunk deterministically. Same splitter, same version, same overlap. Non-deterministic semantic splits fight this pipeline. If you use them, pin the model and the breakpoint threshold, and accept more churn.
  4. Diff IDs. Load existing chunk IDs for that source_id. Embed the set difference. Upsert. Delete IDs that vanished. That last delete is how a removed "90-day refund" paragraph actually leaves the index.
  5. Commit the document hash. Only after upsert and orphan delete succeed. Crash between those steps and you re-run, which is safe because IDs are deterministic.

Upsert without orphan delete is how stale policy survives. I treat delete as part of ingest, not a cleanup cron.

A sketch of the write path:

def ingest_document(doc: Document, store: VectorStore) -> None:
    if store.doc_hash(doc.source_id) == doc.content_hash:
        return

    chunks = splitter.split(doc.text)  # pinned splitter_version
    new_ids = set()
    for chunk in chunks:
        cid = chunk_id(doc.source_id, chunk.text, SPLITTER_VERSION)
        new_ids.add(cid)
        if store.chunk_hash(cid) == hash_text(chunk.text):
            continue
        store.upsert(cid, embed(chunk.text), metadata={
            "source_id": doc.source_id,
            "heading_path": chunk.heading_path,
            "content_hash": hash_text(chunk.text),
            "splitter_version": SPLITTER_VERSION,
        })

    orphans = store.ids_for(doc.source_id) - new_ids
    store.delete(orphans)
    store.set_doc_hash(doc.source_id, doc.content_hash)

On a 2,000-page corpus, a typical weekday is not 2,000 embeddings. It is twelve changed pages, forty new chunks, nine orphans. That is the efficiency: you pay for what moved.

What I Watch After It Ships

  • Embed calls per ingest versus documents processed. If they stay 1:1, IDs are shifting.
  • Orphan delete count. Zero forever means you never retract. Spikes on every run mean the splitter is unstable.
  • Retrieval hit on the edited section within minutes of the webhook. I keep a canary question per critical policy page.
  • Duplicate near-matches. Same paragraph, two IDs, usually a normalize() bug (whitespace, smart quotes).

Rerankers and hybrid search sit on top of this. None of them fix an index that still contains last week's refund window.

Keep the Index Patchable

Efficient RAG is not a clever chunk size. It is a store whose units you can name, replace, and remove.

Chunk so each vector answers one kind of question, then expand to the parent when the model needs room. Identify chunks so the same text produces the same ID tomorrow. Ingest so an edit is a diff: upsert what changed, delete what left, skip everything else.

The chatbot that cited Friday's policy was not an LLM failure. It was a batch job pretending to be a database. Build the incremental pipeline first. The nicer splitter is cheaper to try once updates are a patch, not a rebuild.

New posts by email

I'll email you when I publish something new. Mostly AI agents, frontend, and things I'm learning on the job. Unsubscribe any time.