MB_CORE_LOG
Retrieval & EvaluationStatus: Prototype

Dynamic vs. Fixed K Reranking in RAG

A retrieval experiment comparing fixed context size with relevance-based dynamic document selection.

Tech Stack: Python · Sentence Transformers · FAISS · Cross-Encoder · Jupyter

Problem

A fixed document count can add irrelevant context for simple queries or omit context for broader queries.

What I built

Built a notebook comparison using embeddings, FAISS retrieval, cross-encoder reranking, and fixed versus score-based dynamic context selection.

Key decisions

Select by relative reranker score: Context size follows relevance but becomes sensitive to the selection threshold. Evaluate retrieval independently: Context metrics isolate the effect of document selection.

Results

In the documented synthetic experiment, mean context word count fell from 125.17 to 13.00 while answer presence in context remained 1.00 for both methods.

Limitations

Uses 100 synthetic documents and 30 queries. The notebook uses word count as a token-cost proxy; answer presence measures retrieved context, not generated-answer correctness.

Evidence

The repository contains the comparison notebook, experiment design, evaluation definitions, and aggregate results.

1. Problem Statement

A fixed document count can add irrelevant context for simple queries or omit context for broader queries.

2. Real-World Motivation

Compare selection strategies by measuring both context quality and the amount of text selected.

3. System Architecture

A synthetic notebook experiment compares fixed and score-based context sizes after the same embedding retrieval and cross-encoder reranking stages.

Both strategies retrieve and rerank 50 candidates from 100 synthetic documents. Evaluation checks known-answer presence in selected text; the notebook does not call an LLM to generate answers. Its token-cost proxy is a whitespace word count.

4. Pipeline Data Flow

The comparison holds candidate retrieval and reranking constant while changing how many documents enter the final context.

Fixed selection keeps 10 documents. Dynamic selection uses 80% of the highest raw cross-encoder score, with a top-one fallback if nothing passes. The same 30 labeled synthetic queries drive both evaluations.

5. Failure Modes & Mitigations

ScenarioImpactMitigation Strategy
Threshold discards useful contextSelection could omit information needed for an answer.Compare answer presence, ranking metrics, and context precision.
Synthetic results do not generalizeGains may differ on real document collections.Keep dataset scope explicit and validate separately on new corpora.

6. Design Tradeoffs

DecisionAlternativeRationale
Select by relative reranker scoreAlways keep a fixed document countContext size follows relevance but becomes sensitive to the selection threshold.
Evaluate retrieval independentlyJudge only final generated answersContext metrics isolate the effect of document selection.

7. Validation

Both strategies are compared on the same synthetic documents and queries with known answers.

8. Setup & Delivery

The project provides a notebook and README with setup and execution instructions.

Results & Evaluation

Evaluation Summary
The synthetic experiment compares context quality and word count as a token-cost proxy.
Evaluation Scope
  • Precision: 0.11 fixed K; 0.98 dynamic K.
  • Average selected documents: 10.00 fixed K; 1.03 dynamic K.
  • Average context word count: 125.17 fixed K; 13.00 dynamic K.

Scaling Strategy

The notebook uses a small synthetic corpus and FAISS IndexFlatL2. Large-corpus throughput is outside the measured experiment.

Security Model

The experiment uses synthetic documents. Enterprise document permissions and tenant isolation are outside its scope.

Observability

  • Context quality
    Answer presence, precision, MRR, and NDCG.
  • Context size
    Selected document count and whitespace-delimited word count.

Future Roadmap

Further evaluation would use real corpora, additional query types, and generated-answer measurements.