Objective: Evaluate the impact of retrieval depth (Top-K) and response length (Max Tokens) on answer quality while keeping all other parameters constant.
Evaluation Method: Each configuration was evaluated across five financial test cases using an LLM-as-a-Judge framework that scored responses for Relevance, Groundedness, Accuracy, Completeness, and Risk Awareness.
Experimental Variables: Only Top-K Retrieval and Max Tokens were varied across Configurations A–C. All remaining RAG parameters were held constant.
Selection Rationale — these two parameters directly influence the two primary stages of a RAG pipeline:
- Top-K Retrieval determines how much relevant context is retrieved from the vector database.
- Max Tokens determines how much of that context the LLM can incorporate into its generated response.
- Varying only these parameters enabled a controlled evaluation of how retrieval depth and response length affect answer quality while minimizing the influence of other factors.
Total Evaluations: 15 (5 Test Cases × 3 Configurations)
| Parameter |
Config A |
Config B |
Config C |
| Chunk Size | 500 | 500 | 500 |
| Chunk Overlap | 50 | 50 | 50 |
| Top-K | 8 | 5 | 10 |
| Temperature | 0 | 0 | 0 |
| Top-p | 1 | 1 | 1 |
| Max Tokens | 1000 | 800 | 1500 |
| Test Case |
Config A |
Config B |
Config C |
| Test Case 1 — Amazon | ✓ | ✓ | ✓ |
| Test Case 2 — Apple | ✓ | ✓ | ✓ |
| Test Case 3 — Bitcoin | ✓ | ✓ | ✓ |
| Test Case 4 — JPM | ✓ | ✓ | ✓ |
| Test Case 5 — Nvidia | ✓ | ✓ | ✓ |