Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads
Abstract
Evaluation budgets in agentic retrieval-augmented generation span questions, search trajectories, and repeated answers. We measure allocation precision, reading efficiency, and cost boundaries using a retrieval-feedback comparison on HotpotQA and MuSiQue. At 34.14--34.39M model tokens, broader question coverage lowers standard error by 33\% versus five reads and 12.6\% versus three trajectories. Archived nested and Q-only forecasts predict these allocations within 4.0\% and 3.5\%, respectively. Depth subsets establish no clear forecasting advantage beyond the two-trajectory audit. One-read variance penalties relative to the fitted optimum at the same token budget are 0--9.9\%, with substantial Pro uncertainty. Under recorded model fees, more questions beat more trajectories at search prices of \$0--1 per 1,000 requests; question-versus-read fee rankings remain unresolved. Temperature zero cuts answer disagreement from 14.3\% to 3.4\% while comparison precision stays similar. \par\medskip\noindentKeywords: Agentic RAG; Evaluation budget; Generalizability theory; Repeated sampling.
Community
Budget Allocation Across Questions, Trajectories, and Reads
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Token-Budgeted Escalation for Financial Document QA: Cost Is Predictable, Benefit Is the Bottleneck (2026)
- Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA (2026)
- ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning (2026)
- Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection (2026)
- Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis (2026)
- The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents (2026)
- VIGIL: Verifier-Informed Gated Improvement Loop for Spreadsheet Question Answering (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.05034 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper