Evidence is dispersed.
The right source may be easy to miss, while relevant passages are buried in long PDFs or repeated across background material.
RAG-SCHOLAR is a context-efficient retrieval system for answering research questions across large scholarly corpora. It navigates papers, sections, entities, and figures before assembling only the compact, complementary evidence needed by the answer model.
Result highlight — b = 320: 47.82 Paper F1 and 69.42 Fact F1 with 250.97 total tokens.
97.76% fewer total tokens and higher Paper F1 and Fact F1 versus CompAct.
Research answers often depend on details scattered across papers, sections, figures, and repeated terminology. Sending the entire collection to an answer model is expensive and lets redundant context compete with the evidence that actually resolves the question.
The right source may be easy to miss, while relevant passages are buried in long PDFs or repeated across background material.
RAG-SCHOLAR identifies papers and section roles first, uses scholarly entities and relations to guide retrieval, then keeps a compact, complementary text-and-visual bundle for the final answer.
Paper F1 measures source identification in the final answer; Fact F1 measures factual answering. Both are macro-averaged percentages on a 0–100 scale, evaluated on their respective query sets.
Final-answer paper identification across 116 text queries.
Factual answering on the 24-query factual subset of the text suite.
RAG-SCHOLAR keeps papers, sections, passages, figures, and scholarly entities in a navigable hierarchy. The controller selects a compact, complementary evidence bundle rather than treating a longer context as automatically better.
Recover papers, sections, passages, figures, and normalized scholarly entities.
Locate coherent candidate documents before selecting their local evidence.
Retain relevant and complementary evidence under the chosen controller setting.
Bring paper, textual, and visual evidence together for the final response.
Papers, sections, passages, figures, and local links.
offline parseEntities and cross-paper relations expose evidence topology.
structured retrievalCoarse-to-fine selection limits answer-model input.
query timeEvidence supports selected papers, facts, and claims.
answer timeAt b = 320, RAG-SCHOLAR improves both Paper F1 and Fact F1 over all ten Table 1 baselines, using fewer total tokens than nine of them. Choose a baseline to compare both scores and the measured token cost.
Paper F1 and Fact F1 · macro averages (%)
Against CompAct: +7.96 Paper F1 and +16.11 Fact F1 percentage points.
Total = Pre. + Input, per query.
RAG-SCHOLAR: Pre. 0.00 + Input 250.97.
CompAct: Pre. 11,026.36 + Input 154.96. Total savings include online compression; CompAct retains a smaller input context.
Table 2 adds visual input. RAG-SCHOLAR uses 71.4–72.0% fewer total answer-model input tokens than all three page-level multimodal systems while increasing both Paper F1 and Fact F1.
Multimodal suite · 30 paper queries / 15 reasoning queries
Text 1,090.70 + visual 217.80. Smaller text input is not offset by a larger visual payload.
ICLR 2025 · total input 4,649.97
ICCVW 2025 · total input 4,668.37
ICLR 2025 · total input 4,581.77
The cyan row is the primary compact-quality operating point used throughout this page. All token values are per-query means; the reporting convention keeps online preprocessing, answer-model input, and total cost distinct.
| Method / setting | Paper F1 ↑ | Fact F1 ↑ | Pre. tok. ↓ | Input tok. ↓ | Total tok. ↓ | Prep. (s) ↓ |
|---|---|---|---|---|---|---|
| RAG-SCHOLAR retention settings | ||||||
| b = 100 | 44.83 | 65.56 | 0 | 174.42 | 174.42 | 11.39 |
| b = 240 | 45.64 | 66.00 | 0 | 218.20 | 218.20 | 11.38 |
| b = 320 · compact quality point | 47.82 | 69.42 | 0 | 250.97 | 250.97 | 11.37 |
| b = 480 | 49.01 | 64.44 | 0 | 321.36 | 321.36 | 11.35 |
| b = 640 · Paper F1 peak | 49.49 | 58.47 | 0 | 387.05 | 387.05 | 11.36 |
| Uncapped · Fact F1 peak | 46.74 | 71.67 | 0 | 565.21 | 565.21 | 11.80 |
| Compact evidence | ||||||
| CompAct (EMNLP 2024) | 39.86 | 53.31 | 11,026.36 | 154.96 | 11,181.32 | 46.77 |
| ReComp (ICLR 2024) | 40.39 | 27.00 | 0 | 244.28 | 244.28 | 25.06 |
| Page-level VLM retrieval | ||||||
| ColPali (ICLR 2025) | 31.84 | 45.07 | 0 | 1,854.68 | 1,854.68 | — |
| M3DocRAG (ICCVW 2025) | 31.21 | 42.30 | 0 | 1,854.42 | 1,854.42 | — |
| VisRAG (ICLR 2025) | 29.97 | 28.05 | 0 | 1,854.46 | 1,854.46 | — |
| MDocAgent (arXiv 2025) | 36.67 | 33.19 | 0 | 1,838.84 | 1,838.84 | — |
| Long-context compression | ||||||
| LLMLingua-2 (ACL Findings 2024) | 36.78 | 13.19 | 42,729.27 | 2,687.00 | 45,416.27 | 205.91 |
| LongLLMLingua (ACL 2024) | 30.75 | 28.60 | 24,421.79 | 2,304.91 | 26,726.70 | 448.14 |
| Selective Context (EMNLP 2023) | 27.15 | 17.78 | 14,106.06 | 2,500.00 | 16,606.06 | — |
| LLMLingua (EMNLP 2023) | 19.78 | 11.25 | 23,596.27 | 2,234.35 | 25,830.62 | 418.99 |
Source: updated manuscript, Table 1. Costs are per-query means; Total = Pre. + Input. Pre. counts tokens processed by online LLM compression; Input counts the final serialized context. Generated output is excluded; dashes indicate unreported preparation times.
Paper F1 and Fact F1 describe different parts of a scholarly answer. The controller setting b selects an operating point; the measured context length is reported separately as Input.
Macro F1 of the paper set named in the final answer: 116 text queries and 30 multimodal queries.
Macro F1 on 24 factual text queries and 15 visual-reasoning queries, respectively.
Text suite: Total = Pre. + Input. Multimodal suite: Total = Text + Visual. Generated output is excluded.