Context-efficient scholarly retrieval

Less context.
More answer quality.

RAG-SCHOLAR answers research questions across large scholarly corpora with compact, source-specific evidence. It separates finding the right papers from choosing what the answer model needs to read.

At b = 320: 50.36 Paper F1 · 69.42 Fact F1 · 250.97 total tokens per query.

146-paper core116 text queries30 multimodal queries
Text suite · b = 320Two quality measures
Total tokens / query250.97

97.76% fewer total tokens than CompAct, with +10.50 Paper F1 and +16.11 Fact F1 points.

Total includes online compression plus retained input; it is not the full inference cost.

The problem

Relevant repetition
can crowd out the answer.

Research papers share terminology, background, and established methods. Many passages can match a question while repeating the same concept. Within a limited context, these matches compete with the source-specific details needed to answer.

Source identification

Shared concepts guide the search.

Paper and section relevance establish where to look. Explicit source conditions keep a semantically related paper from replacing the source the question actually requires.

Answer evidence selection

Distinct details earn the context.

Separate limits on each paper and section keep one source from filling the candidate pool. Image regions are also checked for overlap with retained text.

Three connected components

Identify sources broadly.
Select evidence precisely.

RAG-SCHOLAR preserves paper → section → chunk / figure structure and scholarly entity links. The index helps identify sources; the answer model receives a selected evidence bundle, not the entire graph.

01

Hierarchical retrieval

Rank local evidence within its parent paper and section. Per-source limits prevent a paper with more matching passages from taking an unrestricted share of the candidate pool.

Papers → sections → local evidence

Default limits: 5 papers, 2 sections per paper, 2 chunks and 1 figure per section.

02

Graph-guided scholarly index

Combine explicit lexical, entity, and section-role features with a two-layer relation-aware GNN. Text-anchored representations add neighborhood information; rank fusion preserves recognized source constraints.

Explicit features + graph representations

Self-supervised training without query labels. Node vectors are exported offline; queries score the stored vectors.

03

Text-conditioned visual retrieval

Select image regions that align with the query while discounting information already covered by retained text. Overlap and area constraints guide compact visual selection.

Query relevance − text overlap − area cost

Selected windows form one enclosing crop. Visual cost is measured after image processing, not inferred from crop area.

Method: hierarchical retrieval, graph-guided scholarly index, and scholarly-context-conditioned visual retrieval. The evidence limits control source concentration, not total corpus-scoring complexity.

Text-suite evidence

Compact context.
Higher scores on both measures.

At b = 320, RAG-SCHOLAR exceeds all ten evaluated external baselines on Paper F1 and Fact F1, with fewer total tokens than nine of them. Paper F1 uses 116 queries; Fact F1 uses the 24-query factual subset.

Answer quality

Macro-averaged F1 (%) · higher is better

0–100 scale
Compact evidencePage-level VLM retrievalLong-context compression

Against CompAct: +10.50 Paper F1 and +16.11 Fact F1 percentage points.

Total token comparison

Pre. + Input · full-suite per-query mean

Lower is better
Same fixed linear scale for every baseline
RAG-SCHOLAR
b = 320
250.97
CompAct
11,181.32
045,416.27 tokens
97.76% fewer total tokens than CompAct, with higher Paper F1 and Fact F1.

What this comparison counts

Pre. is online compression cost; Input is the retained serialized context. Total adds these reported counts. Output tokens, outer answer instructions, chat framing, and offline indexing are excluded.

b = 320 · compact quality point50.36 Paper F1 · 69.42 Fact F1.
250.97 total tokens · 11.37 s preparation.
b = 480 · Paper F1 peak50.91 Paper F1 at 321.36 total tokens: the highest paper-identification score in the retention sweep.
Uncapped · Fact F1 peak71.67 Fact F1 at 565.21 total tokens. b = 320 uses 55.60% fewer tokens, with Fact F1 lower by 2.25 points.

Text-suite results, main Table 1. The strongest external text baseline on Fact F1 is CompAct. ReComp uses slightly fewer total tokens than b = 320; the selector reports that difference explicitly.

Multimodal evidence

Less input in both modalities.
More accurate answers.

On the multimodal suite, RAG-SCHOLAR improves both F1 scores over all three evaluated page-level systems. Against M3DocRAG, it gains 3.08 Fact F1 points while using 71.97% fewer total answer-model input tokens.

RAG-SCHOLAR

30 queries · 15 figure-reasoning queries

Table 2
Paper F1 ↑100.00
Fact F1 ↑60.00
Total input tokens ↓1,308.50

Text 1,090.70 + visual 217.80. Both modalities use less input than the compared page-level pipelines.

ColPali

4,649.97 input tokens

+43.67Paper F1 pts
+13.71Fact F1 pts
−71.86%input tokens

M3DocRAG

4,668.37 input tokens

+37.56Paper F1 pts
+3.08Fact F1 pts
−71.97%input tokens

VisRAG

4,581.77 input tokens

+55.33Paper F1 pts
+34.67Fact F1 pts
−71.44%input tokens
All multimodal scores and input counts
Main Table 2 · total = text + visual
MethodPaper F1 ↑Fact F1 ↑Text tok. ↓Visual tok. ↓Total tok. ↓
RAG-SCHOLAR100.0060.001,090.70217.801,308.50
ColPali56.3346.293,265.371,384.604,649.97
M3DocRAG62.4456.923,271.371,397.004,668.37
VisRAG44.6725.333,200.711,381.064,581.77
Reported results

The complete text
quality–cost record.

Nine retention settings and ten external baselines. The highlighted b = 320 row is the primary operating point used above.

Main Table 1 and complete retention sweep · costs averaged over 116 queries
Method / settingPaper F1 ↑Fact F1 ↑Pre. tok. ↓Input tok. ↓Total tok. ↓Prep. (s) ↓
RAG-SCHOLAR retention settings
b = 10045.7665.560.00174.42174.4211.39
b = 14045.4766.940.00182.86182.8611.38
b = 18046.5766.270.00196.23196.2311.37
b = 24048.7966.000.00218.20218.2011.38
b = 320 · compact quality point50.3669.420.00250.97250.9711.37
b = 480 · Paper F1 peak50.9164.440.00321.36321.3611.35
b = 64050.5458.470.00387.05387.0511.36
b = 80049.2165.890.00447.53447.5311.40
Uncapped · Fact F1 peak49.3771.670.00565.21565.2111.80
Compact evidence
CompAct (EMNLP 2024)39.8653.3111,026.36154.9611,181.3246.77
ReComp (ICLR 2024)40.3927.000.00244.28244.2825.06
Page-level VLM retrieval
ColPali (ICLR 2025)31.8445.070.001,854.681,854.68
M3DocRAG (ICCVW 2025)31.2142.300.001,854.421,854.42
VisRAG (ICLR 2025)29.9728.050.001,854.461,854.46
MDocAgent (arXiv 2025)36.6733.190.001,838.841,838.84
Long-context compression
LLMLingua-2 (ACL Findings 2024)36.7813.1942,729.272,687.0045,416.27205.91
LongLLMLingua (ACL 2024)30.7528.6024,421.792,304.9126,726.70448.14
Selective Context (EMNLP 2023)27.1517.7814,106.062,500.0016,606.06
LLMLingua (EMNLP 2023)19.7811.2523,596.272,234.3525,830.62418.99

A dash means the manuscript does not report that value. Prep. is online preparation time before answer generation; offline parsing, graph construction, and indexing are excluded. Costs and F1 have different averaging populations, described below.

Evaluation protocol

Paper identification.
Factual answering.

Text suite · b = 320. All scores below are query-level macro averages (%).

Paper F1

50.36

F1 of the paper set identified in the final answer, evaluated across 116 text queries.

Fact F1

69.42

Balances factual precision and recall on 24 factual queries with 72 reference facts. F1 is averaged per query, not recomputed from the macro precision and recall below.

Fact P

73.26

Correct reference facts relative to credited facts plus unsupported or contradicted assertions.

Fact R

68.06

The share of reference facts correctly expressed in the answer.