🎉 The dataset and fine-tuned models are now available on Hugging Face! 🎉

Ratio: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

1The Hebrew University of Jerusalem, 2Allen Institute for AI (Ai2)

Abstract

Retrieved scientific literature can serve as inspiration for both human and AI scientists. Inspiration can take different forms: prior work may directly suggest how to address a problem, or surface directions at different levels of abstraction—zooming out to a more general view or zooming in to a concrete realization. We introduce Ratio (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which we name ideation moves: ADDRESS retrieves potential approaches for stated problems, BROADEN retrieves more general formulations, and SPECIFY retrieves concrete instantiations. Ratio is constructed from millions of full-text scientific papers across CS literature via a general recipe that extends discourse-marker distant supervision—previously used only for classification—to corpus-scale retrieval, combined with extensive LLM and human vetting. Experiments show that operation-specific fine-tuning substantially boosts retrievers but leaves much room for further improvements. Ratio provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scientific inspiration retrieval.

Ideation what?

Engaging with prior work is central to how researchers develop new ideas. Consider a researcher grappling with a problem — prior literature can inspire them in three distinct ways, each a distinct move through the ideation space:

A researcher grappling with
Reward hacking in RL-trained LLMs
ADDRESSaddressing the problem
A method that mitigates it: Reward-model ensembles
BROADENabstracting it
The general principle behind it: Goodhart's law
SPECIFYgrounding it
The same principle elsewhere: Recommender systems gaming engagement metrics

Standard search ranks by topic, so the papers that suggest a move stay buried

Most scholarly retrieval benchmarks evaluate topical relevance or whether a document answers a search query. These criteria do not distinguish between passages that play different ideation roles. Run the query above through BM25 or a dense encoder, and every top hit is about reward hacking:

• Surveys of reward hacking
• Evaluations of reward over-optimization
• Definitions and taxonomies
Every hit is about the topic. Not one of them is a move.

The three inspirations above all sit in the same corpus, but they share almost no wording with the query, so a topical retriever ranks them far below. What we need is the relation between query and candidate, not the subject. Ratio organizes retrieval around exactly these three moves:

ADDRESS
Problem → Solution

The query articulates a problem; the candidate is an approach or insight that can help address it.

BROADEN
Specific → General

The candidate reformulates the query at a broader scope, such that the query can be viewed as a particular instance of it.

SPECIFY
Abstract → Concrete

The candidate instantiates or operationalizes the query through a concrete case, mechanism, or example.

RATIO overview

We construct Ratio from full-text scientific papers using discourse-marker supervision and human–LLM validation to support retrieval across three ideation moves: ADDRESS, BROADEN, SPECIFY.

The Ratio Benchmark

Authors routinely signal how a sentence relates to its predecessor with discourse markers. We exploit this to harvest query–gold pairs directly from full-text papers:

“Our model is prone to overfitting. To address this issue, we apply dropout.”
query = “Our model is prone to overfitting.”  •  marker = “To address this issue,” (stripped, kept only as metadata)  •  gold = “we apply dropout”

These labels are not weak proxies: a marker such as “To address this issue,” is an explicit, author-asserted judgment by a domain expert that the following statement constitutes a solution direction for the stated problem. Unlike prior discourse-marker distant supervision, which is confined to sentence-pair classification, we extend markers to define relation-conditioned ideation retrieval over a corpus of millions of candidates: the marker is stripped from the input and determines which relation a candidate must instantiate with respect to a query, evaluated by ranking against a shared candidate corpus.

Move-specific marker lexicons are built by an iterative multi-stage process combining manual corpus analysis, rule-based (Hearst-style) pattern expansion, and generation by multiple LLMs, followed by expert review and LLM validation. Of 4,252 candidate markers, 809 fired in the corpus; two NLP experts independently vetted all 809 and found every one valid. Markers rejected during curation become distractor markers whose sentences are added to the candidate pool as hard negatives, removing pool-membership shortcuts.

Construction of RATIO

Construction of Ratio. Validated discourse markers identify ADDRESS, BROADEN, and SPECIFY transitions in full-text papers. The marker is removed from the candidate text, publication dates determine the temporal partitions, and all operations retrieve from a shared candidate corpus.

Scale

3.02M
query–gold pairs (all splits)
13.8M
training candidate sentences
1.1M
source papers
809
human-vetted markers

Ratio comprises 3,017,476 query–gold pairs—2,779,177 SPECIFY, 222,707 ADDRESS, and 15,592 BROADEN. Splits are temporal to enforce a strict approach against contamination: training uses papers from 2015 through September 2025, validation the remainder of 2025, and the test set papers from 2026 only—postdating the public release of every evaluated model, so no backbone could have encountered them during pre-training. The candidate corpus of each split is shared by all three relations, so selecting the retrieval operation does not reveal which candidates are relevant.

Relation Train Validation Test
QueriesCandidates QueriesCandidates QueriesCandidates
SPECIFY 2,605,51513,787,834 79,583361,172 94,079404,371
ADDRESS 195,605 13,082 14,020
BROADEN 13,687 495 1,410

Temporal split with a shared candidate pool across relations. Each query is paired with a single gold candidate; the candidate corpus of each split is shared by all relations.

Human-calibrated validation

We complement the mined benchmark with a human-calibrated, LLM-validated silver test set of 17,579 queries (7,327 SPECIFY, 9,668 ADDRESS, 584 BROADEN). For each relation we write 4–6 candidate validation prompts and keep the two that best match expert judgments (F1 of .83–.84 for ADDRESS, .86–.87 for SPECIFY, .76 for BROADEN, against 300 expert annotations); inter-annotator F1 among five experts is .87 / .90 / .82. A pair is kept only if both prompts accept it. We further validate top-10 retrieved candidates to account for false negatives.

Move-specific fine-tuning helps a lot — but the task is far from solved

We evaluate BM25 and three dense retrievers (all-mpnet-base-v2, ModernBERT-embed-large, Stella-en-1.5B-v5), each as a pre-trained baseline and after relation-specific contrastive fine-tuning. Baseline pre-trained models provide limited gains over BM25, whereas move-specific fine-tuning yields substantial improvements, increasing MRR@10 by 1.6×–2.4× for ModernBERT-embed-large. ModernBERT-embed-large is the strongest model on every operation:

MRR@10 (silver test set) SPECIFY ADDRESS BROADEN
BM25 (unigram) 17.6 8.3 10.9
ModernBERT-embed-large — baseline 26.7 10.2 17.3
ModernBERT-embed-large — fine-tuned 46.7 24.5 27.3

Headline results on the silver test set (recommended-prefix setup; full tables for all models, setups, and metrics in the paper). Even the best model fails for most ADDRESS and BROADEN queries, and MRR@100 exceeds MRR@10 by at most 0.8 points—failures are not near-hits.

The mined positive is not the only valid inspiration

One gold per query makes every number a floor: with 404,371 candidates, the corpus almost certainly holds other valid answers the single mined gold misses. To measure what fine-tuning actually learned, we judge each of the top-10 retrieved candidates independently with human-calibrated LLM prompts (400 queries per relation, 24K judged pairs). After fine-tuning, at least one accepted candidate appears in the top 10 for 89.0% of SPECIFY and 76.5% of ADDRESS queries (hard agreement: both prompts accept), with accepted candidates per ADDRESS list rising from 0.61 to 1.67.

On tuned ADDRESS, for 41.2% of queries the gold is not ranked in the top 10 yet the judge accepts an alternative candidate; counting any accepted candidate raises ADDRESS MRR@10 from 24.5 to 41.2. These discovered positives overlap lexically with the query less than the mined positives do, and 88–90% of all accepted candidates come from papers other than the query's—tuned retrievers surface valid solution directions from the broader literature, i.e., inspiration in the intended sense, rather than exploiting adjacency or other shallow shortcuts.

Query-candidate lexical overlap by relation over retrieval depth

Query–candidate lexical overlap by relation, over retrieval depth on the silver test set (ModernBERT-embed-large). Lexical overlap is Jaccard score with unigrams, stopwords removed, snowball-stemmed. Tuned retrieval pulls lower-overlap candidates than the off-the-shelf baseline.

Citation (BibTeX)

@misc{sharon2026ratiobenchmarkretrievaltyped,
      title={RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature},
      author={Maayan Sharon and Tom Hope},
      year={2026},
      eprint={2608.27394},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2608.27394},
}