Publication
arXiv
Stage
Preprint
What we read
Summary of the abstract
Authors
Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

Experts asked 184 senior authors to label 207 recent computer science papers and say which earlier works would have helped their projects. AI found related papers but was slower and less accurate than humans across methods and data sets.

Why it matters

Humans still steer AI toward solid starting points. Teams should mix AI screening with human review when looking at old studies for new ideas.

Experts built a test by asking experienced researchers to name earlier papers that would have helped their work. AI tools struggled to match this judgment.

In practice, use AI to filter a first set of papers and let humans decide which truly provide strong ideas for a new project.

What remains uncertain

This uses only abstracts and preprints, so results may not apply to full papers or to established fields.

Read the paper PDF ↗

Original sources · 1
  1. ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research ↗arXiv · 2026-10-01

Check the original paper for its authors, methods, version and access terms.