Measuring the Gap Between Human and LLM Research Ideas
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM generated ideas from human researchers? To characterize this gap, we build a large scale evaluation framework for ideation from high quality human research papers. For each paper, we reverse engineer a small set of clo...