Judging What We Cannot Solve: Evaluating Research-Level Math
Judging What We Cannot Solve studies evaluation of research-level mathematical solutions when direct verification is difficult. Its method, Consequence-Based Utility, tests whether a candidate solution helps a model solve related questions with verifiable answers.
Authors: Guijin Son, Donghun Yang, Hitesh Laxmichand Patel, Hyunwoo Ko, Amit Agarwal, Sunghee Ahn, Kyong-Ha Lee, Youngjae Yu. Venue: ICML 2026. Recognition: Spotlight.
Evaluation method
Each candidate becomes an in-context example for nearby mathematical questions. Its effect on downstream performance provides a signal for ranking solutions. The study compares this approach with reward models, generative reward models, and LLM judges on problems paired with expert-written and model-generated solutions.
This work contributes to my interest in evaluating reasoning when a model’s ability to judge is limited by its ability to solve the original task.
Paper and citation
Explore my research interests and more publications.
