Judging What We Cannot Solve: Evaluating Research-Level Math

Judging What We Cannot Solve studies evaluation of research-level mathematical solutions when direct verification is difficult. Its method, Consequence-Based Utility, tests whether a candidate solution helps a model solve related questions with verifiable answers.

Authors: Guijin Son, Donghun Yang, Hitesh Laxmichand Patel, Hyunwoo Ko, Amit Agarwal, Sunghee Ahn, Kyong-Ha Lee, Youngjae Yu. Venue: ICML 2026. Recognition: Spotlight.

Evaluation method

Each candidate becomes an in-context example for nearby mathematical questions. Its effect on downstream performance provides a signal for ranking solutions. The study compares this approach with reward models, generative reward models, and LLM judges on problems paired with expert-written and model-generated solutions.

This work contributes to my interest in evaluating reasoning when a model’s ability to judge is limited by its ability to solve the original task.

Paper and citation

Explore my research interests and more publications.