RECOR: Conversational Retrieval with Reasoning

RECOR evaluates information retrieval that depends on both conversation history and reasoning. The benchmark contains 707 conversations and 2,971 turns across eleven domains.

Authors: Mohammed Ali, Abdelrahman Abdallah, Amit Agarwal, Hitesh Laxmichand Patel, Adam Jatowt. Venue: ACL 2026 Findings.

What the benchmark measures

A decomposition and verification pipeline builds conversations from source-grounded facts, with explicit retrieval reasoning for each turn. Evaluation examines how history and reasoning change retrieval quality and identifies remaining difficulty with implicit logical connections.

The paper reports an increase in nDCG@10 from 0.236 to 0.479 when combining conversation history and reasoning in its evaluated setting. The results inform the retrieval component of systems that must follow a user’s information needs across multiple turns.

Paper, data, and code

Explore enterprise retrieval and more publications.