SemEval-2027 · Task 1

RETECO: Reasoning-Oriented Retrieval

A shared task on temporally grounded retrieval and reasoning-intensive conversational retrieval.

Status: Accepted as a SemEval-2027 shared task.

Organizers

  • Abdelrahman AbdallahUniversity of Innsbruck
  • Mohammed AliUniversity of Innsbruck
  • Muhammad Abdul-MageedUniversity of British Columbia
  • Kevin DuhJohns Hopkins University
  • Adam JatowtUniversity of Innsbruck

All listed team members are co-organizers.

Overview

Retrieval systems are commonly evaluated by topical relevance. RETECO instead focuses on cases where a system must reason about when information is valid, how events or facts change over time, and what has already been established in a conversation.

The task brings together two complementary settings grounded in the existing TEMPO and RECOR benchmarks. Participants may enter an individual sub-track or work across retrieval and grounded generation.

Task description

Track 1: Temporal grounded retrieval (TEMPO)

Systems retrieve passages that are topically relevant and temporally aligned with a complex query. The track covers 13 domains.

  • Sub-track 1a — Temporal retrieval: rank evidence for the complete query.
  • Sub-track 1b — Step-wise temporal retrieval: retrieve evidence for decomposed reasoning steps.

Track 2: Conversational reasoning retrieval (RECOR)

Systems use conversation history and discourse dependencies to retrieve reasoning-dependent evidence and, in the generation settings, produce a grounded response. The track covers 11 domains.

  • Sub-track 2a — Conversational retrieval: rank supporting passages for the current turn.
  • Sub-track 2b — Gold-passage generation: generate a response from organizer-provided evidence.
  • Sub-track 2c — Full conversational RAG: retrieve up to five passages and generate a faithful response.

Read the complete task definitions and input/output formats →

Data and evaluation

Public pilot data are available from TEMPO and RECOR. RETECO will use a separate held-out test set for the SemEval evaluation phase.

nDCG@10 is the official retrieval leaderboard metric. Track-specific diagnostics report temporal coverage, temporal precision, conversation-turn depth, domain variation, and grounded generation quality.

Participation

Registration, the official evaluation platform, submission limits, and contact channels will be announced here when confirmed. Participants may use the public pilot repositories to inspect the data formats and reproduce baseline systems in the meantime.

Read the participation and reproducibility guide →

Papers and proposal

RETECO is based on the TEMPO and RECOR benchmarks. The revised task proposal, paper links, and ready-to-copy BibTeX records are available on the papers and citations page.