01 · SAMPLEInspect
Small real examples demonstrate queries, documents, conversations, qrels, and output structures.
RETECO · SemEval-2027 Task 1
RETECO builds on two operational benchmarks, documents every domain collection, and preserves a separate held-out SemEval test.
Release model
01 · SAMPLESmall real examples demonstrate queries, documents, conversations, qrels, and output structures.
02 · TRAIN / DEVPublic TEMPO and RECOR resources support training, validation, and reproducible baselines.
03 · TESTQueries are released for submission while gold judgments remain private.
Curated trial package
The repository includes a compact, human-inspected package copied from the official public releases: five TEMPO examples and four RECOR conversations, every referenced positive passage, TREC-style qrels, and a pinned provenance manifest.
9 task records · 41 supporting passages · 58 positive judgments
Track 01 pilot
TEMPO contains temporal reasoning-intensive queries requiring evidence about periods, trends, events, or change. Its 13 independent domain corpora contain 1,654,055 documents in total.
| Group | Domain | Queries | Documents | Avg. gold/query | Avg. steps |
|---|---|---|---|---|---|
| Blockchain | Bitcoin | 100 | 153,291 | 3.3 | 2.93 |
| Blockchain | Cardano | 51 | 87,201 | 2.5 | 2.84 |
| Blockchain | IOTA | 10 | 10,372 | 3.8 | 3.20 |
| Blockchain | Monero | 65 | 85,093 | 2.6 | 2.72 |
| Social Sciences | Economics | 83 | 93,756 | 3.6 | 3.08 |
| Social Sciences | Law | 35 | 43,288 | 3.0 | 3.23 |
| Social Sciences | Politics | 150 | 183,394 | 2.7 | 3.35 |
| Social Sciences | History | 801 | 356,493 | 4.5 | 3.42 |
| Applied | Quantitative Finance | 34 | 28,785 | 2.4 | 2.68 |
| Applied | Travel | 100 | 177,677 | 2.6 | 3.11 |
| Applied | Workplace | 36 | 64,659 | 2.8 | 2.42 |
| Applied | Genealogy | 115 | 156,228 | 2.8 | 3.78 |
| STEM | History of Science & Mathematics | 150 | 213,818 | 2.5 | 3.25 |
| — | Total | 1,730 | 1,654,055 | — | — |
Counts follow the official TEMPO release and paper. Each domain is a separate dataset split and retrieval corpus.
Track 02 pilot
RECOR combines multi-turn context with reasoning-dependent passage relevance. Six domains originate from BRIGHT collections and five from StackExchange; together they contain 507,141 documents.
| Source | Domain | Conversations | Turns | Corpus documents | Avg. docs/turn |
|---|---|---|---|---|---|
| BRIGHT | Biology | 85 | 362 | 57,359 | 1.56 |
| BRIGHT | Earth Science | 98 | 454 | 121,249 | 1.58 |
| BRIGHT | Economics | 74 | 288 | 50,220 | 2.28 |
| BRIGHT | Psychology | 84 | 333 | 52,835 | 2.16 |
| BRIGHT | Robotics | 68 | 259 | 61,961 | 1.76 |
| BRIGHT | Sustainable Living | 78 | 319 | 60,792 | 1.88 |
| StackExchange | Drones | 37 | 142 | 16,381 | 2.36 |
| StackExchange | Hardware | 46 | 188 | 26,308 | 2.10 |
| StackExchange | Law | 50 | 230 | 20,027 | 2.55 |
| StackExchange | Medical Sciences | 44 | 183 | 23,297 | 2.23 |
| StackExchange | Politics | 43 | 213 | 16,712 | 2.49 |
| — | Total | 707 | 2,971 | 507,141 | 2.01 |
Corpus-document counts are verified from the official Hugging Face dataset metadata; conversation, turn, and relevance statistics follow the published RECOR paper.
Evidence of feasibility
The pilots already include sparse, dense, and reasoning-specialized retrieval systems. These representative nDCG@10 values were reported in the accepted proposal and show substantial remaining headroom.
| Model | Track 1 · TEMPO | Track 2 · RECOR | Family |
|---|---|---|---|
| BM25 | 10.8 | 44.6 | Sparse lexical |
| BGE | 22.0 | 41.1 | Dense encoder |
| ReasonIR | 27.3 | 49.6 | Reasoning retriever |
| DiVeR | 32.0 | 54.5 | Reasoning retriever |
Machine-readable data
Public resources use JSON/JSONL records and TREC-style qrels. Final SemEval input and submission schemas will ship with examples and a format checker before evaluation.
| Resource | Typical fields | Purpose |
|---|---|---|
| TEMPO examples | id, query, gold_ids, gold_answers | Whole-query retrieval |
| TEMPO steps | id, query, gold_ids | Step-wise retrieval |
| RECOR benchmark | id, task, turns, metadata | Conversation container |
| RECOR turn | turn_id, query, history, gold_doc_ids | Target-turn retrieval |
| Documents | document id, content | Retrieval corpus |
| Qrels | query id, document id, relevance | Local scoring |
Responsible release
The SemEval task release is planned for CC BY 4.0 distribution and Zenodo archival, with PII and source-license screening. Code and existing benchmark repositories may carry their own licenses; participants must follow the license displayed with each resource.