Tasks
Overview
The challenge includes two tasks. CQ Generation produces competency questions from a variety of inputs (user stories, datasets, ontologies), and Ontology Generation produces ontologies from competency questions. Each task reuses an existing benchmark and can be entered independently. Each task also has a hidden dataset on which a system that also provides its source code will be evaluated. For the CQ Generation task, participants may choose which input or inputs the system generates CQs from, for example from user stories and PDFs rather than from ontologies.
Task 1 — CQ Generation Active
Objective
Given a specific input, generate competency questions that capture the functional requirements of the intended ontology.
Description
Competency questions (CQs) are natural-language questions that an ontology should answer. Systems generate CQs from different kinds of input and are evaluated with the Bench4KE benchmarking system against its reference competency questions.
Input / Output
Input: user stories, structured or semi-structured data, ontologies, or PDF documents. Output: a set of competency questions in the required format.
benchmarkdataset_input.csv
The input side of the Bench4KE benchmark, published in
dataset/CQGen/:
73 rows, semicolon-separated, UTF-8, one row per distinct input. It contains no competency questions. A system takes
Description, Scenario, Dataset or Link as its input; the reference questions it is
scored against stay with the organisers. Here is what each column contains:
| Column | Description | Example |
|---|---|---|
id | Stable identifier of the input row, CQGEN-001 … CQGEN-073. | CQGEN-001 |
Project Name | The ontology project the row belongs to. | Polifonia |
Name | The persona, use case or module within the project; may be empty. | Medical diagnostics |
Description | Notes about the project's source data; may be empty. | This epidemiological dataset from the Lombardy region captures detailed information about gastrointestinal and foodborne illnesses … |
Scenario | A user story or persona describing what the ontology must support; may be empty. | Amy wants to assess the developments of organ builders. This research includes looking into which organs an organ builder worked on and how their projects developed over time. … |
Dataset | A data sample, or the URL of a data file; may be empty. | codice_ente,ente_controllore,punto_prelievo,tipologia_punto_prelievo,indirizzo_punto_prelievo, … |
Link | The URL of an ontology file (.owl, .ttl, .rdf), a PDF document or a repository; may be empty. | https://raw.githubusercontent.com/D2KLab/llm4ke/…/swo_merged.owl |
Every row carries at least one of Description, Scenario, Dataset or Link.
Participants choose which input or inputs their system generates from, so rows that lack the input a system uses are simply skipped.
submission_output.csv
One CSV file with one row per generated competency question. Here is what each column contains:
| Column | Description | Example |
|---|---|---|
Project Name | The project of the dataset row the CQ was generated from. | Polifonia |
Name | The Name of that dataset row; may be empty. | Music Meta Ontology |
Scenario | The scenario used as input, copied from the dataset row; empty if no scenario was used. | Amy wants to assess the developments of organ builders. This research includes looking into which organs an organ builder worked on … |
Dataset | The dataset value used as input, copied from the dataset row; empty if no dataset was used. | codice_ente,ente_controllore,punto_prelievo,tipologia_punto_prelievo, … |
Link | The ontology or PDF URL used as input, copied from the dataset row; the column can be left out when no such input was used. | https://raw.githubusercontent.com/D2KLab/llm4ke/…/swo_merged.owl |
Generated CQs | One generated competency question per row. | Which organs did the organ builder work on? |
Evaluation and Metrics
The task is evaluated with Bench4KE. Although Bench4KE may compute additional diagnostic lexical and semantic metrics, the official CQ Generation ranking score is based on a combined metric that balances coverage, semantic precision, diversity, and verbosity control.
Coverage / Hit Rate measures the proportion of reference CQs covered by at least one generated CQ above the semantic similarity threshold. Precision MMS measures how cleanly the generated CQs map to the reference CQs by averaging, for each generated CQ, its maximum semantic similarity to the reference CQs. Average Centroid Distance measures the semantic diversity of the generated CQ set. A Verbosity Penalty discourages over-generation and prevents systems from obtaining high coverage by generating excessive numbers of CQs.
The official CQ Generation ranking score is:
S = (2 · (Cov · Prec_MMS) / (Cov + Prec_MMS)) × (1 + α · ACD) × VP
where Cov is Coverage / Hit Rate, Prec_MMS is Precision MMS, ACD is Average Centroid Distance, and VP is an exponential verbosity penalty. The harmonic mean of Coverage and Precision MMS forms the core accuracy score, ACD provides a small diversity bonus, and VP penalises submissions that generate substantially more CQs than the reference set. The exact values of the parameters are fixed in the official task configuration.
Reference benchmark: Bench4KE (github.com/fossr-project/ontogenia-cini), DOI 10.5281/zenodo.17817277
Task 2 — Ontology Generation Active
Builds on the cq4oe-benchmark (CQ4OE), whose reference ontologies are aligned with competency questions across six source ontologies (Wine, AWO, ODRL, SAREF4WATR, VGO, and SWO) spanning three size tiers. It is split into two sub-tasks, and participants may enter either or both.
Both sub-tasks are evaluated against a reference sub-ontology per domain, which stays with the organisers: CQ2Term is compared against its terms, and CQ2Onto against its full axioms.
Sub-task 2a — CQ2Term
Objective. Given a set of competency questions, predict the explicit terms (classes and properties) each CQ requires. It targets the conceptualization step. Evaluated over 99 CQs.
Description. The system identifies the vocabulary an ontology would need to answer the CQs, without yet building the full axiomatisation. The output is a set of candidate classes and properties.
Input / Output. Input: a set of competency questions. Output: a set of terms (classes and properties).
<domain>_cq2term_cqs.json
A JSON list of competency questions, provided by the benchmark in
CQ2Term/competency_question/
(99 CQs over the six domains). Each question has an id and its text in value.
[
{"id": "CQ1", "value": "Which animal eats which other animal?"},
{"id": "CQ2", "value": "Is [this animal] a herbivore?"}
]
<domain>_cq2terms_terms.json
The same list, with the predicted terms added to every question: the classes in
class and the properties in property (singular keys, arrays of strings, empty when none).
Keep id and value unchanged.
[
{"id": "CQ1", "value": "Which animal eats which other animal?",
"class": ["Animal"], "property": ["eats"]},
{"id": "CQ2", "value": "Is [this animal] a herbivore?",
"class": ["Animal", "Herbivore"], "property": []}
]
Evaluation and Metrics. Predicted and reference terms are first aligned using five similarity methods (hard matching, sequence matching, Levenshtein, Jaro–Winkler, and embedding-based semantic similarity). On that alignment the metrics are Precision, Recall, and F1, plus CQ-conditioned coverage (at-least-one, mean, and full) that checks whether each term is recovered under the CQ that requires it.
Sub-task 2b — CQ2Onto
Objective. Given a set of competency questions, generate a full OWL ontology that satisfies them. It tests whether a model recovers implicit and derived terms and expresses them as a coherent ontology. Evaluated over 118 CQs.
Description. Beyond extracting terms, the system produces class and property hierarchies, property characteristics, domain and range axioms, and other restrictions, so that the resulting ontology answers the competency questions.
Input / Output. Input: a set of competency questions. Output: an OWL ontology.
<domain>_cq2onto_cqs.json
A JSON list of competency questions, provided by the benchmark in
CQ2Onto/competency_question/,
in the same id / value format as CQ2Term. The whole set of a domain is the input to one run.
[
{"id": "CQ1", "value": "Which animal eats which other animal?"},
{"id": "CQ2", "value": "Is [this animal] a herbivore?"},
{"id": "CQ6", "value": "Which plants eat animals?"}
]
<domain>_ontology.owl
An ontology written in the OWL 2 language (W3C OWL 2 Web Ontology Language),
one file per domain, generated from the competency questions of that domain. Any RDF serialisation that rdflib can parse
is accepted (RDF/XML as .owl, .rdf);
the file name must start with the domain identifier.
Evaluation and Metrics. Evaluated against the CQ-aligned reference ontology across five targets: term recovery, property characteristics, domain/range triples, TBox axioms, and hierarchy closure. Each reports Precision, Recall, and F1 (global and alignment-conditioned views where applicable), with CQ-conditioned coverage at the axiom level and after HermiT reasoning-based closure recovery.