Tasks

Overview

The challenge includes two tasks. CQ Generation produces competency questions from a variety of inputs (user stories, datasets, ontologies), and Ontology Generation produces ontologies from competency questions. Each task reuses an existing benchmark and can be entered independently. Each task also has a hidden dataset on which a system that also provides its source code will be evaluated. For the CQ Generation task, participants may choose which input or inputs the system generates CQs from, for example from user stories and PDFs rather than from ontologies.

Task 1 — CQ Generation Active

Objective

Given a specific input, generate competency questions that capture the functional requirements of the intended ontology.

Description

Competency questions (CQs) are natural-language questions that an ontology should answer. Systems generate CQs from different kinds of input and are evaluated with the Bench4KE benchmarking system against its reference competency questions.

Input / Output

Input: user stories, structured or semi-structured data, ontologies, or PDF documents. Output: a set of competency questions in the required format.

Input · the challenge dataset benchmarkdataset_input.csv

The input side of the Bench4KE benchmark, published in dataset/CQGen/: 73 rows, semicolon-separated, UTF-8, one row per distinct input. It contains no competency questions. A system takes Description, Scenario, Dataset or Link as its input; the reference questions it is scored against stay with the organisers. Here is what each column contains:

ColumnDescriptionExample
idStable identifier of the input row, CQGEN-001 … CQGEN-073.CQGEN-001
Project NameThe ontology project the row belongs to.Polifonia
NameThe persona, use case or module within the project; may be empty.Medical diagnostics
DescriptionNotes about the project's source data; may be empty.This epidemiological dataset from the Lombardy region captures detailed information about gastrointestinal and foodborne illnesses …
ScenarioA user story or persona describing what the ontology must support; may be empty.Amy wants to assess the developments of organ builders. This research includes looking into which organs an organ builder worked on and how their projects developed over time. …
DatasetA data sample, or the URL of a data file; may be empty.codice_ente,ente_controllore,punto_prelievo,tipologia_punto_prelievo,indirizzo_punto_prelievo, …
LinkThe URL of an ontology file (.owl, .ttl, .rdf), a PDF document or a repository; may be empty.https://raw.githubusercontent.com/D2KLab/llm4ke/…/swo_merged.owl

Every row carries at least one of Description, Scenario, Dataset or Link. Participants choose which input or inputs their system generates from, so rows that lack the input a system uses are simply skipped.

Output · one file per submission submission_output.csv

One CSV file with one row per generated competency question. Here is what each column contains:

ColumnDescriptionExample
Project NameThe project of the dataset row the CQ was generated from.Polifonia
NameThe Name of that dataset row; may be empty.Music Meta Ontology
ScenarioThe scenario used as input, copied from the dataset row; empty if no scenario was used.Amy wants to assess the developments of organ builders. This research includes looking into which organs an organ builder worked on …
DatasetThe dataset value used as input, copied from the dataset row; empty if no dataset was used.codice_ente,ente_controllore,punto_prelievo,tipologia_punto_prelievo, …
LinkThe ontology or PDF URL used as input, copied from the dataset row; the column can be left out when no such input was used.https://raw.githubusercontent.com/D2KLab/llm4ke/…/swo_merged.owl
Generated CQsOne generated competency question per row.Which organs did the organ builder work on?

Evaluation and Metrics

The task is evaluated with Bench4KE. Although Bench4KE may compute additional diagnostic lexical and semantic metrics, the official CQ Generation ranking score is based on a combined metric that balances coverage, semantic precision, diversity, and verbosity control.

Coverage / Hit Rate measures the proportion of reference CQs covered by at least one generated CQ above the semantic similarity threshold. Precision MMS measures how cleanly the generated CQs map to the reference CQs by averaging, for each generated CQ, its maximum semantic similarity to the reference CQs. Average Centroid Distance measures the semantic diversity of the generated CQ set. A Verbosity Penalty discourages over-generation and prevents systems from obtaining high coverage by generating excessive numbers of CQs.

The official CQ Generation ranking score is:

S = (2 · (Cov · Prec_MMS) / (Cov + Prec_MMS)) × (1 + α · ACD) × VP

where Cov is Coverage / Hit Rate, Prec_MMS is Precision MMS, ACD is Average Centroid Distance, and VP is an exponential verbosity penalty. The harmonic mean of Coverage and Precision MMS forms the core accuracy score, ACD provides a small diversity bonus, and VP penalises submissions that generate substantially more CQs than the reference set. The exact values of the parameters are fixed in the official task configuration.

Reference benchmark: Bench4KE (github.com/fossr-project/ontogenia-cini), DOI 10.5281/zenodo.17817277

Task 2 — Ontology Generation Active

Builds on the cq4oe-benchmark (CQ4OE), whose reference ontologies are aligned with competency questions across six source ontologies (Wine, AWO, ODRL, SAREF4WATR, VGO, and SWO) spanning three size tiers. It is split into two sub-tasks, and participants may enter either or both.

Both sub-tasks are evaluated against a reference sub-ontology per domain, which stays with the organisers: CQ2Term is compared against its terms, and CQ2Onto against its full axioms.

Sub-task 2a — CQ2Term

Objective. Given a set of competency questions, predict the explicit terms (classes and properties) each CQ requires. It targets the conceptualization step. Evaluated over 99 CQs.

Description. The system identifies the vocabulary an ontology would need to answer the CQs, without yet building the full axiomatisation. The output is a set of candidate classes and properties.

Input / Output. Input: a set of competency questions. Output: a set of terms (classes and properties).

Input · one file per domain <domain>_cq2term_cqs.json

A JSON list of competency questions, provided by the benchmark in CQ2Term/competency_question/ (99 CQs over the six domains). Each question has an id and its text in value.

[
  {"id": "CQ1", "value": "Which animal eats which other animal?"},
  {"id": "CQ2", "value": "Is [this animal] a herbivore?"}
]
Output · one file per domain <domain>_cq2terms_terms.json

The same list, with the predicted terms added to every question: the classes in class and the properties in property (singular keys, arrays of strings, empty when none). Keep id and value unchanged.

[
  {"id": "CQ1", "value": "Which animal eats which other animal?",
   "class": ["Animal"], "property": ["eats"]},
  {"id": "CQ2", "value": "Is [this animal] a herbivore?",
   "class": ["Animal", "Herbivore"], "property": []}
]

Evaluation and Metrics. Predicted and reference terms are first aligned using five similarity methods (hard matching, sequence matching, Levenshtein, Jaro–Winkler, and embedding-based semantic similarity). On that alignment the metrics are Precision, Recall, and F1, plus CQ-conditioned coverage (at-least-one, mean, and full) that checks whether each term is recovered under the CQ that requires it.

Sub-task 2b — CQ2Onto

Objective. Given a set of competency questions, generate a full OWL ontology that satisfies them. It tests whether a model recovers implicit and derived terms and expresses them as a coherent ontology. Evaluated over 118 CQs.

Description. Beyond extracting terms, the system produces class and property hierarchies, property characteristics, domain and range axioms, and other restrictions, so that the resulting ontology answers the competency questions.

Input / Output. Input: a set of competency questions. Output: an OWL ontology.

Input · one file per domain <domain>_cq2onto_cqs.json

A JSON list of competency questions, provided by the benchmark in CQ2Onto/competency_question/, in the same id / value format as CQ2Term. The whole set of a domain is the input to one run.

[
  {"id": "CQ1", "value": "Which animal eats which other animal?"},
  {"id": "CQ2", "value": "Is [this animal] a herbivore?"},
  {"id": "CQ6", "value": "Which plants eat animals?"}
]
Output · one file per domain <domain>_ontology.owl

An ontology written in the OWL 2 language (W3C OWL 2 Web Ontology Language), one file per domain, generated from the competency questions of that domain. Any RDF serialisation that rdflib can parse is accepted (RDF/XML as .owl, .rdf); the file name must start with the domain identifier.

Evaluation and Metrics. Evaluated against the CQ-aligned reference ontology across five targets: term recovery, property characteristics, domain/range triples, TBox axioms, and hierarchy closure. Each reports Precision, Recall, and F1 (global and alignment-conditioned views where applicable), with CQ-conditioned coverage at the axiom level and after HermiT reasoning-based closure recovery.