Last week, I was at the 23rd European Semantic Web Conference (ESWC 2026). This is one of my regular conferences and I’m always happy to see both friends and meet new faces. ESWC provides a space with research that intersects data, machine learning, knowledge representation and their joint applications. You can sit in a session with presentations on datalog and graph queries and go to another on new models for information extraction for KGs. What’s nice is the shared language and technology around knowledge graphs that brings these topics together. The conference had around 270 participants this year. I think that the 250 – 500 range of participants makes for a good event in terms of meeting people while having enough critical mass to have a variety of content.
I noted last year that the community is very welcoming and does its best to bring people together. This year was no different with lots of opportunities in particular for PhD students to participate in terms of both the doctoral symposium but also poster presentations and being part of the minute madness. The addition of daily pre-conference intro sessions for PhD students to kind of plan the day and meet the session chairs was really positive. One of my faults is to always have questions at talks and this year it was nice that PhD students were prioritised in the Q/A. It’s good to make me wait 😀.
INDElab was well represented at the conference:



We had a main track research paper (Are a Thousand Words Better Than a Single Picture? Beyond Images – A Framework for Multi-Modal Knowledge Graph Dataset Enrichment) led by Pengyu. We had two demos:
- Efficient Adaptive Information Extraction (Video, Code, Model)
- A Prototyping Environment for Comparing Neuro-symbolic Diagnostic Assistants (Video, Code)
At the Workshop on Quality of Knowledge Graphs, we had a paper on multimodal tabular anomaly detection (slides). I also presented some preliminary ideas on Explanation as Evaluation and was on the panel at the TAAPAI workshop. Lastly, I gave a keynote at the TRIPLET workshop thinking about the where knowledge should reside in complex AI systems (in the knowledge graph, in the foundation model, in natural languages, ..). More on this later.
I’m thinking about the following themes coming away from the conference:
Long-lived AI systems

Currently, we have Agentic AI systems which run in the hours and maybe at the extreme in days (see METR time horizons). A central concern coming out of the conference was how do we enable Agentic AI systems that run for weeks, months or longer. To do that there is critical need for a knowledge bases that are trustworthy and can evolve overtime. Here, I would point to the excellent comprehensive tutorial given by Estevam Hruschka from Megagon Labs. Returning to the ideas of that he pioneered with Tom Mitchell and others on never ending learning building knowledge bases from the web overtime. (I’m always surprised that we are using the NELL KB to this day to measure link prediction performance, but anyway…). Estevam argued for both trustworthy knowledge underpinned by explicit knowledge bases but with the ability to continuously learn. The tutorial connected principled notions to the latest in developments in automatically building AI agent harnesses (see also meta-harness), agentic memory, and LLMs writing to knowledge bases like wikis (see Karpathy’s LLM wiki and https://github.com/sciknow-io/skillful-alhazen).
This theme of long running AI systems was also part of the keynote of Prof. Maria Esther Vidal. In particular, I liked her notion of semantic ecosystems where AI systems exist in an entire landscape of resources, other agents, data, and tools, governance layers that are changing and evolving.

A thread of research that is core long-lived AI systems is continual learning. At the conference, I enjoyed the paper by Janna Omeliyanenko and colleagues on Parameter Efficient Continual Automated Knowledge Graph Completion. They develop a number of combinations of PEFT where they adapt pre-trained models for both continual learning of relation extraction and link predication for KGs. I haven’t really seen this combination in the KG setting before.
An important aspect of AI Agents is tool use, hence, it’s important that agents are able to easily make use of knowledge graphs. I’d point to two pieces of work: 1) rudof-MCP which makes the Rust-based rudof semantic web tooling library available through the model context protocol; and 2) an extension of the HeFQUIN federated SPARQL query engine to support Web APIs by Olaf Hartig and Jonas Westman (this won the best demo award). You can also now model agentic AI systems using the AgentO ontology for auditing and documentation.
Evaluation
This long-lived agent viewpoint also brings up the problem of how to evaluate such systems. This is a continuation of one of my points from last year about evaluation with respect to the capabilities of GenAI but expands the number of concerns. I enjoyed Elena Simperl’s talk at the ELMKE: Evaluation of Language Models in Knowledge Engineering workshop which focused on evaluation in the context knowledge engineering teams asking what good looks like?

I also learned from her about the work under the heading of open world evaluations. Likewise, Estevam’s tutorial suggests to move towards evaluation throughout the lifecycle:

This is also an argument for why explanations might be a good way to evaluate AI system performance.
Also on the evaluation front, it was good to see quite a lot of work on producing evaluation benchmarks and also critically interrogating the results of in the research field. Some highlights:
- Link Prediction or Perdition: the Seeds of Instability in Knowledge Graph Embeddings (PDF) – our evaluation metrics for KG embeddings are not capturing the huge dependence on initial conditions. We need better evaluation protocols. The best paper award winner.
- Bench4KE: Benchmarking Automated Competency Question Generation (PDF) – best resources paper
- HERTy-Wiki: A Benchmark for Hierarchical Entity Reasoning and Typing (PDF)
- CUBE-MT: A Cultural Benchmark for Multimodal Knowledge Graph Construction with Generative Models (PDF)
- Evaluating Knowledge Graph Construction from Text Without Supervision (PDF) – a metric similar to bertscore but defined for the KG context
- Ontology Population Using LLMs: Which Factors Matter? (PDF) – what it says in the title.
- BOLD: A Simulation Framework for Dynamic Linked Data Environments and a Benchmark for Linked Data User Agents (repo)
Raphael Troncy’s excellent keynote at the Text2KG workshop highlighted the need for better evaluation as we move from just extracting triples to constructing situated, evidence-grounded, schema (ontology)-aware, multimodal, knowledge graphs. He ended with a pertinent question: what does it mean for a generated KG to be correct?
Where should knowledge reside
In our group, we’ve been thinking about where knowledge resides. For example, Jan-Christoph has been studying with the community large language models as knowledge bases, for example, by running challenges that explore what can be retrieved from the parametric knowledge of LLMs. Brad has been integrating foundation models as source of data in neurosymbolic reasoning. Likewise, our paper at ESWC was looking out how by generating textual representations of images in multimodal knowledge graphs we could improve link prediction performance. Here, we use the encoders on attributes of KGs to include image and textual information in the embedding representation of the KG. This spectrum between KG/database – multimodal KG – foundation model parameters was what I was contemplating in my talk “Leave the modalities alone? When, where and if we should extract knowledge graphs from multiple modalities. (slides).
The importance of this consideration was reinforced at the conference. In particular, the excellent keynote from Isabelle Augenstein: “Understanding LLMs’ Utilisation of Parametric and Contextual Knowledge” (Slides). This looked not at knowledge just in the parameters of a language model but the knowledge conveyed in context and asking which knowledge is actually used for example in answering questions especially in the RAG context:


Second, I had a number of conversations around this theme especially with the prevalence of both multimodal knowledge graphs and the success of GraphRAG and the wider discussion of the importance of semantics for agentic AI as highlighted by Atanas Kiryakov of Graphwise in his keynote (slides):


Cyber-Physical Systems & Knowledge Graphs
Atanas keynote also reminded us of the success of semantic technologies and in particular knowledge graphs for a range of companies and verticals:


One vertical he emphasised was what I would categorise as cyber-physical systems where you have a need for both precision, external knowledge about the domain and the need for interoperability across systems makes knowledge graphs an excellent fit. He showed the example of the Norwegian Electricity Grid Operator (above) and how their semantic digital twin enables natural language question answering over their grid data. Our demo was also about building small scale cyber-physical systems in order to provide an environment to prototype agent based diagnostic systems:


Cyber-physical systems were throughout the conference. I liked the paper by Ivanonvic et al. on exposing the CAN bus data (the communication line that goes through many systems like boats and cars) as a knowledge graph and using that for a variety of diagnostic scenarios. Semantics on a boat 🛥️ is cool! Knowledge graphs are helping with the qualification of components for satellites at ThalesAlenia Space by helping solve data integration challenges such as inconsistent naming, buried identifiers, different schemas, and inconsistent / outdated records. KGs are being used to monitor water systems and at Bosch for helping developing micro-electronic mechanical systems (MEMS). I was also happy to discuss with this year’s general chair Prof. Maribel Acosta on her large project (CONVIDE) on using KGs and SHACL to help with the development of cyber physical systems.
Wrap-up
There was quite some interesting results I didn’t cover here ranging from multimodality to neurosymbolic systems. Also I would be a remise not to mention that competency questions are a thing! This was the first year of ESWC being under the SWSA umbrella and I think it’s already showing results with transfer of reviews from ISWC to ESWC for resubmissions. I’m looking forward to how the two main conferences will be working together in the future. You should follow the new SWSA LinkedIn page to keep up to date.
Random Notes
- Cool work from Anna Lisa and a team from IBM Research showing how to use Wikidata + Wikipedia together to generate conversations that improve the reasoning performance of Vision Language Models. This is build into the data pipeline fro producing the IBM Granite VLM family of models.
- For Antonis, David, Hazar check out the slides at Quality for Knowledge Graphs workshop focused a lot on FAIR data checking and SHACL. Proceedings here.
- DuckDB was popping up at the conference. bikiDATA uses it in the backend to query RDF data. The CAN bus paper I mentioned uses it to store data before exposing it as a KG. I’m quietly convinced that Hannes is implementing a whole alternative semantic web stack in the DuckDB ecosystem: a db in a web browser with WASM, a protocol for communicating data over HTTP (quack), and a shared metadata layer (DuckLake).
- PROV, nanopublications and RO-Crates all in one paper – how can I not like it.
- openlineage.io
- Don’t worry world models were talked about.
- Datalog and SPARQL and RUST. I need to check out the Nemo datalog based rule reasoner.
- Congrats to the SemDH workshop team including our own Lise Stork on winning the Dieter Fensel Visionary Contribution Award.
- Open Data + ML: the Dubrovnik visitor tracker
- Dubrovnik is pretty amazing as a city. Something to see.
- KG based data integration pipeline tutorial
- https://github.com/ANR-kFLOW/CausalSense for Event Relation Extraction.
- Kudos to the organisation team for putting together such an excellent set of research and a fun event.















































Good job!




It was quite good, I hope they put the slides up somewhere. The key notion for me is this idea of instance optimality: by using machine learning we can tailor performance to specific users and applications whereas in the past this was not cost effective because the need for programmer effort. They suggested 4 ways to create instance optimized algorithms and data structures:







