Seminars - page 3
The DIG team holds a seminar about every two weeks with speakers either from the team, or invited.
You can add the seminars to your calendar with this ics file, and get emails about future seminars by subscribing to our mailing-list.
If you would like to present your work at our seminar, please contact Nils.
Upcoming Seminars
Toward Adaptive Intelligence
Tuesday, October 13, 2026 11:45, 4A301
Nilesh Verma (University of Waikato)
Real-world data arrives as an endless stream whose distribution shifts over time. Models that work today may degrade tomorrow, and the human effort required to repeatedly re-tune them does not scale. This talk presents a line of work on adaptive intelligence, where systems select, configure, and revise their own learning pipelines as the data changes. I begin with AutoML for non-stationary data streams, covering online pipeline search, automated drift and outlier handling, and meta-learning for data streams. I then turn to tabular foundation models and ask what happens when in-context learners encounter concept drift. Along the way, I introduce TuiML, an open-source MCP-native ML runtime that enables agents and researchers to run, benchmark, and empirically study machine learning workflows.
From Observations to Explanations: Exploring Methodologies for Identifying Relevant Actual Causes in Complex Systems
Tuesday, October 27, 2026 11:45, 4A301
Samuel Reyd
Complex systems such as multi-agent and cyber-physical systems exhibit emergent behaviors that are hard to understand and explain. Causal models offer a principled framework for this, but building one for a complex system is already difficult: the relevant scope, abstraction scale, and variables are not given in advance. Moreover, such models capture general causation, i.e., which factors influence a class of outcomes, whereas a user asking “why did this happen?” needs an actual cause: the specific facts responsible for a specific event. Actual causation is conceptually and computationally demanding; finding actual causes is intractable in general, and not all causes are equally informative. Producing a useful explanation therefore requires three steps: modeling the system causally, identifying its actual causes, and filtering them for relevance. This model-identify-filter pipeline forms the backbone of the thesis. This thesis first reviews the challenges of causal modeling in complex adaptive systems and examines the choice of abstraction scale through causal emergence, showing that emergence detection depends heavily on the sampling distribution used. It then addresses the identification of actual causes, proposing a data-based method for Markovian multivariate systems and approximate search algorithms with adjustable precision that recover the full set of actual causes in general SCMs, released as an open-source Python module. Finally, it presents a general, tunable method for selecting relevant causes according to user and context, and develops a normality criterion to produce more relevant explanations. Overall, the thesis connects formal causal models, scalable computation, and human-centred explanation into a coherent framework for context-aware explanation in complex systems.
Past Seminars
GPTKB: Comprehensively Materializing Factual LLM Knowledge
Tuesday, April 29, 2025 11:45, 4A301
Simon Razniewski (TU Dresden)
LLMs have majorly advanced NLP and AI, and next to their ability to perform a wide range of procedural tasks, a major success factor is their internalized factual knowledge. Since (Petroni et al., 2019), analyzing this knowledge has gained attention. However, most approaches investigate one question at a time via modest-sized pre-defined samples, introducing an “availability bias” (Tversky and Kahneman, 1973) that prevents the discovery of knowledge (or beliefs) of LLMs beyond the experimenter’s predisposition. To address this challenge, we propose a novel methodology to comprehensively materialize an LLM’s factual knowledge through recursive querying and result consolidation. As a prototype, we employ GPT-4o-mini to construct GPTKB, a large-scale knowledge base (KB) comprising 101 million triples for over 2.9 million entities. This work marks a milestone in two areas: For LLM research, for the first time, it provides constructive insights into the scope and structure of LLMs’ knowledge (or beliefs), and its strengths and weaknesses. For KB construction, it pioneers new pathways for the long-standing challenge of general-domain KB construction. GPTKB is accessible at https://gptkb.org.
ProvSQL: Provenance and Probabilistic Querying in Uncertain Databases
Tuesday, April 08, 2025 11:45, 4A125
Pratik Karmakar (None)
Probabilistic databases provide a powerful framework for managing and querying uncertain data, enabling principled reasoning under uncertainty. ProvSQL extends PostgreSQL to support provenance tracking and probability computation in probabilistic databases, leveraging provenance circuits to efficiently compute probabilities and Shapley-based data valuations. In this talk, we introduce ProvSQL, demonstrate its capabilities, and explore a key use case—content based image retrieval from the COCO dataset. We show how probabilistic query evaluation and data valuation techniques enhance explainability and trust in AI-driven decision-making.
Tabular foundation models: priors for numbers and strings
Tuesday, March 25, 2025 11:45, 4A301
Gaël Varoquaux (INRIA)
Deep-learning typically does not outperform tree-based models on tabular data. Often this may be explained by the small size of such datasets. For images, sound, text, the solution has be pretrained models, leading to foundation models, adapted and reused for many tasks. I will discuss the challenges to bring these ideas to tabular learning, and the progress that we have made, building priors for tables, ie columns of different natures, with numbers and strings.
Neuro-symbolic approaches for the knowledge graph lifecycle
Tuesday, March 18, 2025 11:45, 4A301
Pierre Monnin (INRIA)
In the Web of Data, an increasing number of knowledge graphs (KGs) are concurrently published, edited, and accessed by human and software agents. Their wide adoption makes essential the tasks of their lifecycle: construction, refinement (e.g., matching, link prediction), mining, and usage to support applications (e.g., explainable AI, recommender systems). However, all these tasks require facing the inherent heterogeneity of KGs, e.g., in terms of granularities, vocabularies, and completeness. Besides, scalability issues arise due to their increasing size and combinatorial nature. In my talk, I will present my research on neuro-symbolic approaches for the KG lifecycle, intertwining domain knowledge from ontologies, deductive reasoning, analogical reasoning, and machine learning models. Throughout my presentation, I will show that such approaches enhance models by improving their semantic awareness, frugality, and the semantic interpretability of their latent representation space.
None
Tuesday, March 04, 2025 11:45, 4A301
Ken Satoh (None)
None
None
Tuesday, February 04, 2025 11:45, 4A125
Fabian (None)
None
None
Tuesday, January 21, 2025 11:45, 4A301
Simon Delarue (None)
None
None
Tuesday, December 10, 2024 11:45, 4A125
Lanfang Kong (None)
None
None
Tuesday, December 03, 2024 11:45, 4A125
Gabriel Damay (None)
None
None
Tuesday, November 12, 2024 11:45, 4A125
Cyril Chhun (None)
None