PhD defense Martin Pekar Christensen

#Data#Database #Data #lake
Share

This seminar is about the PhD defense by Martin Pekar Christense



  Date and Time

  Location

  Hosts

  Registration



  • Add_To_Calendar_icon Add Event to Calendar
  • Selma Lagerløfs Vej 300
  • Aalborg East, Nordjyllands Amt
  • Denmark 9220
  • Building: SLV 300
  • Room Number: room 0.2.13

  • Contact Event Host
  • Co-sponsored by Sean Bin Yang


  Speakers

Christensen of Aalborg University

Topic:

Core Techniques for Semantic Data Lake Systems

Data lakes have in recent years significantly proliferated due to the limitations of the wellestablished concept of data warehouses and the need to store heterogeneous, raw data. Where data warehouses require enforcing a schema on data, data lakes are more flexible and store heterogeneous structured, semi structured, and unstructured data in their unprocessed format. However, the flexibility of data lakes poses many challenges with respect to data integration and discovery, i.e., identifying the data relevant to a task. Data discovery therefore poses a main bottleneck in data lakes. In data discovery, the aim is to identify a large set of relevant datasets for a specific task. However, current approaches are limited by their paradigm, and hence, do not sufficiently retrieve all relevant datasets. Therefore, this thesis addresses this problem with a semantics-aware approach that uses a reference knowledge graph to enable entity-centric, example-driven data discovery. We propose a novel definition of a Semantic Data Lake, its architecture, and a semantic table search system, Thetis, for discovering data lake tables that are semantically relevant to an example table query. We address scalability by proposing a search space prefiltering technique based on localitysensitive hashing that significantly improves search runtime. To evaluate semantic table search approaches, we propose a novel real-world, large-scale benchmark consisting of more than 200K and 400K Wikipedia tables from 2013 and 2019, respectively. We have automatically extracted more than 9K and 2K table queries of various sizes in terms of number of tuples with more than 600K and 900K ground truth relevance annotations. Thetis complementing BM25 keyword search in this benchmark obtains a drastically higher recall compared to pure BM25 key wordsearch, and Thetis alone outperforms other baselines for join and union search, as well as representation learning. The thesis demonstrates Jazero, which implements a full semantic data lake with Thetis semantic table search, and it provides a full graphical user interface. The demonstration highlights the usefulness of Jazero in a real world setting for discovering tables that are semantically relevant for a given task.

Biography:

Martin Pekar Christensen is a researcher in the field of data management and semantic data systems. His work focuses on improving data discovery in large and heterogeneous data lakes through semantic technologies, knowledge graphs, and advanced search techniques. In his PhD research, he has developed methods and systems for semantic table search, including Thetis and Jazero, with a particular emphasis on scalable and effective data discovery.

Address:Selma Lagerløfs Vej 300, , Aalborg East, Denmark, 9220