Abstract
Tables are crucial containers of information, but understanding their meaning may be challenging. Over the years, there has been a surge in interest in data-driven approaches based on deep learning that have increasingly been combined with heuristic-based ones. In the last period, the advent of Large Language Models (LLMs) has led to a new category of approaches for table annotation. However, these approaches have not been consistently evaluated on a common ground, making evaluation and comparison difficult. This work uniquely compares Semantic Table Interpretation (STI) approaches with generative and encoder-only LLMs on diverse datasets. In particular, we conduct an extensive evaluation of four STI state-of-the-art (SOTA) approaches — Alligator (formerly s-elBat), TURL, TableLlama, and DAGOBAH (the latter with a partial evaluation due to its high computational demands); Alligator and DAGOBAH belong to the family of heuristic-based algorithms, while TURL and TableLlama are respectively encoder-only and decoder-only LLMs. We also include in the evaluation both GPT-4o and GPT-4o-mini, since they excel in various public benchmarks. The primary objective is to measure the ability of these approaches to solve the entity disambiguation task concerning both the performance achieved on a common-ground evaluation setting and the computational and cost requirements involved, either monetary or in terms of computational resources, with the ultimate aim of charting new research paths in the field.