NL2UNIFOL: From Natural Language Sentences to Uniform First-Order Logic Formulae

   page       BibTeX_logo.png       attach   
Matteo Magnini, Davide Liga, Luca Pasetto
Ha Thanh Nguyen, Francesca Toni, Kostas Stathis, Ken Satoh, Randy Goebel, Francesco Chiariello, Yves Lespérance, Matteo Magnini, Federico Sabbatini, Elena Umili, Nourhan Ehab, Mervat Abu-Elkheir (a cura di)
Proceedings of the Joint Workshop on Statistics and Knowledge Integration for Logic, Learning, Ethical Decisions, and LLMs (SKILLED-LLMs 2026) co-located with the Federated Logic Conference 2026 (FLoC 2026), Lisbon, Portugal, July 18, 2026, pp. 126–143
CEUR Workshop Proceedings 4229
Sun SITE Central Europe, RWTH Aachen University
2026

Automatic translation of Natural Language (NL) sentences into logic representation – such as First-Order Logic (FOL) formulae – is a task of great interest for many communities, including Knowledge Representation (KR), Natural Language Processing (NLP), normative Artificial Intelligence (AI), and AI in general. Recently, an increasing number of works have explored the use of Large Language Models (LLMs) to translate NL into FOL, with promising results. However, existing works mostly focus on the translation of a single NL sentence at a time, resulting in independent formulae that do not share a common predicate and constant name space. This work presents a new framework – namely Natural Language to Uniform First Order Logic (NL2UNIFOL) – that leverages LLMs to translate NL sentences into a uniform FOL theory, where formulae share the same vocabulary. NL2UNIFOL is an end-to-end pipeline that: (i) translates NL sentences into preliminary FOL formulae; (ii) identifies and safely merges predicate and constant names to obtain a uniform vocabulary; (iii) finally generates the final FOL theory. We use NL-FOL pairs from the MALLS dataset to validate our framework. Results show that in the majority of cases NL2UNIFOL is able to correctly create clusters of predicate and constant names to be merged together, which allows to obtain a uniform FOL theory. We also recognise that there are still a numerous amount of cases in which the represented name and the representative name semantically diverge, which motivates future work to further improve the symbol clustering process.

parole chiave   Natural Language Processing, First-Order Logic, Large Language Models, Knowledge Representation
evento origine
world SKILLED-LLMs 2026 @ FLoC 2026
rivista o collana
book CEUR Workshop Proceedings (CEUR-WS.org)