AI Navigate

Llettuce: An Open Source Natural Language Processing Tool for the Translation of Medical Terms into Uniform Clinical Encoding

arXiv cs.CL / 3/13/2026

📰 NewsTools & Practical UsageModels & Research

Key Points

  • Llettuce is an open-source NLP tool designed to translate informal medical terms into OMOP standard concepts.
  • It leverages large language models and fuzzy matching to automate mapping and improve accuracy over Athena and Usagi, which require manual input.
  • The tool emphasizes GDPR compliance and can be deployed locally to protect data in healthcare workflows.
  • It aims to streamline clinical terminology mapping and enhance on-premises data encoding capabilities.

Abstract

This paper introduces Llettuce, an open-source tool designed to address the complexities of converting medical terms into OMOP standard concepts. Unlike existing solutions such as the Athena database search and Usagi, which struggle with semantic nuances and require substantial manual input, Llettuce leverages advanced natural language processing, including large language models and fuzzy matching, to automate and enhance the mapping process. Developed with a focus on GDPR compliance, Llettuce can be deployed locally, ensuring data protection while maintaining high performance in converting informal medical terms to standardised concepts.