Skip to main content

Classification of input texts with predefined labels from the EuroVoc thesaurus

European Parliament

Use case ID: 013

Author: European Parliament

Date: 5 June 2024

Objective:

The Innovation, Intranets and Digital Solutions Unit (INNOVIT) team fine-tuned this EUBERT-based automatic indexing tool in order to allow it to classify input texts with predefined labels from the European Union’s (EU) EuroVoc thesaurus.

 

EuroVoc is a large, multidisciplinary, multilingual (24 languages of the EU), hierarchical thesaurus of more than 7,000 classes covering the activities of EU institutions. Given the number of legal documents produced every day and the huge mass of pre-existing documents to be classified, high-quality automated or semi-automated classification methods are most welcome in this domain.

 

This model, based on the BERT deep neural network, was trained on more than 3,200,000 documents to achieve that task and is used in a production environment via the HuggingFace inference endpoint. This model supports the 24 languages of the EU.

Actors:

  • Innovation, Intranets and Digital Solutions Unit (INNOVIT)
  • Directorate for Publishing, Innovation and Data Management
  • Directorate-General for Innovation and Technological Support (DG ITEC)

 Prerequisites:

  • No hardware or software requirements if the end-user interface is used
  • Coding and package requirements if the model is integrated into an existing solution

 Scenario:

  1. The end user inputs a document to classify.
  2. The system processes the input document and outputs a set of predefined labels with corresponding confidence scores.

 Alternate flows:

  • If the document is in a language not supported by the model, the system will flag it as unsupported.

Expected results:

  • The system provides accurate classification of legal and policy documents based on the EuroVoc thesaurus.
  • Document classification within EU institutions is more efficient and accurate.

 Potential challenges:

  • Ensuring model accuracy across all 24 languages
  • Handling ambiguous or poorly structured text inputs
  • Maintaining performance and speed with large volumes of documents

 Data requirements:

  • Input: text documents in any of the 24 supported languages of the EU
  • Predefined EuroVoc labels for classification (the EuroVoc thesaurus was developed by the European Parliament (EP), in collaboration with the Publications Office of the EU (OP), and has more than 7,000 classes)
  • Historical data for model training and fine-tuning (EP Public Register of Documents)

 Integrations with other systems:

  • The model is integrated in some of the applications developed in-house, such as Monitor Partners’ Interest, a content aggregation solution for scraping, translating and indexing websites and sources of interest.

Success metrics:

  • Micro F1 score (threshold value: 0.46)
  • High normalized discounted cumulative gain (NDCG) scores
  • User satisfaction and feedback from end users regarding classification accuracy and relevance

 

The Use cases for AI in parliaments collection is published by the IPU’s Centre for Innovation in Parliament as part of the Parliamentary Data Science Hub’s project to create guidelines for AI governance in parliaments.

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International licence. It may be freely shared and reused with acknowledgement of the author and the IPU. 

A use case describes how a system should work. It is used to plan, develop and measure implementation. A use case is not the same as a case study, which is a descriptive text of an actual project’s implementation. Please note that this use case is provided “as is” and neither the IPU nor the author accepts any responsibility for its use.

For more information about the IPU’s work on artificial intelligence, please visit www.ipu.org/AI or contact [email protected]