Skip to main content

Text classification based on multiple thesauruses

Italy - Senate

Use case ID: 028

Author: Senate of Italy

Date: 12 June 2024 

 

Objective: 

Automatically classify parliamentary texts according to multiple thesauruses (such as TESEO and Eurovoc), enhancing the organization, searchability and retrieval of information.

Actors: 

  • Senate document management and research teams
  • Senate IT and web development team
  • Artificial intelligence (AI) system for text classification

Prerequisites: 

  • Existing database of parliamentary texts
  • Trained large language model (LLM)-based AI model for text classification
  • Access to thesauruses such as TESEO and Eurovoc
  • Internet accessibility for users

Scenario: 

  1. The user accesses a parliamentary text on the Senate website.
  2. The user initiates the classification process by clicking a button or issuing a command.
  3. The LLM-based AI model processes the text to classify it according to the selected thesauruses (e.g. TESEO, Eurovoc).
  4. The system suggests appropriate categories or descriptors for the text from the thesauruses, explaining the motivations behind the suggestions.
  5. The user can choose whether or not to approve the suggested descriptors.
  6. The classified text is displayed with the approved thesaurus categories marked or listed.
  7. Users can search and filter documents based on these classifications.

Alternate flows: 

  • If the AI model cannot confidently classify the text, it prompts the user for manual verification and correction.
  • For texts that span multiple categories, the system may offer detailed subclassification.

Expected results: 

  • Accuracy and efficiency in classifying parliamentary texts according to various thesauruses is improved.
  • The organization and searchability of parliamentary documents are enhanced.
  • The time and effort required for manual classification is reduced.
  • Consistency and reliability in the application of thesaurus categories are increased.

Potential challenges: 

  • Ensuring the LLM-based AI model can accurately classify diverse types of texts across different thesauruses
  • Handling documents with complex language or multiple relevant categories
  • Continuously updating the LLM-based AI model to adapt to changes in thesauruses and new document types

Data requirements: 

  • Historical texts with annotated thesaurus classifications for training and improving the LLM-based AI model
  • Current database of parliamentary texts
  • Access to updated versions of thesauruses (TESEO, Eurovoc, etc.)

Integrations with other systems: 

  • Existing Senate website infrastructure
  • LLM-based AI processing systems and models
  • Thesauruses databases (TESEO, Eurovoc, etc.)

Success metrics: 

  • Accuracy rate of text classifications
  • User satisfaction ratings and feedback
  • Reduction in time spent on manual text classification
  • Increase in the number of texts classified automatically
  • Consistency and reliability of classifications across different documents

 

The Use cases for AI in parliaments collection is published by the IPU’s Centre for Innovation in Parliament as part of the Parliamentary Data Science Hub’s project to create guidelines for AI governance in parliaments.

 

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International licence. It may be freely shared and reused with acknowledgement of the author and the IPU. 

 

A use case describes how a system should work. It is used to plan, develop and measure implementation. A use case is not the same as a case study, which is a descriptive text of an actual project’s implementation. Please note that this use case is provided “as is” and neither the IPU nor the author accepts any responsibility for its use.

 

For more information about the IPU’s work on artificial intelligence, please visit www.ipu.org/AI or contact [email protected]