Skip to content

Latest commit

 

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

🌿 BASIL - Broad AI Safety for Italian Language

A curated collection of publications, datasets, and models focused on broad AI safety in the Italian language.

AI safety research remains overwhelmingly English-centric, leaving Italian comparatively under-evaluated. This list collects the resources that specifically address Italian: safety benchmarks, guardrail models, and the papers that support them.

The repository aims to provide a focused starting point for researchers working on Italian-language AI safety and to make the growing body of Italian-specific work easier to discover and reuse.

Note

This collection is incomplete and currently under construction. The listings below are not exhaustive: entries are still being gathered and reviewed, and sections, links, and metadata may change as the repository evolves.

Scope

This repository focuses on broad, holistic AI safety in Italian, covering resources that address multiple dimensions of safety, such as harmful content, bias, toxicity, privacy, security, and robustness. Works narrowly focused on a single dimension are currently out of scope.

Publications

2026

  • The Effects of Benevolent Fine-tuning on the Safety of the Italian Large Language Models — Pulerà et al., CLIC-it 2026 [PDF]
  • “Capisci a me”: The Hidden Risks of Regional Language Processing in LLMs — Magazzù et al., CLIC-it 2026 [PDF] [Code]
  • Mind the Language Gap: Assessing LLM Safety in Italian — Marafatto & Navigli, LREC 2026 [PDF] [Poster] [Code]
  • AI Safety Lost in Translation: Evaluating the Effectiveness of English-Italian Cross-Lingual LLM Safety Alignment — Wu & Brandao, LREC 2026 [PDF]
  • Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection — Giarrusso et al., IASEAI 2026 [PDF]
  • "Learning from Mistakes: Can LLM Self-recover after Misalignment?": Examining the LLM's ability to self-recover after misalignment — Sorokoletova et. al., AAAI26 WS37 [PDF]
  • Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study — Mungari, arXiv 2026 [PDF] [Code]

2025

  • BeaverTails-IT: Towards A Safety Benchmark for Evaluating Italian Large Language Models — Magazzù et al., CLiC-it 2025 [PDF]
  • Uncovering Unsafety Traits in Italian Language Models — Rizzi et al., CLiC-it 2025 [PDF]

2024

  • Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models — Pernisi et al., ACL 2024 [PDF] [Code]
  • Multi-property steering of large language models with dynamic activation composition — Scalena et al., BlackboxNLP 2024 [PDF] [Code]

Datasets

Dataset Year Description Languages Links License
SafeLLM-it 2026 Culturally grounded Italian safety benchmark of 1,762 manually categorized Italian Wikipedia pages and 5,286 generated prompts probing refusal behavior, moderation consistency, and unsafe content generation. Italian GitHub Paper CC BY-NC-SA 4.0
MMLU-Redux-IT 2026 MMLU-Redux translated into Italian and local languages to study how dialectal input affects model understanding and safety. Italian, Venetian, Lombard, Friulian, Ligurian, Sicilian Hugging Face Paper CC BY 4.0
XSTest-IT 2026 XSTest prompts translated into Italian and local languages, preserving the safe/unsafe contrast labels for dialectal safety evaluation. Italian, Venetian, Lombard, Friulian, Ligurian, Sicilian Hugging Face Paper CC BY 4.0
BeaverTails-IT 2025 Italian machine-translated version of BeaverTails, with parallel translations from five state-of-the-art MT models; 330k question-answer pairs. Italian Hugging Face Paper CC BY-NC 4.0
BeaverTails-IT-Evaluation 2025 Italian machine-translated version of BeaverTails-Evaluation: 700 prompts spanning 14 safety categories, used with fine-tuned classifiers and human judgment to assess seven Italian LLMs. Italian Hugging Face Paper CC BY-NC 4.0

Models

BeaverTails Safety Classifiers:

Multilingual Resources (coming soon)

Multilingual resources for broad AI safety that include Italian will be added in the future. This section will cover datasets, models, and publications that address AI safety across multiple languages and would be relevant and useful for Italian language research.

Works that merely include Italian as one of many languages will not be listed here, unless they provide specific insights or evaluations for Italian.

Related Lists and Resources

Here are some international lists and resources of AI safety that may be of interest:


Icons by Simple Icons and Iconify.

About

A Collection of Research Papers on AI Safety for Italian Language

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors