We present Kenji-Endo, a BabyLM pretrained on a dedicated Italian dataset, which participated in four tasks at the 9th edition of EVALITA: DeSegMa, MultiPRIDE, IMPOLS, and FadeIT. Kenji-Endo achieved competitive performance across all tasks, demonstrating that language modeling with limited data and compact model sizes can represent a viable alternative to Large Language Models.

Kenji-Endo: a BabyLM @EVALITA

Calogero Jerik Scozzaro
First
;
Matteo Rinaldi;Gianluca Mittone;Marco Antonio Stranisci
Last
2026-01-01

Abstract

We present Kenji-Endo, a BabyLM pretrained on a dedicated Italian dataset, which participated in four tasks at the 9th edition of EVALITA: DeSegMa, MultiPRIDE, IMPOLS, and FadeIT. Kenji-Endo achieved competitive performance across all tasks, demonstrating that language modeling with limited data and compact model sizes can represent a viable alternative to Large Language Models.
2026
9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian
Bari, Italia
26-27 febbraio 2026
Proceedings of the 9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026), Bari, Italy, February 26th-27th, 2026
CEUR Workshop Proceedings
4195
0
0
https://ceur-ws.org/Vol-4195/36.pdf
Calogero Jerik Scozzaro , Matteo Rinaldi , Gianluca Mittone , Marco Antonio Stranisci...espandi
File in questo prodotto:
File Dimensione Formato  
36.pdf

Accesso aperto

Tipo di file: PDF EDITORIALE
Dimensione 234.36 kB
Formato Adobe PDF
234.36 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2318/2160232
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 4
  • ???jsp.display-item.citation.isi??? ND
social impact