
Gaélique et breton : Même combat face à l'IA
Réalisation : ABP – 116 vues
Tomás Ó Síocháin, un Irlandais a présenté son analyse en français. Il est le CEO at Údarás na Gaeltachta une agence du gouvernement irlandais pour promouvoir le gaélique. Intervention à la fin du député Paul Molac.
A language can survive among people and disappear from machines. In the face of the rise of artificial intelligence, Irish and Breton must now have vast digital resources to avoid being marginalized by new technologies like AI.
Artificial intelligence represents a tremendous opportunity for minority languages, but it could also accelerate their disappearance from the digital world. This is the warning issued by Tomás Ó Síocháin, the director general of Údarás na Gaeltachta, the Irish public agency responsible for the development of Gaelic-speaking regions, during the **Interceltic Business Forum** in Lorient. The issue is no longer just about saving a language by increasing the number of its speakers: **it is now necessary to provide it with enough digital data for it to exist in machines**. Large AI models are trained on vast corpora of texts, documents, and recordings in which English and major languages are overrepresented. Languages with few digital resources are therefore at risk of being left behind.
Who owns the corpora?
Ireland has already built significant reserves of digital data in Irish. But Ó Síocháin raises a second question: should these corpora be entrusted to American AI giants? He cites two different strategies. Iceland has provided data to OpenAI to improve the performance of its models in Icelandic, while the Māori of New Zealand have prioritized sovereignty over their data and the development of their own tools. Údarás is currently pursuing both paths: working with major platforms while supporting local and European solutions. It evaluates the performance of the main AI models in Irish each month and is experimenting with HubChat, an assistant intended for public services.
For Ó Síocháin, it is not enough for an artificial intelligence to produce approximately correct Irish. During the questions, he emphasizes that a solution that is simply "good enough" is not sufficient for a minority language: the models must also convey the context, nuances, and culture specific to the language, instead of mechanically transposing a worldview learned from a dominant language. **He hopes that the European Union will take more responsibility for this issue and announces a conference on linguistic diversity during the next Irish presidency of the Union.**
Breton facing the same challenge
Speaking at the end of the presentation, Breton deputy Paul Molac immediately drew a parallel with Breton. Spell checkers, GPS, voice recognition, or machine translation: the possibility of using a language daily now depends on its presence in these new tools. He recalls the work carried out by the Public Office of the Breton Language, particularly on voice recognition, which should eventually allow for direct translation of a person speaking Breton into French or English. "It is absolutely necessary if we want our languages to endure [...] for the coming century," says the deputy, who links biodiversity to "glottodiversity," the diversity of languages, which is also threatened today.
Building the corpora requires resources
The problem has been known for several years in Brittany. As early as 2023, also during the Interceltic Festival in Lorient, linguist Mélanie Jouitteau, a researcher at CNRS, sounded the alarm the effort must focus on collecting texts, videos, and recordings in Breton to build the corpora essential for new artificial intelligence applications. However, the project for a sustainable corpus platform presented by Mélanie Jouitteau and Reun Bideault to the General Delegation for the French Language and Languages of France (DGLFLF), attached to the Ministry of Culture, was not selected for funding.
Digitizing books and archives, collecting and transcribing thousands of hours of speech, organizing and making this data accessible represents a considerable amount of work. Other governments have recognized the stakes: the Spanish state and the Basque government https://abp.bzh/l-intelligence-artificielle-apprendra-le-basque-un-plan-de-10-5-mi-75761 have invested 10.5 million euros in a program aimed at developing the digital linguistic resources of Euskara.
Yann Le Cun made the same observation as early as 2023. In a conversation with ABP, he suggested "convincing the Minister of Culture to collect texts in regional languages and to fund the training of an open-source LLM with this data."
If we want Breton to exist in the artificial intelligence of tomorrow, the creation of sufficiently large corpora thus becomes a public policy issue.
evezhiadennoù (0)
Evezhiadenn ebet c'hoazh. Bezit ar re gentañ o reiñ ho soñj !