San Millán 
Spanish Oral Corpus

COE San Millán

The key strategic asset for establishing Spanish as a native language in 
Artificial Intelligence

WHAT IS IT?

What is the San Millán 
Spanish Oral Corpus?

The San Millán Spanish Spoken Corpus (COE San Millán) has been 
designed to provide a large-scale, high-quality and technically 
advanced linguistic resource that captures the reality of spontaneous speech in all its diversity.

This corpus was created with the aim of establishing a unified corpus of spoken Spanish that will drive the New Language Economy and ensure the competitiveness of Spanish-based artificial intelligence applications.

 

COE San Millán is shaping the Spanish language for the digital future.

 

HOW IS IT ORGANISED?

How is it organised?

To ensure that the data serves as ‘premium fuel’ for training Language Models (LLMs), the corpus is structured across three levels:

Layer 1

Transcript layer

It captures exactly what happens in the conversation: pauses, hesitations, interruptions, or people speaking at the same time. Everything is accurately recorded in accordance with the EAGLES standard.

EAGLES

Layer 2

Morpho-syntactic layer

It provides an in-depth analysis based on Universal Dependencies to understand how sentences are actually constructed in everyday speech.

Universal Dependencies

Layer 3

Semantic layer:

Using WordNet / MCR 3.0, the COE San Millán incorporates semantic information that helps interpret the meaning and context of what is being said. In this way, the models do not merely ‘read’ words, but gain a better understanding of what they convey, enabling the models to grasp the meaning and context of speech.

WordNet / MCR 3.0

HOW IS IT BUILT?

How is the San Millán COE built?

The development of the corpus follows a methodology based on subtraction; in other words, an ideal coverage is defined (coverage matrix), material that has already been licensed and catalogued is subtracted (acquisition matrix), and only the identified gaps are recorded.

To achieve this, three acquisition channels are used:

Existing corpus

integration of materials through agreements with institutions.

Public sources:

audio recordings from the media, debates or public speeches.

Our own recordings (UNIR model):

involvement of students and teachers from a global network to ensure a wide range of backgrounds and reduce dialectal bias.

International standards 
and sustainability

The San Millán COE has been designed in accordance with international standards 
to ensure its durability and legibility:

TEI P5 container (XML)

Regarded as the “gold standard” for preservation, it combines synchronised audio, text and annotations into a single file.

CMDI metadata (ISO 24622-1)

They allow for detailed searches (by age, region or gender of the speaker) without the need to download the entire corpus.

International connection

The corpus is integrated into the CLARIN ecosystem via the Virtual Language Observatory (VLO), ensuring its visibility to researchers across the European Union.

THE STRATEGIC VALUE OF THE MODEL

The strategic value of the model

This corpus is not merely a database, but a FAIR (Findable, Accessible, Interoperable and Reusable) infrastructure. Its alignment with DCAT-AP profiles ensures its integration into the European Single Market for data. In this way, it positions the Valle de la Lengua Data Space as the global hub for academic and linguistic knowledge on the Spanish language.

Findable

Localizable

Accessible

Accesible

Interoperable

Interoperable

Reusable

Reutilizable

CORPUS SAMPLE

Download a sample from the corpus

An actual recording excerpt with its corresponding transcript, ready for assessing the quality and format of the data from the San Millán COE.

Central American Mexican dialectal variety

Recording of an academic speech with a phonetic transcription and annotations
Mexican and Central American Woman Aged 55 or over Reading aloud Higher education

 

Duration: 16:13 minutes
Audio format: WAV, 44.1 kHz
Transcript format: XML
Layers included: Enriched orthographic and morphological transcription
 

Download audio (.wav) Download transcript (.xml)