San Millán
Spanish Oral Corpus
COE San Millán
The key strategic asset for establishing Spanish as a native language in
Artificial Intelligence
WHAT IS IT?
What is the San Millán
Spanish Oral Corpus?
The San Millán Spanish Spoken Corpus (COE San Millán) has been
designed to provide a large-scale, high-quality and technically
advanced linguistic resource that captures the reality of spontaneous speech in all its diversity.
This corpus was created with the aim of establishing a unified corpus of spoken Spanish that will drive the New Language Economy and ensure the competitiveness of Spanish-based artificial intelligence applications.
COE San Millán is shaping the Spanish language for the digital future.
HOW IS IT ORGANISED?
How is it organised?
To ensure that the data serves as ‘premium fuel’ for training Language Models (LLMs), the corpus is structured across three levels:
Layer 1
Transcript layer
It captures exactly what happens in the conversation: pauses, hesitations, interruptions, or people speaking at the same time. Everything is accurately recorded in accordance with the EAGLES standard.
EAGLES
Layer 2
Morpho-syntactic layer
It provides an in-depth analysis based on Universal Dependencies to understand how sentences are actually constructed in everyday speech.
Universal Dependencies
Layer 3
Semantic layer:
Using WordNet / MCR 3.0, the COE San Millán incorporates semantic information that helps interpret the meaning and context of what is being said. In this way, the models do not merely ‘read’ words, but gain a better understanding of what they convey, enabling the models to grasp the meaning and context of speech.
WordNet / MCR 3.0
HOW IS IT BUILT?
How is the San Millán COE built?
The development of the corpus follows a methodology based on subtraction; in other words, an ideal coverage is defined (coverage matrix), material that has already been licensed and catalogued is subtracted (acquisition matrix), and only the identified gaps are recorded.
To achieve this, three acquisition channels are used:
integration of materials through agreements with institutions.
audio recordings from the media, debates or public speeches.
involvement of students and teachers from a global network to ensure a wide range of backgrounds and reduce dialectal bias.
International standards
and sustainability
The San Millán COE has been designed in accordance with international standards
to ensure its durability and legibility:
THE STRATEGIC VALUE OF THE MODEL
The strategic value of the model
This corpus is not merely a database, but a FAIR (Findable, Accessible, Interoperable and Reusable) infrastructure. Its alignment with DCAT-AP profiles ensures its integration into the European Single Market for data. In this way, it positions the Valle de la Lengua Data Space as the global hub for academic and linguistic knowledge on the Spanish language.
Localizable
Accesible
Interoperable
Reutilizable
CORPUS SAMPLE
Download a sample from the corpus
An actual recording excerpt with its corresponding transcript, ready for assessing the quality and format of the data from the San Millán COE.
Recording of an academic speech with a phonetic transcription and annotations
Mexican and Central American Woman Aged 55 or over Reading aloud Higher education
Duration: 16:13 minutes
Audio format: WAV, 44.1 kHz
Transcript format: XML
Layers included: Enriched orthographic and morphological transcription