Specialised terminology dictionaries in the Valle de la Lengua Data Space

USE CASE

A multi-sectoral repository that takes specialised terminology from international reference sources — established in English — and translates it into the eight major varieties of Spanish, using real speech to document how each concept is referred to in medicine, law, technology and the environment.

One concept, one international standard.  
 Eight ways to refer to it in Spanish.

Specialist terminology worldwide is established in English: organisations such as the WHO, the IEEE and the ISO define the concept and its standard term. When that term is translated into Spanish, the major reference sources offer a single equivalent.

But the same concept is not referred to in the same way in Mexico, Chile, the Caribbean or Spain. That final stage — from the standardised term to the word actually used in each variety — is not documented by any international resource. And that is where this repository comes in.

"Influenza" in the textbook. 
And in the doctor’s surgery? 
Flu, the flu, a cold – depending on who you ask.

This is what the Valle de la Lengua Data Space repository documents: from the technical reference term to the word that is actually used in each variety of Spanish.

+500M

native Spanish speakers around the world

8

documented dialectal macro-varieties

4

specialist sectors covered

+300

terms collected and documented

WHAT IS IT?

From the international standard 
to spoken Spanish

The repository is based on standardised specialist terminology — as established in English by organisations such as the WHO, the IEEE and the ISO — and, for each concept, documents the variants actually used across the eight macro-varieties of Spanish. It covers four sectors: medicine, technology, law and the environment.

The difference from these standards lies not in the concept itself—which is respected—but in the final stage: whereas an international resource offers a single Spanish equivalent, this repository compiles the living forms used in each variety, with examples from real-life conversation cross-referenced against the most authoritative lexicographical sources.

Example entry · Medicine

Ultrasound scan

Spain – Southern Cone            ultrasound scan

Mexico                                      ultrasound

Caribbean                                 sonogram

Reference                                 SNOMED CT · ICD-11

WHY IT MATTERS

Five gaps that no 
international mechanism can address

MeSH, SNOMED CT, ICD-11, the IEEE and ISO glossaries, Black’s Law Dictionary, IATE and the IPCC glossaries are high-quality resources. However, they share five structural limitations that cannot be remedied by making incremental improvements, as these limitations stem from decisions made with other purposes in mind.

They stop at a single equivalence

They define the concept in English and provide a single term in Spanish, without including the actual terms used in each variety.

They treat Spanish as a uniform

They do not document how vocabulary varies between Mexico, Chile, the Caribbean and Spain.

They rely solely on what is written

None of them record the term used by the doctor or lawyer when speaking, which does not always match that used in the official document.

They do not show actual usage

They provide definitions, but no contextualised examples indicating who uses each variant and in what register.

They have a geographical bias

Strictly speaking, several of these are peninsular or institutional variants that under-represent the Spanish spoken in the Americas.

The repository completes that section

It adheres to the concept of the international standard and adds what is missing: the actual names used for each variety.

What makes it different

Four decisions that make 
all the difference

From international concept to actual use

Each entry starts with the standardised technical term and works its way down to what is actually said — "hora médica" in Chile, "baumanómetro" in Mexico —, recapturing what the single equivalent leaves out.

The eight macro-varieties

It records the variants of the eight dialectal regions, drawing on VARILEX-R and CORPES XXI from the RAE and ASALE.

Based on spoken corpus data

The basis is spoken language, the most direct source of actual usage, which reflects variation across registers better than written text.

In line with the standards

Each entry links to leading international resources. It does not replace them: it adds the dialectal dimension that they lack.

Applications

From professional consultancy to 
AI training

Direct use

A searchable resource that improves the quality and consistency of specialist work across different regions.

  • Specialist translation and software localisation
  • Technical writing and document standardisation
  • Training for healthcare and legal professionals
  • Communication between public authorities and citizens from different backgrounds

Track 1 · Public service and the profession

Infrastructure for language models

Structured terminology is clean, verified training data with controlled dialectal coverage.

  • Fine-tuning models for specialised domains
  • Retrieval and generation (RAG) systems
  • Ontologies for semantic reasoning
  • Correcting the dialectal asymmetry of current models

 

Track 2 · Artificial intelligence

Scope

Four sectors, one 
contribution

For each sector, the repository aligns with the international standard and adds the layer that none of them cover.

Medicine and health

Reference

MeSH · SNOMED CT · ICD-11

 

Contributes

Real-world clinical variants by type and oral record

Engineering and technology

Reference

Glossaries: IEEE 610 · ISO/IEC 24765

 

Contributes

A list of Anglicisms and country-specific adaptations

Law and legislation

Reference

Black’s Law Dictionary · IATE

 

Contributes

Equivalents in Spanish-speaking legal systems

Natural sciences and the environment

Reference

IPCC AR6 Glossary · UNEP

 

Contributes

Environmental and risk terminology by region

Trust

Methodological rigour and 
institutional support

The repository draws on sources of the highest authority and is based on the San Millán Corpus of Spoken Spanish. It is managed and maintained by the Foundation for the Transformation of La Rioja, which determines its scope and the terms of access.

Standards define the concept. This 
repository recovers the Spanish that 
truly captures its essence

From the single technical term to the eight vivid ways of expressing it. That is the section we are beginning to document.