EDVAL’s architecture and infrastructure
EDVAL constitutes a strategic infrastructure for Spanish in artificial intelligence, articulated through four complementary technical blocks that guarantee interoperability, linguistic quality, and computational scalability. The first block corresponds to the federated data space, which implements a resource catalog, interoperable connectors, and clearing house mechanisms aligned with European standards. The second block encompasses the oral corpus of Spanish, recognized as the largest oral corpus currently available, covering multiple dialectal varieties, validated transcription, advanced linguistic annotation, and structured metadata for traceability. The third block comprises artificial intelligence models specialized in voice, developed through Deep Learning techniques and trained specifically on the oral corpus for automatic speech recognition (ASR) and natural language processing (NLP). The fourth block constitutes the supporting technological infrastructure, based on GPU-accelerated computing, high-capacity unified storage, and GitOps architecture for continuous deployment. This modular structure guarantees the horizontal scalability of the system and its adaptation to European regulatory requirements such as GDPR and the Data Governance Act.
A living ecosystem with european interoperability
The EDVAL data space implements an ecosystem platform that facilitates the secure and traceable exchange of linguistic resources among project participants. The resource catalog acts as a centralized registry of available assets, including textual corpora, oral corpora, language models, training datasets, and linguistic processing services, all described using standardized metadata that ensures semantic interoperability. The connectors constitute fundamental technical components that allow the integration of external systems with the data space, implementing authentication, authorization, and encryption protocols that ensure traceability and granular access control to each resource. The clearing house function provides mechanisms for validation, registration, and auditing of transactions between data providers and consumers, guaranteeing transparency in the exchange and compliance with established license agreements. Technical interoperability is based on the adoption of European standards promoted by GAIA-X and the Data Spaces Support Centre (DSSC), ensuring compatibility with other sectoral data spaces and facilitating the cross-border federation of linguistic resources. This architecture guarantees data sovereignty, distributed governance, and regulatory compliance according to current European regulations.
| Component | Technical Function | Standard |
|---|---|---|
| Resource Catalog | Centralized registry of linguistic assets with standardized metadata (textual corpora, oral corpora, language models, training datasets, processing services) | GAIA-X, DSSC |
| Interoperable Connectors | External system integration components with authentication, authorization, and encryption protocols for granular access control | GAIA-X, DSSC |
| Clearing House | Validation, registration, and auditing mechanisms for transactions between providers and consumers with full traceability and license compliance | GAIA-X, DSSC |
Transcription, annotation, and temporal alignment
The oral corpus of Spanish constitutes the central linguistic asset of EDVAL, recognized as the largest oral corpus currently available for the Spanish language. The transcription of the corpus implements rigorous validation protocols that guarantee phonetic and orthographic precision across multiple dialectal varieties of Peninsular and Latin American Spanish, covering formal, informal, specialized, and colloquial contexts.
The linguistic annotation incorporates morphosyntactic tagging, prosodic analysis, marking of dialectal phonetic phenomena, and segmentation into minimal linguistic units, facilitating advanced computational analysis and language model training. The temporal alignment synchronizes textual transcriptions with corresponding audio segments using precise timestamps, allowing phonetic-acoustic analysis, the training of speech synthesis systems, and the development of automatic speech recognition models adapted to specific dialectal varieties. This richness of annotation and geographical diversity positions the corpus as a differentiating resource to train AI models in Spanish with native quality, avoiding biases introduced by transfer from models trained predominantly in English.
Join us, share and promote Spanish
EDVAL invites developers of AI systems, technology institutions, universities and researchers specialising in natural language processing to join the technical ecosystem, by contributing linguistic resources, validating voice models or deploying advanced services on the shared infrastructure. The adoption of open standards and transparent governance positions EDVAL as the leading platform for native Spanish in artificial intelligence, facilitating technical collaboration and generating sustainable value for the Spanish-speaking ecosystem.