Speaker
Description
Organisations in security-critical sectors such as healthcare and manufacturing hold data of high analytical value, but regulatory, competitive and privacy constraints prevent them from pooling it to train artificial intelligence models. The SOTERIA project introduced here addresses this gap through the research and development of trustworthy data spaces that allow AI to be deployed on sensitive data without requiring participants to give up control over it.
The technical approach of SOTERIA combines federated learning with Trusted Execution Environments (TEEs). In the federated architecture data never leaves the owner's premises: only model updates are exchanged, and they travel over secure communication channels. Aggregation and orchestration run inside TEEs based on Intel SGX, Intel TDX, AMD SEV-SNP and NVIDIA H100, so that neither the infrastructure provider nor the other users can inspect the computation. Before contributing, each participant verifies the execution environment through remote attestation, obtaining cryptographic evidence of platform and code integrity. These guarantees are complemented by secure multi-party computation, homomorphic encryption and data minimisation applied to the most sensitive exchanges.
Trust at the computation layer is extended to the identity layer through decentralised access control based on Self-Sovereign Identity, verifiable credentials anchored in DLT, and zero-knowledge proofs that allow attributes to be proven without disclosing the underlying data. The platform is completed by an intelligent monitoring layer designed under Zero Trust principles, in which no node, identity or workload is trusted by default and anomalous behaviour is detected and responded to continuously.
The SOTERIA solution will be validated in healthcare and industrial scenarios. The clinical use case will cover the detection of cerebral microbleeds in magnetic resonance imaging in the context of Alzheimer's disease, while the industrial use case will be focused on modelling key parameters from production data across plants operating a common machine archetype to help building a reliable digital twin. Both address interoperability with existing standards: HL7/FHIR and DICOM in the clinical domain, MES, OPC UA and B2MML in the industrial one. The evaluation will compare local, federated and TEE-backed federated approaches in terms of data available for training, performance, scalability, privacy, exposure surface, communication cost, environment verifiability and resilience against vulnerabilities, as starting hypotheses to be tested experimentally.
In conclusion, SOTERIA aims to deliver an integrated infrastructure in which data remains under the control of its owners while models are trained in a distributed manner, workloads run in secure and verifiable environments, access is managed in a decentralised way, and security threats are monitored and mitigated continuously. The expected outcome is a technological foundation for secure collaboration between organisations and a contribution to the development of trustworthy European data spaces.
This work has been supported by the SOTERIA project (EXP 00177383 / ITC-20251151), coordinated by Bahía Software, funded by CDTI through the Consorcios Regionales INNTERCONECTA-STEP 2025 programme and co-financed by the European Union through the European Regional Development Fund (ERDF) 2021-2027.