Speaker
Description
Distributed quantum computing (DQC) partitions a circuit across interconnected QPUs that exchange entanglement rather than qubits: a promising answer to the scalability limits of monolithic processors, and a natural fit to the multi-node model HPC centres already run. Physical testbeds remain scarce, so emulation is where protocols, partitioning strategies and communication costs are studied, yet the tooling is fragmented. Many emulators are in circulation, each with its own network abstraction, entanglement primitives and resource accounting. Several are unmaintained, depend on another backend, or target quantum communication rather than computation. A center deciding what to deploy has no comparable numbers to work from. We present a benchmarking framework that produces them, and results from applying it on production HPC resources at CESGA.
We first surveyed the available tools following the taxonomy of Caleffi et al. [1], assessing each against required criteria (teleportation support, gate coverage), important ones (discrete-event simulation, multi-node execution, network and noise modelling) and useful ones (native metrics, active development). The resulting capability matrix is directly reusable by other centres and a result in itself: only four surveyed tools met the minimum requirements, and the distinction between full, partial and backend-dependent support is not visible from their own documentation.
The workload is a distributed inverse quantum Fourier transform (iQFT), a canonical subroutine of phase estimation and Shor's algorithm that demands high entanglement and global connectivity, and whose output on a Fourier basis state is deterministically verifiable, so correctness can be checked rather than assumed. The circuit is factorised recursively into local iQFT blocks and inter-node controlled phase gradients executed through gate teleportation with cat-entangle and cat-disentangle primitives. Because all controlled-phase gates sharing a control within a layer teleport together [2], the entangled-pair count for n = mk qubits over k nodes drops from C(mk,2) to approximately m*C(k,2), improving communication scaling from O(m2k2) to O(mk).
Five emulators are benchmarked: Qiskit Aer, a monolithic baseline with teleportation emulated via Bell pairs and conditional correction; SquidASM, the most hardware-aware, with discrete-event scheduling, decoherence and link-level noise; Interlin-q, which partitions automatically over QuNetSim; SQUANCH, an agent-oriented density-matrix simulator; and CUNQA [3], developed in the Spanish quantum ecosystem and deployed on QMIO and FinisTerrae III. Sweeping 4-20 qubits and 1-8 nodes, we measure execution time, peak memory and fidelity against the monolithic output. Heterogeneous APIs and inconsistent resource reporting make this hard to automate, so each emulator sits behind an adapter exposing a common node interface, and runs are generated into SLURM array jobs with isolated environments per emulator on QMIO.
Emulators diverge by orders of magnitude in cost and, more consequentially, in correctness: some preserve the expected noiseless output across the sweep while others degrade silently as size or node count grows. Cross-emulator disagreement is the only practical oracle at these scales. The framework [4] is thus a reusable instrument rather than a single measurement; all code and results are released openly.
[1] Caleffi et al., Comput. Netw. 254, 110672 (2024).
[2] Cardama et al., arXiv:2605.10710.
[3] cesga-quantum-spain.github.io/cunqa
[4] github.com/gdiazcamacho/DQC-Emulator-Benchmarking