

In our latest paper, we’ve taken a big step toward large scale fault-tolerant quantum computing, squeezing up to 94 error-detected qubits (and 48 error-corrected qubits) out of just 98 physical qubits, a low-fat encoding that cuts overhead to the bone. With 64 of our logical qubits, we were able to simulate quantum magnetism at a scale that can be exceedingly difficult for classical computers.
The "holy grail" of quantum computing is universal fault-tolerance: the ability to correct errors faster than they occur during any computation. To realize this, we aim to create “logical qubits,” which are groups of entangled physical qubits that share quantum information in a way that protects it. Better protection leads to lower “logical” error rate and greater ability to solve complex problems.
However, it’s never that easy. An unofficial law of physics is “there’s no such thing as a free lunch”. Creating high quality, low error-rate logical qubits often costs many physical qubits, thus reducing the size of calculations you can run, despite your new, lower-than-ever error rates.
With our latest paper, we are thrilled to announce that we have hit a key milestone on the Quantinuum roadmap: an ultra-efficient method for creating logical qubits, extracting a whopping 48 error-corrected and 64 error-detected logical qubits out of just 98 physical qubits. Our logical qubits boasted better than “break-even” fidelity, beating their physical counterparts with lower error rates on several different fronts. And still that isn’t the end of the story: we used our 64 error-detected logical qubits in a large-scale quantum magnetism simulation, laying the groundwork for future studies of exotic interactions in materials.
To get this world-leading result, we employed a neat trick: ‘nesting’ super efficient quantum error-detecting codes together to make a new, ultra-efficient error-correcting code. Dr. DeCross, a primary author on the paper, said this nesting is like “braiding together ropes made out of ropes made out of ropes”. Physicists call this ‘code concatenation’, and you can think of it as adding layers of protection on top of each other.
To begin, we took the now-famous ‘iceberg code’, a quantum error detection code that gives an almost 1:1 ratio of physical qubits to logical qubits. The iceberg code only detects errors, however, which means that instead of actually correcting errors it lets you throw out bits where errors were detected. To make a code that could both detect and correct errors, we concatenated two iceberg codes together, giving a code that can correct small errors while still boasting a world-record 2:1 physical:logical ratio (physicists call this a “high encoding rate”).
The team then benchmarked the logical qubits, checking large system-scale operations and comparing them to their physical counterparts. This introduces a crucial hurdle to clear: oftentimes, researchers end up with logical qubits that perform *worse* than their physical counterparts. It’s critical that logical qubits actually beat physical ones, after all – that is the whole point!
Thanks to some clever circuit design and our natively high fidelities, the new logical qubits outperformed their physical counterparts in every test we performed, sometimes by a factor of 10 to 100.
Of course, the whole point is to use our logical qubits for something useful, the ultimate measure of functionality. With 64 error-detected qubits, we performed a simulation of quantum magnetism; a crucial milestone that validates our roadmap.
The team took extra care to perform their simulation in 3 dimensions to best reflect the real-world (often, studies like this will only be in 1D or 2D to make them easier). Problems like this are both incredibly important for expanding our understanding of materials, but are also incredibly hard, as their complexity scales quickly. To make qubits interact as if they are in a 3D material when they are trapped in 2D inside the computer, we used our all-to-all connectivity, a feature that results from our movable qubits.
Breaking the encoding rate record and performing a world-leading logical simulation wasn’t enough for the team. For their final feat, the team generated 94 error-detected logical qubits, and entangled them all in a special state called a “GHZ” state (also known as a ‘cat’ state, alluding to Schrödinger’s cat). GHZ states are often used by experts as a simple benchmark for showcasing quantum computing’s unique capacity to use entanglement across many qubits. Our best 94-logical qubit GHZ state boasted a fidelity of 94.9%, crushing its un-encoded counterpart.
Taken together, these results show that we can suppress errors more effectively than ever before, proving that Helios is capable of delivering complex, high-fidelity operations that were previously thought to be years away. While the magnetism simulation was only error-detected, it showcases our ability to protect universal computations with partially fault-tolerant methods. On top of that, the team also demonstrated key error-corrected primitives on Helios at scale.
All of this has real-world implications for the quantum ecosystem: we are working to package these iceberg codes into QCorrect, an upcoming tool that will help developers automatically improve the performance of their own applications.
This is just the beginning: we are officially entering the era of large-scale logical computing. The path to fault-tolerance is no longer just theoretical—it is being built, gate by gate, on Helios.
Quantinuum, the world’s largest integrated quantum company, pioneers powerful quantum computers and advanced software solutions. Quantinuum’s technology drives breakthroughs in materials discovery, cybersecurity, and next-gen quantum AI. With over 500 employees, including 370+ scientists and engineers, Quantinuum leads the quantum computing revolution across continents.
Quantum computing is entering a new era. As systems move from Noisy Intermediate-Scale Quantum (NISQ) toward Fault-Tolerant Application-Scale Quantum (FASQ), traditional metrics like qubit count, gate fidelity, and gate speed are no longer enough to describe what a machine can actually deliver.
Developed by Sandia National Laboratories, with input from Quantinuum and NVIDIA, QUOPS—the Quantum Universal Operations Performance System—is a common, architecture-agnostic benchmark for measuring quantum performance across both physical- and logical-qubit systems on the path toward quantum utility.
QUOPS can be applied to different architectures, codes, modalities, and levels of fault tolerance. QUOPS runs the same randomized workloads across different computational shapes, measures whether each workload succeeds, identifies the boundary of a system’s capability region, and reports two summary metrics:
The result is a direct measure of how much computation a system can perform and how quickly it can do so. Together, these measurements provide a two-dimensional view of capability while reducing system performance to a common currency: quantum operations.
Component-level metrics remain essential for engineering. Qubit count, two-qubit fidelity, and gate speed can reveal control errors, crosstalk, leakage, connectivity constraints, and other system limitations. But they do not necessarily predict system-level performance.
Fault tolerance makes this gap even larger. Physical operations become logical computation with the addition of logical encoding, syndrome measurement, decoding, logical gate construction, magic-state production, routing, and control. Ultimately, this means that fault tolerance expands the relevant currencies of computation. Code distance, logical fidelity, magic-state throughput, decoding, connectivity, and space-time volume can matter far more for performance than raw qubit count or individual gate speeds.
This creates a growing challenge for buyers, governments, and researchers. As organizations move from experimentation toward larger-scale and potentially on-premise quantum systems, they need to know a simple thing:
What computation can a machine actually execute successfully?
QUOPS addresses that question by measuring the integrated system rather than inferring performance from individual components.
This is particularly important as the field considers workloads requiring roughly 10⁹–10¹² operations on thousands of qubits. Today's measured capabilities are still orders of magnitude smaller; QUOPS turns that gap into a measurable quantity.
QUOPS can also provide a practical layer for quantum procurement and planning.
HPC centers need to understand when quantum computing will become useful for real workloads. Customers may have a goal of procuring a system that can, for example, run a trillion error-free operations. Today, answering these questions can require complex resource estimates that depend on hardware modality, QEC code, magic-state factories, decoding, compilation, and other architectural choices.
In both cases, QUOPS provides a simpler system-level reference point: Q describes the size of computation a machine can execute, while Ω describes its effective throughput. Furthermore, because QUOPS is architecture-neutral and includes anti-gaming provisions, it can also help buyers compare competing systems without relying solely on vendor-selected metrics or announcements.
While QUOPS is a new benchmark, it has already been measured on several vendors’ hardware. This marks an important step for our industry: we can now compare vendors directly, assessing their capabilities in a way that flattens the differences introduced by modality and architecture choices.
Figure 1. The QUOPS capability region and score for state-of-the-art processors from Quantinuum, Google, and IBM (adapted from Figure 2 of the scientific publication co-authored by Quantinuum, Sandia National Laboratories, and NVIDIA). QUOPS specifies a random circuit construction that can be built for a specified width (number of qubits) and size (number of quantum gates). A set of circuits is run at several width and size points and the average fidelity of those circuits are measured and compared to a predefined threshold. Each labeled point above represents experimental data from QUOPS circuits that passed the threshold with high confidence. The lines are filled capability limits of each machine between the points. The stars indicate the QUOPS score (Q), which is the experimental data point that passes the threshold with maximum size inside the shaded cone of width2 ≤ size ≤ width3.
Figure 2. The QUOPS score (Q) vs rate (Ω) for state-of-the-art processors from Quantinuum, Google, and IBM (adapted from Figure 2 of the QUOPS scientific publication co-authored by Quantinuum, Sandia National Laboratories, and NVIDIA). Each point is the maximum QUOPS circuit size that passes the threshold within the specified cone and rate that it was run. The dashed lines indicate the extrapolated effect of error mitigation, which attenuates the rate by including the shot overhead needed for general-purpose error mitigation. The gradient lines show the estimated runtime of a circuit at a given score and rate.
Figures 1 and 2 show how QUOPS quantifies the capability tradeoffs between different systems. Willow and Boston are superconducting systems with very fast gate speeds but limited connectivity, while Helios is a trapped-ion QCCD system with effective all-to-all connectivity but much slower gates. Willow and Boston have smaller capability regions and QUOPS scores but higher QUOPS rates; while Helios reaches larger capability regions and QUOPS scores but lower QUOPS rates. All three systems have the ability to trade speed for larger circuits with error mitigation. This is commonly assumed in the community but is nicely quantified with the QUOPS rate, which accounts for the corresponding sampling overheads of general error mitigation techniques (as shown by the dashed lines in Figure 2).
QUOPS will not replace every quantum benchmark. The field will continue to need application-specific suites, component-level measurements, hybrid-HPC benchmarks, and independent verification.
QUOPS instead serves as a common system-level yardstick that can make roadmaps more comparable, procurement more objective, and progress easier to track.
We are calling on vendors to report QUOPS metrics (Q, Ω) and capability regions alongside existing metrics, buyers and agencies to consider QUOPS thresholds in RFPs, and researchers to contribute fault-tolerant architectures and resource estimates.
As quantum computers become fault tolerant, success will no longer be defined simply by how many qubits a machine contains or how low its error rates are.
It will be defined by the computation the machine can deliver.
QUOPS is a step toward measuring that capability—and toward giving the quantum industry a benchmark built for the era ahead.
Building a quantum computer is one thing. Showing that it is genuinely using quantum mechanics is another.
A new experiment, just published in Nature Communications, takes a fresh approach to that question. Instead of relying on entanglement or the complex calculations often used to benchmark quantum computers, researchers designed a simple game (initially published in Physical Review Letters) that tests something more fundamental: quantum superposition.
Using superposition, the team constructed a game where quantum mechanics provides a provable advantage over classical approaches. Once the game was set, the team ran it on real hardware. The results showed a clear performance gap between the best possible classical system and our System Model H2 – a gap that only grew as the test became more difficult.
The game is played by a single player with access to a computer. The player receives a quantum state representing a set of numbers—for example, {0, 1, 5, 7}. Their goal is to return a number that belongs to the complement of that set: {2, 3, 4, 6}.
That sounds simple. But as the size of the sets grows, something remarkable happens.
A classical strategy needs to test many numbers to succeed. A quantum strategy, however, succeeds in one step. The authors show that the quantum strategy has a score that grows exponentially faster.
Importantly, this isn't based on an assumption that this problem is difficult for classical computers. The separation is mathematically proven. In other words, the researchers can show that the quantum advantage exists without relying on unproven assumptions from complexity theory.
Using our System Model H2, the experimenters were able to confirm the theoretically derived separation between the quantum and the classical strategy (up to the largest sizes they could fit on the quantum processor) with high confidence – showing that the violation remained close to exponential.
Many famous experiments testing quantum behavior rely on entanglement and non-locality, where multiple parties share parts of a quantum system.
This experiment is different.
There is only one player, who has access to the entire quantum system. The advantage comes from superposition—the ability of a quantum system to exist in a combination of states until it is measured.
That distinction matters because it provides another way to ask whether a quantum computer is actually behaving quantum mechanically.
The researchers turned their game into an experimental test and ran thousands of different circuits on Quantinuum's System Model H2. The scores they observed were close to the theoretical predictions for a quantum strategy.
One of the challenges with existing quantum-computing demonstrations is figuring out whether the machine really produced the result it was supposed to produce.
For example, random circuit sampling can be extremely difficult to verify classically as systems become larger. That creates a tension: you want to demonstrate that a quantum computer is doing something a classical computer cannot easily reproduce, but you also need a practical way to check the result.
The complement-sampling game offers a different approach. The violation of classical performance can be efficiently verified with a classical computer.
That makes the test potentially more scalable: you don't need to reproduce the entire quantum computation on a classical computer just to determine whether the machine demonstrated non-classical behavior.
The deeper message of the experiment is that demonstrating a quantum computer isn't simply about having qubits.
A convincing demonstration should show that the machine is exploiting properties that genuinely distinguish quantum computation from classical computation. Here, the researchers focus on one of those defining properties—superposition—and construct a game where quantum mechanics provides a provable advantage.
This first experimental demonstration of complement sampling doesn't close every possible loophole, which is common for this sort of experiment – closing the major experimental loopholes in Bell-inequality tests took decades—a body of work that ultimately contributed to the 2022 Nobel Prize in Physics. The researchers explicitly note that the implementation relies on assumptions about how the input state is prepared, so the experimental results should be interpreted with some caution.
Still, the work provides a new way to probe the boundary between classical and quantum computation.
And that may be the most interesting part: rather than asking only “How many qubits does the machine have?”, we can ask a more meaningful question—
“What can this machine do that only a quantum system can?”
That is ultimately what it takes for a quantum computer to actually be quantum.
The NISQ1 era is coming to an end. At Quantinuum, we’ve already demonstrated numerous QEC codes, all the primitives needed for logical computation, steadily declining logical error rates, and full computations at the logical level.
But there’s still a way to go. One of the defining challenges over the coming years will be putting it all together into a usable – and scalable – fault tolerant architecture. Today, we are excited to announce that we have experimentally validated one of our own leading candidates for such an architecture, the Helix code.
With this demonstration, we have put all the pieces together: logical memory, logical computation, a heterogenous code architecture that optimizes for magic vs gates, all with super efficient operations and record-breaking2 fidelity.
The delicate nature of qubits gives them their strength – they can be entangled, placed into superpositions, and even teleported. However, this comes at a cost: on the hardware level, quantum bits (qubits) will always be noisier than classical bits.
Enter quantum error correction (QEC). QEC moves us past prohibitive physical noise to fidelities that really matter; where industrial workflows and scientific discovery live. Our field has been hard at work to realize - and optimize - QEC, and we are finally starting to reap the fruits of that labor.
However, for the most part, this work has taken shape only a few pieces at a time: a demonstration of fault tolerant gates here or memory there, sometimes even a full fault-tolerant algorithm, but rarely do we see demonstrations at the architectural scale needed to build our next generation of machines.
The difficulty is that encoding and performing fault-tolerant computation costs considerable space (qubit number) and time (circuit complexity), which QEC researchers summarize with a “spacetime volume”.
The Helix code was custom-designed to usher in the next generation of fault tolerance.
Using our reconfigurable qubits, we designed Helix to minimize its spacetime volume by employing more exotic entanglement schemes compared to traditional codes (imagine cat’s cradle compared to a simple, 2D net). This entanglement complexity is impossible with processors that don’t have reconfigurable connectivity.
Ultimately, this translates to a code that requires fewer physical qubits per logical qubit, while also giving you fast and simple computing.

In general, gates between logical qubits can be quite difficult because logical qubits are composed of physical qubits that are entangled together in some specific way. Sometimes, a single physical qubit may even be shared between multiple logical qubits, as is the case with codes that offer lots of logical qubits per physical qubit. Performing gates across these complex structures can be tricky, and can take a lot of individual operations on physical qubit pairs to accomplish.
There are two major exceptions. The first, called a transversal gate, is where the logical operation maps directly onto the physical one: you just perform a regular 2-qubit gate between each physical qubit in each logical qubit.
The second is simpler still: gates can be accomplished by simple software-level qubit relabeling (eg simply renaming qubit A to qubit B), combined with easy, single qubit gates. This type of automorphism, or permutation-based gate, is particularly elegant.
The Helix code makes heavy use of transversal and automorphism gates, making it considerably faster and easier to compute with than a lot of other options. Ultimately, this translates to a significant reduction in both space (qubit) and time (circuit complexity) overheads: less space is needed for block encoding and ancilla; and time is drastically reduced when simple software relabeling or transversal gates are performed in the place of expensive protocols like lattice surgery.

Experiment 1: Logical Memory
The team started by showing that the Helix code can successfully preserve encoded quantum information for extended periods of time.
To show this, the team started with their logical qubits in a given state. Then, they performed 20 rounds of syndrome extraction, paying special attention to leakage (a dominant source of error on Helios). To remove leakage, the team leveraged Helios’ new leakage repump capacity, as well as circuit-level leakage reduction units.
Result: per qubit, per round, they achieved an error rate of 4.6 x 10-5, with no post selection.
This amounts to a block logical error per round of 9.3 x 10-5, with no post selection. With a small amount of post selection (0.5%), the block logical error per round was reduced to 1.9 x 10-5.
Quantum memory is one of the most fundamental building blocks of a fault-tolerant quantum computer. A useful quantum processor must be able to preserve quantum information long enough to perform the computation, error correction, and communication required by larger algorithms.
These results prove that encoded quantum information can be preserved with a lower error rate than the underlying physical operations – all without needing post selection.
Experiment 2: Logical Computation
A central feature of the Helix logical architecture is that encoding multiple logical qubits does not require correspondingly expensive logical computation. By construction, this code has a variety of logical gates all implementable with only physical single-qubit gates and qubit relabeling. These ‘SWAP-transversal’, or ‘automorphism’, gates provide the ability to do some logical circuits essentially for free, as permutations are realized by simple ion-transport and software level relabeling.
The team experimentally tested the code’s computational abilities by benchmarking the complete logical Clifford group (i.e., all gates except for T gates) while interleaving up to 27 rounds of active adaptive syndrome extraction.
Result: 2.8 x 10-4 logical error rate per Clifford gate, a significant improvement (4.28x) over Helios’ physical 2-qubit Clifford error rate, again achieved without post selection.
This impressive result is partially enabled by the team’s clever adaptive syndrome extraction (ASE) technique. Their ASE technique reduces the number of physical gates required per logical gate by about 33%. This pruning also shortens the physical run time, reducing the ‘wall clock duration’ by about 23%. Both gates and idling contribute significantly to errors, so these reductions translate to a lower logical error rate.
This experiment proves the Helix code’s ability to compute, all while showing significant improvement over the physical level with no post selection. In addition, this marks the first demonstration of randomized benchmarking on a code encoding more than one logical qubit, an important milestone for our community.
Experiment 3: Universality via Logical Entanglement Across Different Codes
Clifford gates alone are insufficient for universal fault-tolerant computation; our QEC architecture must also provide access to non-Clifford resource states (often called “magic”). While the Helix code has many desirable features in terms of Clifford gates, it’s sub-optimal for preparing magic states. Rather than forcing the Helix code to work in this regime, we developed our architecture to employ two codes; one for Clifford gates and memory, and one for magic state preparation. Using different codes each optimized for their own tasks, called a heterogenous architecture, makes the entire assembly considerably more efficient and cost effective.
The trick that makes it all possible is something called chain-mapping, that allows the QPU to smoothly switch between underlying encoding schemes. To test this, the team used a rotated surface code for magic state generation, which would then be injected into the computational (Helix) code to generate non-Clifford gates (enabling fault tolerant universal computation).
Rather than performing magic state injection directly, the team wanted to benchmark the interface (the chain-map). To do this, they used their chain-mapped gates to prepare a three logical qubit GHZ state that spans the two different codes. The resulting GHZ state contained 1 logical qubit from the surface code and two from the Helix code, making for a truly heterogeneous structure.
Result: The logical GHZ state had a fidelity lower bound of 99.925%, and an upper bound of 99.975%. The lower bound exceeded the physical baseline, making all three experiments better than their physical counterparts.

This experiment demonstrates one of our key architectural advantages: using our reconfigurable connectivity to employ multiple QEC encodings in a single fault tolerant architecture, improving our efficiency and reducing qubit costs.
The quantum computing industry has proposed many approaches to error correction. This paper marks one of the first experimentally validated plans for a QEC architecture. With the low logical error rates (all improving on the physical baseline), the practical logical operations (which drastically reduce qubit costs and compute time), and real commercial hardware performance, this result helps to prove that we will deliver on our roadmap.
A crucial element of this demonstration is that these results were obtained on our commercial hardware. This is not a result from a testbed, or a result from hardware that has limited functionality. This is a result from the same computer as our customers use for their own research.
Furthermore, simulations indicate that the improvements in physical fidelity we expect from moving from Helios to Apollo will bring logical error rates in line with our roadmap targets. Because logical error rates depend strongly on physical error rates, Apollo’s expected improvements at the physical level should translate directly into lower logical error rates. Importantly, we expect to achieve these gains without increasing the code distance or using additional physical qubits per logical qubit.

Building a fault-tolerant quantum computer requires solving multiple engineering challenges. In this result we have proven our path to a scalable QEC architecture, showing not just one-off results on test stand hardware, but a harmonious whole consisting of:
✓ A candidate architectural code
✓ Logical memory
✓ Logical computation
✓ A path to magic
✓ Multiple encodings in one architecture
✓ Commercial hardware compatibility
By validating our fault-tolerant architecture on real commercial hardware, Quantinuum has taken a significant step toward Apollo - and toward quantum computers capable of solving meaningful problems at scale.

1 Noisy Intermediate-Scale Quantum
2 Based on a study of existing literature