Quantinuum researchers tackle AI’s ‘interpretability problem’, helping us build safer systems

‍

June 26, 2024
The Artificial Intelligence (AI) systems that have recently permeated our lives have a serious problem: they are built in a way that makes them very hard - and sometimes impossible - to understand or interpret. Luckily, our team is tackling this problem, and we’ve just published a new paper that covers the issue in detail.


It turns out that the lack of explainability in machine learning (ML) models, such as ChatGPT or Claude, comes from the way that the systems are built. Their underlying architecture (a neural network) lacks coherent structure. While neural networks can be trained to effectively solve certain tasks, the way they do it is largely (or, from a practical standpoint, almost wholly) inaccessible. This absence of interpretability in modern ML is increasingly a major concern in sensitive areas where accountability is required, such as in finance and the healthcare and pharmaceutical sectors. The “interpretability problem in AI” is therefore a topic of grave worry for large swathes of the corporate and enterprise sector, regulators, lawmakers, and the general public. 

These concerns have given birth to the field of eXplainable AI, or XAI, which attempts to solve the interpretability problem through so-called ‘post-hoc’ techniques (where one takes a trained AI model and aims to give explanations for either its overall behavior or individual outputs). This approach, while still evolving, has its own issues due to the approximate nature and fundamental limitations of post-hoc techniques.  

The second approach to the interpretability problem is to employ new ML models that are, by design, inherently interpretable from the start. Such an interpretable AI model comes with explicit structure which is meaningful to us “from the outside”. Realizing this in the tech we use every day means completely redesigning how machines learn - creating a new paradigm in AI. As Sean Tull, one of the authors of the paper, stated: “In the best case, such intrinsically interpretable models would no longer even require XAI methods, serving instead as their own explanation, and one of a deeper kind.”

At Quantinuum, we’re continuing work to develop new paradigms in AI while also working to sharpen theoretical and foundational tools that allow us all to assess the interpretability of a given model. In our recent paper, we present a new theoretical framework for both defining AI models and analyzing their interpretability. With this framework, we show how advantageous it is for an AI model to have explicit and meaningful compositional structure.

The idea of composition is explored in a rigorous way using a mathematical approach called “category theory”, which is a language that describes processes and their composition. The category theory approach to interpretability can be accomplished via a graphical calculus which was also developed in part by Quantinuum scientists, and which is finding use cases in everything from gravity to quantum computing. 

A fundamental problem in the field of XAI has been that many terms have not been rigorously defined, making it difficult to study - let alone discuss - interpretability in AI. Our paper presents the first known theoretical framework for assessing the compositional interpretability of AI models. With our team’s work, we now have a precise and mathematically defined definition of interpretability that allows us to have these critical conversations.    

After developing the framework, our team used it to analyze the full spectrum of ML approaches. We started with Transformers (the “T” in ChatGPT), which are not interpretable – pointing to a serious issue in some of the world’s most widely used ML tools. This is in contrast with (sparse) linear models and decision trees, which we found are indeed inherently interpretable, as they are usually described.  

Our team was also able to make precise how other ML models were what they call 'compositionally interpretable'. These include models already studied by our own scientists including DisCo NLP models, causal models, and conceptual space models.    

Many of the models discussed in this paper are classical, but more broadly the use of category theory and string diagrams makes these tools very well suited to analyzing quantum models for machine learning. In addition to helping the broader field accurately assess the interpretability of various ML models, the seminal work in this paper will help us to develop systems that are interpretable by design. 

This work is part of our broader AI strategy, which includes using AI to improve quantum computing, using quantum computers to improve AI, and – in this case - using the tools of category theory and compositionality to help us better understand AI. 

About Quantinuum

Quantinuum, the world’s largest integrated quantum company, pioneers powerful quantum computers and advanced software solutions. Quantinuum’s technology drives breakthroughs in materials discovery, cybersecurity, and next-gen quantum AI. With over 500 employees, including 370+ scientists and engineers, Quantinuum leads the quantum computing revolution across continents. 

Blog
|
corporate
September 28, 2026
Quantum World Congress 2026: Cutting Through the Noise to Deliver the Next Era of Compute

For years, the quantum landscape has been crowded with headline hype, speculative roadmaps, and vanity qubit counts.

While that noise hasn't completely disappeared, Quantum World Congress 2026 proved that the market's tolerance for hype is rapidly wearing thin. While benchmark simulations and roadmap announcements still generate headlines, the real conversation on the ground has shifted toward a far more demanding question: who is actually delivering useful execution today and making meaningful progress toward fault-tolerant quantum computing?

At Quantinuum, our stance has never wavered. Progress is measured by delivering integrated systems that perform useful computation reliably on real hardware and in live enterprise environments. Prove what you can do and publish the results.

What was most encouraging this year at Quantum World Congress was seeing much of the industry increasingly align around that view, with the conversation converging around several priorities in particular.

1. Fault tolerance proven on real hardware is the only credible path to scale
The industry is converging around a shared reality: useful quantum computing will be defined by fault tolerance.

But fault tolerance is not a single metric, a logical-qubit count, or a standalone error-correction demonstration. It requires a complete end-to-end architecture that can detect, correct, and manage errors across every layer of computation, from logical operations to resource-state preparation, and execution of useful algorithms on live hardware.

That is why context matters when evaluating industry announcements.

Demonstrating more logical qubits is not the same as demonstrating fault tolerance. Demonstrating real-time decoding in a simulation is not the same as demonstrating real-time quantum error correction on real hardware. Logical qubits, decoders, syndrome extraction, and resource-state preparation are all necessary ingredients. But proving an individual ingredient in isolation, particularly through simulations or benchmark environments, is fundamentally different from proving that a complete fault-tolerant architecture works on a live quantum computer.

The industry should not confuse progress on individual building blocks with proof of a scalable fault-tolerant system. The real test is whether all these capabilities work together to improve computational performance beyond the physical layer and in a way that can scale.

This is exactly what Quantinuum's Helix architecture is designed to do.

Running on our Helios system, Helix integrates the full stack required for fault-tolerant quantum computing and has already delivered world-class results on live commercial hardware. Helios achieves physical two-qubit gate error rates of 8×10⁻⁴, while Helix has demonstrated a logical compute error rate of 2.8×10⁻⁴, representing a 4.28× improvement over the physical Clifford gate error rate without post-selection, meaning no cherry-picked results.
‍
Just as importantly, Helix improves efficiency as well as reliability. Through adaptive syndrome extraction, it reduces physical gate requirements by 33% and wall-clock execution time by 23%. Today, the architecture is already supporting full computations using 64 error-detected logical qubits with better-than-physical performance and a highly efficient 1:1.5 encoding rate.

These results matter because they demonstrate more than isolated milestones. They show that a scalable fault-tolerant architecture is operating on real hardware today and delivering measurable improvements in computational reliability. As physical systems grow, Helix provides the framework that enables additional scale to translate into increasingly reliable computation.

The path to fault tolerance will continue to evolve. New codes, encoding approaches, and implementation techniques will emerge over time. What remains constant is the need for an architecture capable of orchestrating those innovations into a practical, scalable system. We believe Helix is well positioned to serve as that architecture.

2. The industry must measure performance with useful, standardized metrics
Customers cannot make informed buying decisions if every vendor continues to grade their own homework. Standardized benchmarks are critical because they shift the conversation from theoretical promises to actual, measurable usefulness.

While earlier benchmarks like Quantum Volume were helpful for NISQ-era systems, they don’t keep pace as fault-tolerant systems scale beyond classical simulation limits. That is why Quantinuum is helping drive industry alignment around QUOPS, developed by Sandia National Laboratories with input from NVIDIA and Quantinuum.

QUOPS measures true computational capability by evaluating both computational Size (QUOPS) and Speed (QUOPS/sec). It’s the difference between rating an engine by theoretical horsepower versus measuring how fast and far a vehicle can drive on an actual track.

Most importantly, QUOPS evaluates the capability delivered by the complete computing stack. As the industry moves toward fault-tolerant systems, useful benchmarks must measure the performance of the entire architecture rather than individual subsystems. They must remain transparent as workloads scale, allowing enterprises to clearly evaluate the real-world utility of a system for their specific problem sets.

3. The industrial manufacturing race is on
Proving a fault-tolerant concept on a laboratory bench is only the first step. The transition to utility-scale computing is fundamentally an engineering, supply chain, and manufacturing race.

Our roadmap clearly defines how we scale from a single chip with a 2D-grid layout to larger multi-chip packages. Rather than reinventing the wheel, we are leveraging proven semiconductor manufacturing techniques that built the modern microelectronics industry. We have backed this architecture with a world-class industrial ecosystem, partnering with commercial foundries and component leaders.

The companies that win the next phase of quantum computing will not simply invent breakthrough technologies. They will demonstrate the ability to manufacture, deploy, and scale them reliably.

4. Growing the ecosystem requires a powerful developer platform built for the fault-tolerant era
A high-performing QPU is only part of the equation. You also need the software layer required to run it. For quantum to succeed, it must integrate into the workflows developers and enterprise researchers are already using today.

Nexus is our developer platform designed to be the purpose-built enterprise integration layer, designed from the ground up to squeeze maximum power and accuracy out of Quantinuum’s architectures. Nexus provides seamless interoperability across native Guppy, NVIDIA CUDA-Q, and Microsoft Q#, directly connecting quantum hardware to classical AI and HPC environments.

This is not an experimental platform. It is already supporting meaningful commercial adoption today.

Adopted by more than 200 organizations, representing approximately 45% growth, and used by more than 1,000 active developers, Nexus has seen annual job submissions increase 10x when comparing August 2024-August 2025 with August 2025-August 2026.

As quantum computing enters the fault-tolerant era, the winning platform will be the one that makes advanced quantum capabilities accessible within the workflows enterprises already depend on.

5. Real progress is measured in live production deployments and real-world research
The future of high-performance computing isn't quantum versus AI or HPC. It is all three working together as a unified compute stack.

But we believe the true test of market leadership isn't a self-funded trial or a desktop simulation that claims hybrid integration. It is putting live hardware into mission-critical operational environments to solve real enterprise problems and advance meaningful science.

That is why we are focused on embedding our quantum systems directly into global computing infrastructure. Through our partnership with Oracle, we expect to bring commercial-scale quantum computing directly into an Oracle Cloud Infrastructure AI data center, with the goal of letting enterprise customers run quantum applications within their existing cloud footprint.

Similarly, our deep, multi-year collaboration with NVIDIA continues to advance the frontiers of hybrid compute: from co-founding the NVAQC center in Boston to physically integrating NVIDIA GPUs into our Helios hardware for real-time error correction, to running live workflows with Pfizer to help accelerate pharmaceutical discovery.

Ultimately, as we scale these capabilities from the system level to the enterprise, we believe quantum computing can serve as the crucial engine to unlock far greater value from classical AI. By integrating QPUs alongside GPUs and supercomputers, we are working to move beyond the limitations of classical compute alone.

This hybrid approach enables enterprises to generate higher-quality quantum training data, build more powerful models, and tackle complex chemistry and materials challenges that remain beyond the reach of classical AI by itself.

Looking Ahead‍

Quantum World Congress 2026 proved that the market is finally asking the right questions:

  • How do we achieve true fault tolerance at scale?
  • How do we measure objective computational capability?
  • How do we deploy quantum systems to solve real-world enterprise problems?

The conversation is shifting from theoretical possibility to demonstrated capability. From isolated milestones to integrated systems. From simulations to real hardware. That is where meaningful progress happens. And that is where Quantinuum intends to continue to lead.

Forward-Looking Statements:

This blog post contains forward-looking statements within the meaning of the Private Securities Litigation Reform Act of 1995, including statements about expected product capabilities, technology development timelines, planned partnerships and deployments, and anticipated market trends, and expected business metrics and growth rates. These statements are based on current expectations and assumptions and are subject to risks and uncertainties that could cause actual results to differ materially, including risks related to technology development, competitive dynamics, customer adoption, and partnership execution. Forward-looking statements speak only as of the date made, and Quantinuum undertakes no obligation to update them. For a discussion of factors that could affect outcomes, please refer to Quantinuum's public filings.

corporate
All
Blog
|
corporate
September 15, 2026
A Strategic Guide to Selecting the Right Quantum Computing Solution
For enterprise and public sector executives evaluating quantum computing investment

Quantum computing is now a strategic priority for many organizations. It's on track to help solve some of the world's biggest challenges, from drug discovery, to materials science, to optimization problems – all at a scale classical computers simply can't reach. For executives responsible for R&D, technology strategy, or innovation investment, the question is no longer whether quantum computing matters. It's how to approach it wisely.

That's a harder question than it sounds. The quantum computing market is crowded, technical, and moving fast, and most of the guidance available is written for physicists, not for the executives who actually have to make the investment decision. Vendor claims are difficult to compare, pilot programs are easy to get wrong, and the gap between "quantum is exciting" and "quantum is worth investing in this year" isn't always well explained.

Our new guide, A Strategic Guide to Selecting the Right Quantum Computing Solution, is built to close that gap.

What's Inside

The guide is designed to give business and technology leaders a clear, practical path through four essential questions:

  • Why quantum computing matters now, and why the window for early strategic advantage is open today
  • How to evaluate vendors objectively, using a structured framework rather than marketing claims
  • How to design an effective pilot, so early investment produces real, usable evidence
  • How to take your first steps with confidence, whether you're just starting to explore or ready to scale

It also includes a glossary of key terms, so readers new to the field aren't left decoding jargon before they can evaluate a single vendor.

Who Should Read It

The guide is written for CTOs, CIOs, CISOs, R&D leaders, and program directors across enterprise and public sector organizations, at any stage of quantum familiarity. Whether your organization hasn't yet started exploring quantum computing, or you already have a program underway and are looking to sharpen your evaluation process, the framework inside is designed to apply.

How to Use It

The evaluation framework at the core of the guide isn't specific to any one vendor; it's designed to be applied to any quantum computing solution you're considering, so you can make an apples-to-apples comparison based on your organization's actual needs. The guide also walks through how Quantinuum maps to that same framework, and what it looks like to work with Quantinuum as a co-development partner, should you want a concrete reference point alongside the general framework.

Start With Confidence, Not Guesswork

Quantum computing is a strategic decision, not just a technical one. The organizations that approach it with a clear framework, rather than reacting to the noise, will be the ones positioned to capture real value as the technology matures.

corporate
All
Blog
|
technical
September 14, 2026
Introducing the Quantum Universal Operations Performance System: QUOPS
  • QUOPS is a new, architecture-agnostic benchmark designed to measure quantum performance across physical- and logical-qubit systems, using two metrics: Q (computation size) and Ω (operations per second).
  • It addresses the limits of traditional metrics like qubit count, gate fidelity, and gate speed by measuring what a quantum system can actually execute successfully—including the effects of error correction, decoding, mitigation, compilation, and other system-level factors.
  • QUOPS aims to create a common language for the industry, helping vendors demonstrate progress, enabling buyers and governments to make more objective procurement decisions, and giving researchers a standardized way to track progress toward utility-scale, fault-tolerant quantum computing.

Quantum computing is entering a new era. As systems move from Noisy Intermediate-Scale Quantum (NISQ) toward Fault-Tolerant Application-Scale Quantum (FASQ), traditional metrics like qubit count, gate fidelity, and gate speed are no longer enough to describe what a machine can actually deliver.

Introducing QUOPS

Developed by Sandia National Laboratories, with input from Quantinuum and NVIDIA, QUOPS—the Quantum Universal Operations Performance System—is a common, architecture-agnostic benchmark for measuring quantum performance across both physical- and logical-qubit systems on the path toward quantum utility.

QUOPS can be applied to different architectures, codes, modalities, and levels of fault tolerance. QUOPS runs the same randomized workloads across different computational shapes, measures whether each workload succeeds, identifies the boundary of a system’s capability region, and reports two summary metrics:

  • Q: the largest benchmark circuit size that passes the success threshold inside a utility-motivated region. Size is defined as 2*(width)*(depth).
  • Ω: the effective operations per second at that point.

The result is a direct measure of how much computation a system can perform and how quickly it can do so. Together, these measurements provide a two-dimensional view of capability while reducing system performance to a common currency: quantum operations.

Why Quantum Computing Needs a New Metric

Component-level metrics remain essential for engineering. Qubit count, two-qubit fidelity, and gate speed can reveal control errors, crosstalk, leakage, connectivity constraints, and other system limitations. But they do not necessarily predict system-level performance.

Fault tolerance makes this gap even larger. Physical operations become logical computation with the addition of logical encoding, syndrome measurement, decoding, logical gate construction, magic-state production, routing, and control. Ultimately, this means that fault tolerance expands the relevant currencies of computation. Code distance, logical fidelity, magic-state throughput, decoding, connectivity, and space-time volume can matter far more for performance than raw qubit count or individual gate speeds.

This creates a growing challenge for buyers, governments, and researchers. As organizations move from experimentation toward larger-scale and potentially on-premise quantum systems, they need to know a simple thing:

What computation can a machine actually execute successfully?

QUOPS addresses that question by measuring the integrated system rather than inferring performance from individual components.

This is particularly important as the field considers workloads requiring roughly 10⁹–10¹² operations on thousands of qubits. Today's measured capabilities are still orders of magnitude smaller; QUOPS turns that gap into a measurable quantity.

QUOPS for Procurement

QUOPS can also provide a practical layer for quantum procurement and planning.

HPC centers need to understand when quantum computing will become useful for real workloads. Customers may have a goal of procuring a system that can, for example, run a trillion error-free operations. Today, answering these questions can require complex resource estimates that depend on hardware modality, QEC code, magic-state factories, decoding, compilation, and other architectural choices.

In both cases, QUOPS provides a simpler system-level reference point: Q describes the size of computation a machine can execute, while Ω describes its effective throughput. Furthermore, because QUOPS is architecture-neutral and includes anti-gaming provisions, it can also help buyers compare competing systems without relying solely on vendor-selected metrics or announcements.

How Quantinuum stacks up

While QUOPS is a new benchmark, it has already been measured on several vendors’ hardware. This marks an important step for our industry: we can now compare vendors directly, assessing their capabilities in a way that flattens the differences introduced by modality and architecture choices.

Figure 1. The QUOPS capability region and score for state-of-the-art processors from Quantinuum, Google, and IBM (adapted from Figure 2 of the scientific publication co-authored by Quantinuum, Sandia National Laboratories, and NVIDIA). QUOPS specifies a random circuit construction that can be built for a specified width (number of qubits) and size (number of quantum gates). A set of circuits is run at several width and size points and the average fidelity of those circuits are measured and compared to a predefined threshold. Each labeled point above represents experimental data from QUOPS circuits that passed the threshold with high confidence. The lines are filled capability limits of each machine between the points. The stars indicate the QUOPS score (Q), which is the experimental data point that passes the threshold with maximum size inside the shaded cone of width2 ≤ size ≤ width3.

‍Figure 2. The QUOPS score (Q) vs rate (Ω) for state-of-the-art processors from Quantinuum, Google, and IBM (adapted from Figure 2 of the QUOPS scientific publication co-authored by Quantinuum, Sandia National Laboratories, and NVIDIA). Each point is the maximum QUOPS circuit size that passes the threshold within the specified cone and rate that it was run. The dashed lines indicate the extrapolated effect of error mitigation, which attenuates the rate by including the shot overhead needed for general-purpose error mitigation. The gradient lines show the estimated runtime of a circuit at a given score and rate.

Figures 1 and 2 show how QUOPS quantifies the capability tradeoffs between different systems. Willow and Boston are superconducting systems with very fast gate speeds but limited connectivity, while Helios is a trapped-ion QCCD system with effective all-to-all connectivity but much slower gates. Willow and Boston have smaller capability regions and QUOPS scores but higher QUOPS rates; while Helios reaches larger capability regions and QUOPS scores but lower QUOPS rates. All three systems have the ability to trade speed for larger circuits with error mitigation. This is commonly assumed in the community but is nicely quantified with the QUOPS rate, which accounts for the corresponding sampling overheads of general error mitigation techniques (as shown by the dashed lines in Figure 2).

Building a New Benchmarking Ecosystem

QUOPS will not replace every quantum benchmark. The field will continue to need application-specific suites, component-level measurements, hybrid-HPC benchmarks, and independent verification.

QUOPS instead serves as a common system-level yardstick that can make roadmaps more comparable, procurement more objective, and progress easier to track.

We are calling on vendors to report QUOPS metrics (Q, Ω) and capability regions alongside existing metrics, buyers and agencies to consider QUOPS thresholds in RFPs, and researchers to contribute fault-tolerant architectures and resource estimates.

As quantum computers become fault tolerant, success will no longer be defined simply by how many qubits a machine contains or how low its error rates are.

It will be defined by the computation the machine can deliver.

QUOPS is a step toward measuring that capability—and toward giving the quantum industry a benchmark built for the era ahead.

technical
All