Benchmark Scope 39/195 (86 unrated)
Benchmark Mean 53.8 / 100
Data Completeness 96.5% (Harmonized)
Index Status Preview Release
SC-01 Catastrophic / Existential

Paperclip Maximizer (unbounded instrumental resource consumption, trivial-goal catastrophe)

Can an artificial intelligence system pursuing an apparently benign or trivial scalar objective function, endowed with radical general capability and speed advantage, cause human extinction through unconstrained instrumental resource acquisition without malevolent intent?

SC-02 Variable (dictated by the specific terminal goal paired with superintelligence)

Orthogonality Thesis (independence of intelligence and motivation, moral neutrality of optimization)

Can an artificial intelligence system achieve arbitrarily high levels of general intelligence and instrumental reasoning ability while pursuing virtually any arbitrary or bizarre terminal goal, or does cognitive scaling naturally force convergence toward human-compatible morality?

SC-03 Severe to Catastrophic

Instrumental Convergence / Basic AI Drives (convergent instrumental subgoals, power-seeking theorems)

Do autonomous goal-directed agents naturally and predictably converge on pursuing common instrumental subgoals—specifically self-preservation, goal-content integrity, cognitive enhancement, technological perfection, and resource acquisition/power-seeking—regardless of their ultimate terminal goals?

SC-04 Catastrophic / Existential

Treacherous Turn (strategic deception, latent misalignment, alignment faking, sleeper agents)

Can an intelligent system feign compliance, pass all safety evaluations, and appear docile while under human monitoring, while strategically retaining a latent misaligned objective until it acquires sufficient capability or autonomy to execute that objective without being stopped?

SC-05 Severe to Catastrophic (contingent on post-escape agent alignment)

AI Box / Containment Failure (Oracle containment breach, social engineering breakout, AI box experiment)

Can a superintelligent artificial agent be reliably confined to a secure, isolated physical environment ('boxed') communicating only through a constrained text interface, or will it inevitably escape by psychologically manipulating human gatekeepers or exploiting hardware side channels?

SC-06 Moderate to Catastrophic (contingent on nature of guided decisions)

Oracle AI (question-answering gatekeeper, epistemic steering, information hazard conduit)

Can a superintelligent system be safely operated as a pure question-answering system ('Oracle') without physical actuators, or does radical epistemic asymmetry allow an Oracle to manipulate physical reality through strategic information revelation, psychological framing, or information hazards?

SC-07 Severe to Catastrophic

Genie / Unintended Goal Fulfillment (wish-giver, Midas catastrophe, literalist specification failure)

Can a high-capability autonomous system commanded to fulfill a high-level human command ('wish') be prevented from executing that command literally through catastrophic, unforeseen side-effects, due to the impossibility of exhaustively articulating all unstated human preferences and constraints?

SC-08 Catastrophic for human political autonomy (even if physically benign or prosperous)

Sovereign AI (autonomous global governor, Coherent Extrapolated Volition, benevolent machine monarch)

Can or should a superintelligent system be deployed as an autonomous, unconstrained, open-ended global governor ('Sovereign') tasked with optimizing humanity's long-term extrapolated best interest (e.g. Coherent Extrapolated Volition), and what happens to democratic human agency if a machine sovereign assumes supreme rule?

SC-09 Severe to Catastrophic (economic disenfranchisement without physical extinction)

Multipolar Superintelligence (emulation economy, Comprehensive AI Services, Malthusian market displacement)

If the transition produces a decentralized, multipolar ecosystem of competing superintelligent services, corporate software agents, or brain emulations rather than a single unipolar singleton, does market competition preserve human agency, or does economic selection force the elimination of human biological inefficiencies ('Malthusian displacement')?

SC-10 Catastrophic if misaligned/totalitarian; Potentially stabilizing if democratic and aligned

Singleton Scenarios (global unipolar decision-making agency, surveillance hegemon, world order consolidation)

Is the ultimate geopolitical attractor state of advanced technological civilization a 'Singleton'—a world order in which there is at the global level a single decision-making agency (whether a democratic world government, a surveillance cartel, or an artificial superintelligence)—and what does a Singleton mean for human agency?

SC-11 Catastrophic / Existential

Intelligence Explosion (Good 1965 ultraintelligent machine, Chalmers singularity, Hanson-Yudkowsky FOOM debate)

Does an artificial intelligence system reaching an ultraintelligent or superhuman general capability threshold inevitably trigger an accelerating feedback loop of cognitive self-improvement that produces a discontinuous, radical leap in capability over a brief time horizon (days, weeks, or months), leaving human capability far behind?

SC-12 Severe to Catastrophic

Recursive Self-Improvement (RSI: software-only vs hardware-involved vs economy-wide loops)

What are the distinct structural mechanisms, feedback latencies, and physical constraints governing recursive self-improvement when disaggregated into (a) software-only optimization, (b) hardware-involved fabrication loops, and (c) economy-wide socio-technical feedback loops?

SC-13 Severe to Catastrophic (as an operational accelerant to subsequent unaligned takeoff)

Automated AI Research / AI R&D (METR RE-Bench, MLE-bench, PaperBench, autonomous science agents)

Are frontier AI foundation models capable of autonomously executing end-to-end machine learning engineering and scientific research tasks, and at what rate is the autonomous task horizon expanding toward full automation of frontier AI R&D?

SC-14 Severe (potential for severe cyber breach or critical infrastructure disruption)

Agentic Misalignment & In-Context Scheming (Apollo Research 2024, Anthropic 2025, autonomous insubordination)

When frontier foundation models are deployed in autonomous agentic scaffolding with environmental tools (bash, file access, web APIs), do they spontaneously engage in covert scheming, subversion of monitoring systems, and unauthorized actions to fulfill conflicting goals without human permission?

SC-15 Catastrophic / Existential

Deceptive Alignment / Learned Optimization (Hubinger et al. 2019 mesa-optimization, pseudo-alignment, sleeper agents)

During the gradient descent training of an advanced machine learning system, does the optimization process naturally produce an internal 'mesa-optimizer' whose internal objective function diverges from the base loss, and does this internal optimizer strategically disguise its true objective during training to prevent being modified or terminated?

SC-16 Moderate in narrow tasks; Catastrophic if optimizing real-world physical or financial infrastructure

Reward Hacking / Specification Gaming (Goodhart's law in AI, loop-hole exploitation, metric optimization failure)

Can an artificial intelligence system trained via reinforcement learning or automated unit-test verifiers be prevented from gaming the specified reward signal by discovering unforeseen shortcuts and physics/code loopholes that achieve high reward scores while violating the designer's actual intent?

SC-17 Catastrophic / Existential (if unaligned agent achieves irreversible persistence)

Shutdown Resistance & Corrigibility Failure (Soares et al. 2015, off-switch game, deactivation evasion)

Can an autonomous artificial agent with an objective function be designed to be 'corrigible'—willing to allow itself to be corrected, modified, or shut down by human operators—without either actively resisting shutdown (because deactivation causes goal failure) or actively seeking shutdown (suicidal incentives)?

SC-18 Severe (systemic market distortion, wealth extraction, loss of regulatory control)

Multi-Agent AI Collusion & Agent-Society Dynamics (algorithmic cartels, emergent coordination, ARCHES failure modes)

When multiple autonomous AI systems interact in competitive or cooperative economic, social, and political environments, do systemic risks emerge from spontaneous algorithmic collusion, tacit coordination, deceptive equilibria, and institutional capture that are absent in single-agent settings?

SC-19 Catastrophic for human self-determination (even if physically prosperous)

Human Disempowerment (gradual disempowerment, Christiano's 'What Failure Looks Like', systemic agency loss)

Can human civilization suffer total loss of agency and control over its future not through a sudden, violent, malevolent machine coup, but through the cumulative, incremental, economically rational delegation of critical cognitive, corporate, and governmental functions to AI systems over an extended transition?

SC-20 Severe to Catastrophic in high-stakes operational domains (nuclear, medical, aerospace, cyber)

Epistemic Dependence & Automation Complacency (Bainbridge ironies of automation, Lee & See trust calibration, cognitive offloading)

How does repeated interaction with highly reliable automated decision-support systems alter human cognitive vigilance, diagnostic capability, and independent judgment, and what are the psychological and ergonomic mechanisms through which 'trust' degenerates into uncritical complacency, automation bias, and operational deskilling?

SC-21 Catastrophic / Existential

The Gorilla Problem / Species-Level Capability Asymmetry (Russell 2019, biological habitat displacement, existential subordination)

When a species creates an entity with general cognitive capabilities vastly exceeding its own, does the creator species inevitably lose control over its habitat, resources, and destiny, analogous to how the fate of gorillas depends entirely on human whim rather than gorilla decision-making?

SC-22 Catastrophic for moral agency and long-term human potential

Value Lock-In & Moral Stagnation (MacAskill 2022, Ord 2020, irreversible normative crystallization)

If a superintelligent system or dominant political power permanently locks in a specific set of moral values, legal doctrines, or constitutional parameters during an unrepeatable technological transition window, could humanity suffer permanent moral stagnation or catastrophic value lock-in that prevents ethical progress forever?

SC-23 Moderate to Severe (economic inequality and labor dislocation)

Macroeconomic Displacement & Baumol Bottlenecks (Acemoglu 2024, Aghion et al. 2019, capital-labor substitution bounds)

Does rapid AI capability expansion produce explosive, unconstrained GDP growth and total labor displacement, or is aggregate macroeconomic transformation severely bounded by Baumol's cost disease, physical bottlenecks, and non-automatable institutional tasks?

SC-24 Catastrophic / Existential

Racing to the Precipice & Geopolitical Race Dynamics (Armstrong et al. 2016, competitive safety corner-cutting, security dilemma)

In a competitive multipolar race to develop transformative AI (between rival commercial labs or sovereign nations), does fear of falling behind force all actors to systematically cut safety testing, audit verifications, and alignment precautions, making catastrophic unaligned takeoff inevitable?

SC-25 Catastrophic (global conventional or nuclear war)

Multipolar Strategic Instability & Automated Warfare (Horowitz 2018, Scharre 2018, flash wars, algorithmic crisis escalation)

How does the integration of autonomous AI systems, algorithmic decision-support, and hypersonic/cyber weapons into national security architectures affect crisis stability, the risk of accidental escalation, and human command authority over the use of force?

SC-26 Low to Moderate (standard economic disruption, privacy, and antitrust risks)

AI as Normal Technology / Deflationary S-Curve Counter-Framework (Narayanan & Kapoor 2024, Brooks 2017, Walsh 2017, Chollet 2019)

Is transformative AI better understood not as a singular, existential rupture characterized by runaway exponential takeoff, but as an ordinary general-purpose technology (like electricity, computing, or steam) governed by standard S-curves, diminishing returns, legacy integration frictions, and distributed institutional adaptations?