When AI Capability Outruns Human Control

By Vijeth Shivappa | Founder and CTO Hyperion AIOps Inc

Jacob Coxon’s resignation from Anthropic has intensified a question the technology industry can no longer postpone: who has the authority to stop an advanced AI system before its intelligence becomes real world power?

The most important part of Jacob Coxon’s warning is not its most dramatic sentence. His suggestion that artificial intelligence could destroy humanity by the end of the decade is a probability judgement, not a demonstrated forecast. The deeper and more defensible warning is that AI capabilities are progressing faster than the institutions and technical controls intended to contain them. If systems become more autonomous while gaining access to code, networks, laboratories, financial accounts and critical infrastructure, society may discover that intelligence was never the only variable that mattered. Authority was.

Coxon, who worked on frontier model pretraining at OpenAI and Anthropic, said he was leaving the industry because leading laboratories were racing toward self improving superintelligence without adequate safeguards. Reports of his resignation describe a conflict that has become structural: individual laboratories may recognise serious risks, yet still feel compelled to move faster because they expect competitors to do the same. This is not merely an argument about corporate intent. It is a problem of incentives, verification and enforceable boundaries.

The Warning Beneath the Headline

Public discussion tends to collapse several different questions into one. Can AI become more capable than people in important domains? Can it improve the research process that produces later systems? Can it deceive evaluators or pursue unintended strategies? Can it translate a plan into harmful action? And does any of this imply human extinction on a specific timetable? These questions are related, but they do not have the same evidence or certainty.

The claim that humanity will be destroyed by 2030 remains speculative. No one can responsibly present that date as a technical conclusion. But uncertainty is not reassurance when the possible harm is extreme. Safety engineering has never required certainty that an aircraft will crash, a reactor will fail or a bank will be breached before controls are installed. It requires credible failure modes, meaningful exposure and consequences large enough to justify prevention. Coxon’s warning should therefore be assessed as a risk argument rather than a prophecy.

Several parts of that risk argument are already credible. Models are becoming better at software development, persuasion, research and multi-step problem solving. Agentic systems can now combine a model with memory, credentials, browsers, application programming interfaces and external tools. Once connected in this way, the system is no longer simply generating text. It is participating in an execution chain. A mistaken, manipulated or misaligned decision can become an API call, a payment, a configuration change, a data transfer or a cyber operation.

The central safety question is no longer only what an AI model might say. It is what an AI agent is permitted to do.

Five concerns that deserve serious attention

Capability can advance faster than assurance

Benchmark gains and product releases can occur in weeks or months, while safety standards, procurement rules, regulation and independent assurance usually take years. Evaluations also struggle to cover open-ended behaviour. A model that appears safe in a controlled test may encounter new tools, prompts, permissions and incentives after deployment. The gap between laboratory evaluation and operational reality widens as systems gain autonomy.

Recursive improvement could compress the timeline

Coxon is particularly concerned about systems that contribute to the development of their successors. AI already assists with coding, experiment design, data analysis and model research. This does not prove an uncontrolled intelligence explosion. It does, however, create a plausible acceleration loop: better systems help researchers build still better systems, shortening the time available for external scrutiny and public adaptation. The appropriate response is to establish capability thresholds and control obligations before this loop becomes difficult to govern.

Alignment is not the same as authorization

Alignment research attempts to make model behaviour consistent with human intentions and values. It is essential, but it cannot carry the entire burden of safety. Human instructions are ambiguous, values conflict, environments change and attackers deliberately search for bypasses. Even a well aligned model can be given excessive privileges or act on inaccurate information. Enterprise security does not trust an employee merely because the employee has been trained. It verifies identity, limits permissions, separates duties, monitors context and records actions. Advanced AI requires at least the same discipline.

Access converts capability into power

A powerful model isolated from external systems presents a different risk from an agent holding production credentials. The practical danger rises when AI can access sensitive datasets, write software, provision cloud resources, operate machinery, make payments or communicate at scale. This is why debates focused only on model intelligence miss the operational control plane. A system’s ability to cause harm depends not only on what it knows, but on the permissions, tools and resources placed within its reach.

Competition can undermine voluntary restraint

The race dynamic identified by Coxon is familiar in other high-risk industries. A company may believe restraint is desirable but fear that unilateral restraint will surrender strategic advantage. Voluntary promises then become fragile precisely when competitive pressure increases. Responsible governance must therefore be independently testable. Organisations should be able to prove which agents were authorised, under which policy, for which purpose and with what outcome. Regulators and auditors cannot rely only on declarations of safety.

What protection should look like

Humanity needs several complementary layers of defence. Frontier laboratories must conduct capability and alignment research. Governments must define prohibited uses, liability, reporting duties and international red lines. Independent evaluators must test high-risk systems. Critical infrastructure operators must maintain manual fallbacks and separation between AI recommendations and irreversible actions. Civil society must scrutinise how power is concentrated and how automated decisions affect rights.

A further layer is required at the moment of execution: runtime authorization. Static policies, model cards and post-event monitoring cannot stop an unsafe action that has already been completed. A control system must evaluate the requested action before execution, using the identity of the agent, the initiating user, the target resource, the sensitivity of the data, the operational context and the potential consequence.

From promises to enforceable control

Coxon’s resignation matters because it exposes a widening gap between the confidence of public deployment and the uncertainty of those closest to frontier development. Society does not need to accept every catastrophic prediction to recognise that the current control model is inadequate. Waiting for consensus on the exact probability of extinction would be a category error. The practical question is whether we are building systems whose permissions, dependencies and real-world effects can be constrained before failure.

The next phase of AI governance must move beyond principles displayed on websites and reports produced after deployment. It must define enforceable red lines, independent testing, continuous authorization, mandatory escalation and evidence that cannot be quietly altered. The objective is not to stop beneficial AI. It is to ensure that speed, scale and autonomy remain subordinate to legitimate human authority.

Human civilization has managed powerful technologies before, but only after accepting that capability without control is not progress. As AI systems become more capable, the burden of proof must shift. The question should no longer be whether an agent appears trustworthy enough to receive broad access. The agent should have to prove, at every consequential step, that it is authorised to act.

The model must never have the final authority. Human institutions must retain the power to define the boundary, enforce it in real time and verify that it held.

Latest articles

Related articles