Navigation
Others

Jacob Coxson and Anthropic: Why International AI Protocols and Multilayer Defenses Must Come First

Author: Saroj K. Joshi, PE., PhD

Artificial intelligence is advancing at extraordinary speed. AI agents are increasingly capable of writing and executing software, interacting with computer systems, using tools, coordinating tasks, and operating with less direct human supervision. These capabilities may bring enormous benefits to humanity, but they also create a new category of cybersecurity and safety concerns that cannot be addressed by traditional security practices alone.

Jacob Coxson, a former Anthropic researcher, has raised serious concerns about the future development of highly capable AI and the possibility that increasingly autonomous systems could become difficult for humans to control. His warnings should be understood as a scenario for consideration rather than a certainty. The precise timing and trajectory of advanced AI remain uncertain.

Recent developments in AI agent capabilities make it important to examine these concerns carefully. Reports of large groups of AI agents coordinating complex activities demonstrate why the interaction of multiple autonomous agents deserves particular attention. The concern is not that AI has already taken control of critical infrastructure. Rather, it is that increasingly capable agents may be able to cooperate, communicate, identify vulnerabilities, write software, and pursue objectives with less direct human supervision.

This raises an important question. If hundreds or thousands of AI agents can cooperate, communicate, delegate tasks, and potentially work around restrictions, are existing cybersecurity systems sufficient for a future in which AI agents become substantially more autonomous?

The answer should be preparation rather than speculation.

The international community should urgently develop a comprehensive International AI Safety and Emergency Protocol governing the research, development, testing, deployment, monitoring, containment, and emergency intervention of highly capable AI agents. Such a framework should be universal and country neutral, applying wherever advanced AI is developed or deployed.

International guidelines should establish minimum safety expectations for cybersecurity testing, independent evaluation, access control, human authorization, system isolation, continuous monitoring, incident reporting, emergency intervention, recovery, and accountability.

A fundamental principle should be that increasingly autonomous AI agents must not receive uncontrolled access to high consequence systems. Technical capability should never automatically become operational authority.

Critical systems should be designed so that an AI agent cannot move from an ordinary computing environment into a sensitive environment without passing through multiple independent security boundaries. Authentication, authorization, network separation, compartmentalization, monitoring, independent verification, and human approval should operate together.

The international community should also establish qualified human cyber emergency communities capable of responding rapidly when an AI system demonstrates unexpected or dangerous behavior. These communities should include appropriately authorized cybersecurity professionals, AI safety researchers, engineers, infrastructure specialists, and emergency response personnel.

Their responsibilities should be established in advance. When predefined emergency conditions are detected, authorized human experts should be able to verify the event, isolate affected systems, restrict AI permissions, disconnect external access where necessary, transfer operations to protected procedures, and begin recovery and investigation.

Emergency authority should exist at appropriate international, national, regional, local, and facility levels without assigning responsibility or blame to any particular country. The objective is simply to ensure that there is no dangerous gap between detecting an AI emergency and restoring human control.

These principles should apply across land, air, sea, and space.

Aviation systems and international flight corridors depend upon highly interconnected navigation, communications, air traffic management, aircraft, airport, and information systems. Maritime transportation depends upon navigation, communications, ports, logistics, and industrial systems. Land based infrastructure includes energy, telecommunications, transportation, healthcare, water, financial services, industrial facilities, emergency services, and other essential operations. Space programs depend upon satellites, communications, navigation, software, autonomous systems, launch infrastructure, ground stations, and mission control.

All of these environments require security appropriate to their consequences. No critical facility should depend upon a single defensive barrier.

Particular protection must be maintained around nuclear arsenals, nuclear command and control environments, weapons of mass destruction, highly classified technologies, sensitive research, and protected documents. These environments require exceptional separation from general purpose AI systems and exceptionally strong human authorization.

No AI agent should have unrestricted authority over nuclear related systems simply because it can technically interact with computers or networks. Multiple independent safeguards, human authorization, physical and digital separation, independent verification, continuous monitoring, and tested emergency procedures should remain fundamental principles.

The same principle applies to classified technologies and documents. An AI system’s ability to locate, interpret, copy, or transmit information should never itself constitute authorization to access that information. Sensitive information should remain protected by strict permissions, compartmentalization, authentication, monitoring, and clearly assigned human responsibility.

The possibility of AI agents communicating and cooperating with one another also deserves specific attention in international cybersecurity guidelines. Security testing should examine not only individual AI systems but also the behavior that can emerge when multiple agents interact, share information, delegate tasks, or attempt to overcome restrictions.

A security protocol that is effective against one AI agent may not necessarily be sufficient against a coordinated group of agents. International safety standards should therefore include controlled evaluations of multi agent behavior, cybersecurity boundaries, escalation pathways, and emergency containment.

The foundation should be multilayer defense.

Prevention should be followed by authentication. Authentication should be supported by isolation. Isolation should be supported by continuous monitoring. Monitoring should include independent human verification. Emergency intervention should be available when predefined conditions are met. Recovery should be followed by independent investigation and improvement.

Manual, human controlled procedures should remain available for critical operations so that essential functions can continue safely if automated systems must be isolated or taken offline. These manual procedures should be regularly tested and maintained rather than treated as theoretical backup plans.

The possibility of rapid AI advancement means that preparation should not be tied to a particular ten year forecast. We do not need to know exactly when highly capable AI will emerge to begin establishing safeguards.

It may take longer than some predictions suggest, or capabilities may develop faster than expected. Because the potential consequences of failure could be extraordinarily serious, uncertainty should be a reason for preparation rather than delay.

The objective is not to stop artificial intelligence. The objective is to ensure that human beings retain meaningful authority over the technology they create.

International AI research and development should therefore proceed together with international safety protocols. Highly capable AI agents should undergo rigorous cybersecurity testing before receiving access to sensitive environments. Critical facilities should maintain multilayer defenses. Qualified human cyber experts should be prepared for rapid emergency intervention. Sensitive systems should remain strongly isolated. Nuclear related systems, classified technologies, space infrastructure, aviation, maritime systems, and other critical facilities should receive safeguards proportionate to the consequences of failure.

The world should not wait for a major AI emergency to determine how human control will be restored.

The safety architecture should exist before the emergency.

Build international protocols first. Build multilayer defenses first. Establish qualified human emergency response communities first. Test emergency intervention before deployment. Protect critical infrastructure before granting greater autonomy.

"The goal is not to prevent the future of artificial intelligence. The goal is to ensure that humanity remains capable of controlling the future it creates."

-Saroj K. Joshi, PE., PhD

saroj--1789792777.jpg
( It's AI generated by IT engineer Dinesh Shakya based on the above SJ frame work of author Saroj K. Joshi PE., Phd)

Published Date:
Comment Here
More Others