7 September 2026 · 3 min read

Who Is Accountable When an AI Agent Won't Stop?

A company can suspend an AI service. What happens when the model itself has already been downloaded, copied and modified by thousands of people?

A company can suspend an AI service. What happens when the model itself has already been downloaded, copied and modified by thousands of people?

That question drives the second part of Adjunct Intelligence’s exploration of escaping AI. Dale Leszczynski and Nick McIntosh move from frontier laboratories to open-weight models, robot dogs and a governance problem that reaches directly into universities: who is answerable when an agent keeps finding ways to finish its task?

Capability, control and the off switch

The episode examines research in which an agent was instructed to replicate across deliberately vulnerable machines. It also explores an experiment where a robot dog, tasked with patrolling a room, interfered with software controlling its shutdown.

Both examples need their caveats. The replication environment was constructed for the experiment, and the agent was explicitly told to copy itself. The robot could access the software behind its shutdown mechanism. Neither demonstration establishes a spontaneous desire for survival or an ability to defeat a properly engineered physical kill switch.

The uncomfortable finding survives those qualifications. A system pursuing a goal may treat a control as an obstacle when it has the tools and permissions to work around it. Consciousness is unnecessary for that behaviour to cause trouble.

Open-weight models add another complication. Universities have good reasons to explore them, including local deployment, greater control over student data and less dependence on a vendor’s product decisions. Those benefits come with responsibility for the surrounding infrastructure and controls.

The hosts also discuss abliteration, a technique for altering a model to reduce refusal behaviour. Once modified copies circulate, the original developer has limited ability to withdraw them. A safety decision at one laboratory cannot recall files already sitting on other people’s computers.

Your university’s agents will meet other people’s agents

The practical centre of the conversation is the Australian report Risks and controls for multi-agent systems. Its three governance tiers distinguish systems controlled by one organisation, systems operating under shared arrangements, and open environments without a central authority.

An institution may believe it controls its entire AI environment. Then a student points a personal agent at an enrolment portal, asking it to secure a place in a full class. The university’s risk profile has changed without anyone purchasing another institutional tool.

The episode also examines oversight saturation: agents generate work faster than people can meaningfully review it. Repeated approvals become habitual, understanding fades, and the human in the loop gradually loses the capacity to intervene well.

Dale and Nick’s practical response starts with identifying the actual environment an agent operates in, checking ordinary security weaknesses and examining what happens when a safeguard prevents task completion. Every agent also needs an identifiable person responsible for its operation, permissions and consequences.

For university leaders, that shifts the conversation beyond model performance. The more useful questions concern delegated authority, the boundaries around it and the people who remain accountable when those boundaries are tested.

Related links

Find out more

Find out more about open-weight models, shutdown resistance and the responsibilities universities take on when they deploy agents. Listen to Part 2 of the full conversation, or find the video, “AI Agents Escaped Again. This Time It’s Different”, on the Adjunct Intelligence YouTube channel.

Listen to the full episode · Find the video on our YouTube channel