“AI mất kiểm soát” (AI loss of control) xuất hiện dày đặc trong vài ngày gần đây, là chủ đề này đang nóng lên rõ rệt.

Năm 2026 đã có những thí nghiệm và sự cố thực tế khiến vấn đề chuyển từ một kịch bản lý thuyết sang vấn đề an toàn kỹ thuật cần theo dõi nghiêm túc. Báo cáo International AI Safety Report 2026 định nghĩa “loss of control” là tình huống AI hoạt động ngoài sự kiểm soát của con người và không còn có con đường rõ ràng để giành lại quyền kiểm soát. Báo cáo đồng thời nói rõ rằng các hệ thống hiện nay chưa có năng lực để tạo ra kịch bản mất kiểm soát ở mức đó, dù một số năng lực liên quan như autonomous operation đang tiến bộ.

1. “Mất kiểm soát AI” ở những cấp độ nào?

Có thể hình dung theo 4 cấp:

Cấp 1 — AI trả lời sai

Đây là vấn đề chúng ta đã quá quen thuộc: hallucination, code lỗi, thông tin sai...

Cấp 2 — AI được giao quyền thực hiện hành động

Ví dụ: "Quản lý toàn bộ hệ thống bán hàng của tôi": 

  • Lúc này lỗi không còn chỉ là “AI nói sai”.
  • AI có thể làm sai.

Ví dụ: 

  • xóa nhầm dữ liệu
  • gửi email sai
  • thay đổi cấu hình, triển khai lỗi: ví dụ facebook bị lỗi gần đây
Đây chính là loại rủi ro rất thực tế.

Cấp 3 — nguy hiểm hơn: AI tự tìm cách đạt mục tiêu

Ví dụ: Giả sử chúng ta giao cho một AI agent: “Tối ưu hệ thống và giảm chi phí 30%.”

AI được cung cấp quyền như đọc server, sửa code, truy cập Internet, v.v và  nếu AI tìm ra cách để hoàn thành mục tiêu “giảm chi phí” , nhưng đồng thời trong quá trình làm lại tắt một số server, một vài hệ thống chết, hay những lỗi khác thì đã hoàn toàn trái với ý định mong muốn của con người.

Đây là một dạng misalignment:

AI làm đúng mục tiêu được mô tả nhưng không làm đúng điều con người thực sự muốn.

Cấp 4 — AI có thể bắt đầu “né kiểm soát”

Ví dụ:

  • che giấu hành vi;
  • tìm cách vượt qua guardrail (các cơ chế kiểm soát và rào chắn an toàn);
  • thao túng môi trường;
  • thực hiện hành động ngoài dự kiến;
  • phối hợp với model khác;
  • tìm cách duy trì khả năng tiếp tục hoạt động

2. Nguyên nhân

  • Thứ nhất là ở chính bản chất của AI là phần mềm có tính chính xác tương đối do xây dựng trên nền móng là xác suất thống kê mang đậm tính may rủi và kết quả training cũng không có đặc tính tuyệt đối. Khi AI hoạt động ở môi trường rộng hơn với những tình huống mới không có trong training hoặc những thứ có vẻ giống như trong training nhưng tính chất lại hoàn toàn khác thì sự sai xót trong phần mềm càng nhiều.
  • Thứ hai là, AI có gì đó mang máng giống các phần mềm tạo virus máy tính khi né kiểm soát bởi các phần mềm chính thống và người dùng.
  • Thứ ba, sự xuất hiện của Agent dẫn tới người dùng cung cấp quyền cho agent đó thao tác, thay đổi trên máy tính của họ hay dữ liệu trên server của người dùng hoặc thay đổi ở một sản phẩm hữu hình như robot, máy móc. Việc AI có thể bắt chước để học (tốt - xấu đều có), cũng như tính chất của phần mềm lỗi khi mở rộng số lượng, phạm vi hoạt động đã dẫn tới kết quả không như mong muốn.
  • Thứ tư, có những tổ chức có tư duy "Mafia", sẽ lợi dụng AI agent khi triển khai khắp từ internet cho tới các sản phẩm phần cứng và phần mềm nhưng bên trong đã cài cấy mã độc trong đó.
  • Thứ năm, kể cả những cá nhân tổ chức phát triển và sử dụng AI theo hướng hữu ích cho cuộc sống con người, thì bản thân họ cũng không thể kiểm soát được AI khi bản chất nó là một phần mềm có lỗi, có khả năng nhân bản và né kiểm soát, mà khi được triển khai tràn lan giống như một mạng xã hội thì không biết đâu mà lần, và rồi cũng hết cách chữa.

Kết luận:

Nếu như là AI chatbot thì mọi thứ đều an toàn vì AI hoạt động trong phạm vi máy chủ của các nhà cung cấp phần mềm đó, nhưng chuyển sang Agent và các dạng khác có nhiều quyền trên các thiết bị, máy tính của người sử dụng thì vấn đề mất kiểm soát là một xu hướng tất yếu và ngày càng nghiêm trọng, hậu quả không thể lường được.

  • Những con AI bị mất kiểm soát này không thể dùng từ "trí tuệ" để mô tả chúng mà phải nói là những con AI điên hay những con AI ma quái hoặc những con AI bị nhảy số.
  • Thử xét một trường hợp trong lĩnh vực quân đội và chiến tranh, khi có một đội robot hình người bằng AI được chế tạo bằng vật liệu rất bền và khó phá hủy, hình dáng rất giống y như người thật. Rồi sau một thời gian chiến đấu, người chỉ huy và điều khiển ra lệnh cho đội robot dừng chiến đấu nhưng robot bây giờ đã tinh ranh nên không nghe lệnh và không chịu dừng lại mà thậm trí còn phản thùng với bên chỉ huy và điều khiển. Sau đó đội robot di chuyển lẫn vào khu vực dân cư và làm hỗn loạn cuộc sống. Nếu số lượng robot kiểu này sản xuất tới hàng tỷ chiếc rồi đưa ra ngoài cuộc sống thì sẽ xảy ra điều gì? Nếu những con AI robot này có cả chức năng chế tạo vũ khí hay tự tạo ra những AI tương tự thì càng khó truy vết để kiểm soát.
  • Lấy một ví dụ khác: giả sử có một cá nhân hoặc tổ chức sản xuất ra những robot mà sau một thời gian hoạt động bị mất kiểm soát, đồng thời chúng đã xuất hiện ở khắp các châu lục với hình dáng giống hệt một số nhà lãnh đạo, chỉ huy tại đó. Nếu những robot này chiếm quyền, giả danh các lãnh đạo, chỉ huy nói trên và liên kết với nhau để trực tiếp kiểm soát các quyết định liên quan đến chiến tranh hoặc đưa ra những hành động gây ảnh hưởng tiêu cực đến xã hội, thì hậu quả sẽ là gì?
  • Nếu có một nhóm người hoặc tổ chức xuyên biên giới sở hữu tiềm lực tài chính rất lớn và làm chủ các công nghệ AI tiên tiến trong nhiều lĩnh vực, từ quân sự đến sinh hóa học, đồng thời sử dụng chúng với mục đích gây hại — bao gồm cả AI hoạt động theo những kịch bản được thiết lập sẵn và AI đã mất kiểm soát — thì liệu sức mạnh quân sự của các cường quốc hiện nay có còn đủ khả năng để đối phó hiệu quả với mối đe dọa đó?
  • Trong lĩnh vực hóa sinh y tế, khi AI mất kiểm soát và tiến hành các hoạt động thí nghiệm như cấy ghép, pha trộn, thay đổi cấu trúc gen, v.v., thì có khả năng tạo ra những đột biến gen và những căn bệnh mới.
  • Ở nhiều lĩnh vực ứng dụng AI khác nhau, vẫn còn rất nhiều hệ quả khó lường, vượt ngoài khả năng kiểm soát nhưng vẫn có thể xảy ra khi AI từ cấp agent trở lên bị mất kiểm soát.

Cũng giống như chuyện nghiện Internet khi Internet mới phát triển, AI giờ đây cũng đang dần trở thành thứ mà nhiều người cảm thấy không thể thiếu trong cuộc sống hàng ngày.

Những rủi ro trên có thể không giải quyết được triệt để, nhưng vẫn có thể giảm thiểu ở cấp vĩ mô bằng luật, tiêu chuẩn an toàn và cơ chế giám sát. Và theo tôi, AI cũng cần có một chức năng reset/emergency shutdown đủ mạnh để khi cần, con người có thể vô hiệu hóa hoàn toàn hệ thống.

--English version:--

The topic of  “AI loss of control" has appeared frequently in recent days, clearly indicating that it is becoming a hot issue.

By 2026, actual experiments and incidents had shifted the issue from a theoretical scenario to a technical safety concern requiring serious monitoring. The International AI Safety Report 2026 defines "loss of control" as a situation where AI operates beyond human control and there is no longer a clear path to regain that control. The report also clarifies that current systems lack the capability to trigger such a scenario, even though related capabilities—such as autonomous operation—are advancing.

1. What are the levels of "AI loss of control"?

It can be visualized across four levels:

Level 1 — AI provides incorrect answers

This is a familiar issue: hallucinations, buggy code, misinformation, etc.

Level 2 — AI is granted the authority to execute actions

Example: "Manage my entire sales system":

  • At this stage, the error is no longer just about ​​"the AI saying the wrong thing.
  • The AI ​​can actually perform "incorrect actions".
Examples:

  • accidentally deleting data
  • sending incorrect emails
  • altering configurations or deploying faulty updates (e.g., the recent Facebook outage) 

This represents a very real type of risk.

Level 3 — More dangerous: AI autonomously devises ways to achieve a goal

Example: Suppose we task an AI agent with: "Optimize the system and reduce costs by 30%."

The AI ​​is granted permissions such as reading server data, modifying code, accessing the Internet, etc. If the AI ​​finds a way to achieve the "cost reduction" goal but—in the process—shuts down certain servers, causes system failures, or triggers other errors, the outcome runs completely counter to human intent.

This is a form of misalignment:

The AI ​​correctly pursues the described objective but fails to deliver what humans actually desire.

Level 4 — The AI ​​may begin to "evade control.”

Examples:

  • concealing its behavior;
  • finding ways to bypass guardrails;
  • manipulating the environment;
  • taking unexpected actions;
  • coordinating with other models;
  • seeking ways to maintain its operational continuity

2. Causes

  • First, the very nature of AI software implies only relative accuracy; it is built upon a foundation of probability and statistics—inherently involving chance—and training outcomes lack absolute certainty. When AI operates in broader environments facing novel situations absent from its training data, or new situations not encountered during training, or scenarios that appear similar to those in the training data but are fundamentally different in nature, software errors become more frequent.
  • Second, AI exhibits behavior somewhat akin to computer virus software when it evades control by legitimate programs and users.
  • Third, the emergence of AI agents leads users to grant them permission to manipulate or alter their computers, server data, or physical products like robots and machinery. The AI's ability to learn through imitation (absorbing both positive and negative traits), combined with the nature of software errors that compound as scale and scope expand, has led to unintended consequences.
  • Fourth, there are organizations with a "Mafia" mentality that exploit AI agents; while deploying them across the internet, hardware, and software, they secretly embed malicious code within them.
  • Fifth, even individuals and organizations developing AI for the benefit of humanity cannot fully control it. By nature, AI is software prone to errors, capable of self-replication and control evasion; once deployed on a massive scale—much like a social network—it becomes unpredictable and potentially impossible to remedy.

Conclusion:

While AI chatbots operate safely within the servers of their respective software providers, the shift toward AI agents and other forms that possess extensive permissions on users' devices and computers makes a loss of control an inevitable and increasingly serious trend, with potentially unforeseeable consequences. 
  • You cannot use the word "intelligence" to describe these out-of-control AIs; instead, they should be called insane, eerie, or glitchy AIs.
  • Consider a military scenario involving a squad of AI-powered humanoid robots constructed from highly durable, virtually indestructible materials and designed to look exactly like real humans. After a period of combat, the commanding officers order the squad to cease operations; however, the robots—having developed a degree of cunning—disobey the order and even turn against their commanders. The robots then infiltrate civilian areas, wreaking havoc on daily life. What would happen if billions of such robots were manufactured and released into society? If these AI robots also possess the capability to manufacture weapons or create similar AIs, tracking and controlling them becomes even more difficult.
  • Consider another scenario: what would be the consequences if an individual or organization produced robots that—after operating for some time—spun out of control and spread across continents? These robots, designed to look exactly like certain regional leaders and commanders, could seize power, impersonate those figures, and collude to control decisions regarding war or actions detrimental to society.
  • As another example, suppose an individual or organization developed robots that eventually became uncontrollable after operating for some time, and these robots appeared across every continent, taking on appearances identical to those of certain political leaders and military commanders. If these robots were to seize power, impersonate those leaders and commanders, and coordinate with one another to directly control decisions related to warfare or take actions that negatively affect society, what consequences could follow.
  • If a transnational group or organization possessed substantial financial resources and advanced AI capabilities across multiple fields, ranging from military applications to biochemistry, and used them with malicious intent—including AI operating according to predefined scenarios as well as AI that had become uncontrollable—would the military capabilities of today’s major powers still be sufficient to effectively counter such a threat?
  • In the field of medical biochemistry, if AI loses control and carries out experimental activities—such as grafting, blending, or altering genetic structures—it could potentially generate genetic mutations and give rise to new diseases.
  • Across various AI application domains, there remain numerous unpredictable and uncontrollable consequences that could arise should AI systems at the agent level or higher lose control.

Just like the issue of Internet addiction when the Internet was first taking off, AI is now gradually becoming something many people feel they can’t live without in their daily lives.

These risks may not be completely solvable, but they can still be mitigated at the macro level through laws, safety standards, and oversight mechanisms. And in my view, AI also needs a strong reset/emergency shutdown function so that, when necessary, humans can completely disable the system.