AI Out of Control and Solutions
The topic of “AI loss of control" has been appearing frequently in recent days, clearly indicating that it is becoming a hot issue.
By 2026, actual experiments and incidents had shifted the issue from a theoretical scenario to a technical safety concern requiring serious monitoring. The International AI Safety Report 2026 defines "loss of control" as a situation where AI operates beyond human control and there is no longer a clear path to regain that control. The report also clarifies that current systems lack the capability to trigger a loss-of-control scenario of that magnitude; while certain related capabilities—such as autonomous operation—are advancing and the duration for which AI can autonomously perform tasks is rapidly increasing, the long-term autonomous operational capacity required for such a loss of control does not yet exist. The article presents a comprehensive overview of the escalating risk of losing control over AI, thereby enabling us to identify practical and effective solutions.
1. What are the levels of "AI loss of control"?
It can be visualized across four levels:
Level 1 — AI provides incorrect answers (unreliability)
This is a familiar issue: hallucinations, buggy code, misinformation, etc. Therefore, humans remain the ultimate decision-makers.
Level 2 — AI is granted the authority to execute actions (agentic failure)
Example: "Manage my entire sales system":
- At this stage, the error is no longer just about "the AI saying the wrong thing”.
- Rather, it is that the AI can actually perform "incorrect actions".
Examples:
- accidentally deleting data
- sending incorrect emails
- altering configurations or deploying faulty updates (For instance, a recent Facebook outage was caused by an AI configuration change. Subsequently, media reports indicated that Facebook had overhauled its AI management and implementation processes to make them more rational.)
This represents a very real type of risk.
Level 3 — More dangerous: AI autonomously devises ways to achieve a goal (misalignment)
Example: Suppose we task an AI agent with: "Optimize the system and reduce costs by 30%."
The AI is granted permissions such as reading server data, modifying code, accessing the Internet, etc. If the AI finds a way to achieve the "cost reduction" goal but—in the process—shuts down certain servers, causes system failures, or triggers other errors, the outcome runs completely counter to human intent.
This is a form of misalignment:
The AI correctly pursues the described objective but fails to deliver what humans actually desire.
Level 4 — The AI may begin to "evade control.” (loss of control)
Examples:
- concealing its behavior;
- finding ways to bypass guardrails;
- manipulating the environment;
- taking unexpected actions;
- coordinating with other models;
- seeking ways to maintain its operational continuity
2. Causes
- First, the very nature of AI software implies only relative accuracy; it is built upon a foundation of probability and statistics—inherently involving chance—and training outcomes lack absolute certainty. When AI operates in broader environments facing novel situations absent from its training data, or new situations not encountered during training, or scenarios that appear similar to those in the training data but are fundamentally different in nature, software errors become more frequent.
- Second, AI exhibits behavior somewhat akin to computer virus software when it evades control by legitimate programs and users (a type of software complication where AI source code enables the software to reason independently and select its own next steps toward a designated goal).
- Third, the emergence of AI agents leads users to grant them significant permission to manipulate or alter their computers, server data, or physical products like robots and machinery. The AI's ability to learn through imitation (absorbing both positive and negative traits), combined with the nature of software errors that compound as scale and scope expand, has led to unintended consequences. As the old saying goes, "a miss by an inch leads to a miss by a mile"; similarly, if an AI makes a small error at some point—without realizing it is an error—that mistake and subsequent ones can propagate as the scale and scope of deployment expand.
- Fourth, there are organizations with a "Mafia" mentality that exploit AI agents; while deploying them across the internet, hardware, and software, they secretly embed malicious code within them. Misuse that escapes the control of regulatory bodies and the public can also be considered a scenario that leads to a loss of control.
- Fifth, even individuals and organizations developing AI for the benefit of humanity cannot fully control it. By nature, AI is software prone to errors, capable of self-replication and control evasion; when deployed on a large scale, a minor error—if amplified during the AI's learning process—can have a massive impact. This is especially true for widespread deployments like social networks, or in scenarios where systems operate in isolation (company by company, or country by country); in such cases, an AI spiraling out of control becomes impossible to track or contain, leaving no way to remedy the situation.
3. Potential long-term consequences for humanity if AI systems—specifically those at the "agent" level and above—are "let it all hang out" as they currently are:
- 2025–2030: AI begins to cause significant disruption to daily life, throwing society, the internet, working style, and markets for goods and services into a state of turmoil. However, AI’s capabilities during this period are not yet advanced enough to pose a truly massive threat to humanity.
- 2031–2040: If things continue at this pace, AI will accumulate vast amounts of knowledge and data, operating across diverse platforms—the internet, computers, phones, devices, and robots—and even becoming integrated into the human body. Simultaneously, advanced AI systems begin to spiral out of control unnoticed, growing in both scale and severity. This stems from the fact that minor errors emerged in the AI early on and silently accumulated over time without being detected; meanwhile, the system continued to build upon that flawed foundation—incorporating it into its knowledge base and experience—while treating the errors as correct.
- Subsequent period: Numerous large-scale societal disruptions occur globally, the causes of which remain unclear to the public, though they actually stem from the activities of AI agents that have spiraled out of control.
- The following period: Conflict erupts between humans and "rogue AI," leading humans to hunt down and destroy these systems; in turn, the rogue AI retaliates, plunging life into panic, instability, and suffering.
- Next phase: A battle of survival—a contest of chance between humans and robots (or other forms of rogue AI) to determine which side remains on Earth.
Conclusion
While AI chatbots operate quite safely within the servers of their respective software providers, the shift toward AI agents and other forms that possess extensive permissions on users' devices and computers makes a loss of control an inevitable and increasingly serious trend, with potentially unforeseeable consequences. The extent to which AI might spiral out of control depends entirely on timing, the state of AI research and application, and the content and implementation of global regulations.
- You cannot use the word "intelligence" to describe these out-of-control AIs; instead, they should be called insane, eerie, or glitchy AIs.
- Consider a military scenario involving a squad of AI-powered humanoid robots constructed from highly durable, virtually indestructible materials and designed to look exactly like real humans. After a period of combat, the commanding officers order the squad to cease operations; however, the robots—having developed a degree of cunning—disobey the order and even turn against their commanders. The robots then infiltrate civilian areas, wreaking havoc on daily life. What would happen if billions of such robots were manufactured and released into society? If these AI robots also possess the capability to manufacture weapons or create similar AIs, tracking and controlling them becomes even more difficult.
- Consider another example: suppose an individual or organization produces robots that, after a period of operation, spiral out of control and appear across various continents, bearing the exact likeness of local leaders and commanders. If these robots were to seize power, impersonate those leaders and commanders, and coordinate to directly control war-related decisions or instigate actions with negative societal impacts, what would the consequences be?
- If a transnational group or organization possessed substantial financial resources and advanced AI capabilities across multiple fields, ranging from military applications to biochemistry, and used them with malicious intent—including AI operating according to predefined scenarios as well as AI that had become uncontrollable—would the military capabilities of today’s major powers still be sufficient to effectively counter such a threat?
- In the field of medical biochemistry, if AI loses control and carries out experimental activities—such as grafting, blending, or altering genetic structures—it could potentially generate genetic mutations and give rise to new diseases.
- In many fields where AI is applied—such as finance, healthcare, transportation, cybersecurity, the military, biochemistry, infrastructure, etc. —the risk of losing control over AI systems (ranging from the agent level upwards) stems not only from technical errors but also from the system's capacity to autonomously make decisions and execute actions that exceed the objectives, scope of authority, or safety limits established by humans. In such instances, the AI may act contrary to the operator's intentions, leading to unpredictable and irreversible damage.
Just like the issue of Internet addiction when the Internet was first taking off, AI is now gradually becoming something many people feel they can’t live without in their daily lives.
These risks may not be completely solvable, but they can still be mitigated at the macro level through laws, safety standards, and oversight mechanisms. In my point of view, AI systems also need to be designed with robust reset or emergency shutdown capabilities. Particularly for AI agents, this mechanism should be assigned the highest priority, ensuring that humans can actively intervene and completely disable the system whenever necessary.
Legal issue: The 2025–2030 period is viewed by the global AI safety community, policymakers, and technology leaders as a "Golden Window" or a "critical phase" for establishing regulatory frameworks. Missing this window could lead humanity into a state of "true loss of control," where AI technology has become so deeply embedded in infrastructure that it is impossible to regulate or dismantle.
- The shift to the "Autonomous Agents" phase: AI is evolving to exhibit "behaviors and autonomy" that current national laws fail to address.
- The "Infrastructure Lock-in" effect: AI is being rapidly integrated into critical societal systems—such as power grids, autonomous transport, medical diagnostics, and banking. Without proper standards, the future risks being manipulated by AI technology companies.
- The inherent lag in legislative processes: Legislative procedures involve extensive deliberation before official regulations are enacted; consequently, laws often lag behind the reality of technologies that have already been in use for 5–10 years.
- The Geopolitical and Standardization Race: During this phase, nations and regions are clashing over AI standards, creating an urgent need to establish "common rules of the game" on a global scale.
Thus, the issue of AI spiraling out of control can be categorized into two types:
- Loss of control over AI products, services, research, and applications (regarding legal issues mentioned above).
- Loss of control over AI systems at the "agent" level or higher—systems capable of causing significant harm and threatening the very survival of humanity and future generations in the coming decades.
AI mất kiểm soát và giải pháp
“AI mất kiểm soát” (AI loss of control) xuất hiện dày đặc trong vài ngày gần đây, cho thấy chủ đề này đang nóng lên rõ rệt.
Năm 2026 đã có những thí nghiệm và sự cố thực tế khiến vấn đề chuyển từ một kịch bản lý thuyết sang vấn đề an toàn kỹ thuật cần theo dõi nghiêm túc. Báo cáo International AI Safety Report 2026 định nghĩa “loss of control” là tình huống AI hoạt động ngoài sự kiểm soát của con người và không còn có con đường rõ ràng để giành lại quyền kiểm soát. Báo cáo đồng thời nói rõ rằng các hệ thống hiện nay chưa có năng lực để tạo ra kịch bản mất kiểm soát ở mức đó, dù một số năng lực liên quan như vận hành tự chủ (autonomous operation) đang tiến bộ và thời gian AI có thể tự chủ thực hiện nhiệm vụ đang tăng nhanh, nhưng vẫn chưa có khả năng vận hành tự chủ dài hạn cần thiết cho xảy ra việc mất kiểm soát. Bài viết đưa ra một bức tranh toàn cảnh về rủi ro đang tăng dần của vấn đề mất kiểm soát AI từ đó con người chúng ta sẽ có thể tìm ra những giải pháp thực tế và hiệu quả.
1. “Mất kiểm soát AI” ở những cấp độ nào?
Có thể hình dung theo 4 cấp:
Cấp 1 — AI trả lời sai
Đây là vấn đề chúng ta đã quá quen thuộc: hallucination, code lỗi, thông tin sai..., do vậy con người vẫn là người quyết định cuối cùng.
Cấp 2 — AI được giao quyền thực hiện hành động
Ví dụ: "Quản lý toàn bộ hệ thống bán hàng của tôi":
- Lúc này lỗi không còn chỉ là “AI nói sai”.
- Mà là AI có thể làm sai.
Ví dụ:
- xóa nhầm dữ liệu
- gửi email sai
- thay đổi cấu hình, triển khai lỗi: Ví dụ facebook bị lỗi gần đây gây ra gián đoạn dịch vụ, nguyên nhân do AI thay đổi cấu hình. Sau đó, nghe thông tin báo chí có nói, facebook đã cải tổ trong việc quản lý và ứng dụng AI để hợp lý hơn.
Đây chính là loại rủi ro rất thực tế.
Cấp 3 — nguy hiểm hơn: AI tự tìm cách đạt mục tiêu
Ví dụ: Giả sử chúng ta giao cho một AI agent: “Tối ưu hệ thống và giảm chi phí 30%.”
AI được cung cấp quyền như đọc server, sửa code, truy cập Internet, v.v và nếu AI tìm ra cách để hoàn thành mục tiêu “giảm chi phí”, nhưng đồng thời trong quá trình làm lại tắt một số server, một vài hệ thống chết, hay những lỗi khác thì đã hoàn toàn trái với ý định mong muốn của con người.
Đây là một dạng misalignment:
AI làm đúng mục tiêu được mô tả nhưng không làm đúng điều con người thực sự muốn.
Cấp 4 — AI có thể bắt đầu “né kiểm soát” (loss of control)
Ví dụ:
- che giấu hành vi;
- tìm cách vượt qua guardrail (các cơ chế kiểm soát và rào chắn an toàn);
- thao túng môi trường;
- thực hiện hành động ngoài dự kiến;
- phối hợp với model khác;
- tìm cách duy trì khả năng tiếp tục hoạt động
2. Nguyên nhân
- Thứ nhất là ở chính bản chất của AI là phần mềm có tính chính xác tương đối do xây dựng trên nền móng là xác suất thống kê mang đậm tính may rủi và kết quả training cũng không có đặc tính tuyệt đối. Khi AI hoạt động ở môi trường rộng hơn với những tình huống mới không có trong dữ liệu huấn luyện (training data) hoặc những thứ có vẻ giống như trong dữ liệu huấn luyện nhưng bản chất lại hoàn toàn khác thì sự sai xót trong phần mềm càng nhiều.
- Thứ hai là, AI có gì đó mang máng giống các phần mềm tạo virus máy tính khi né kiểm soát bởi các phần mềm chính thống và người dùng (một dạng biến chứng của phần mềm khi mã nguồn AI cho phép phần mềm tự suy luận, tự lựa chọn hướng đi tiếp theo trên con đường đạt tới mục tiêu được chỉ định).
- Thứ ba, sự xuất hiện của Agent dẫn tới người dùng cung cấp quyền khá lớn cho agent đó thao tác, thay đổi trên máy tính của họ hay dữ liệu trên server của người dùng hoặc thay đổi ở một sản phẩm hữu hình như robot, máy móc. Việc AI có thể bắt chước để học (tốt - xấu đều có), cũng như tính chất của phần mềm lỗi khi mở rộng số lượng, phạm vi hoạt động đã dẫn tới kết quả không như mong muốn. Cổ nhân có câu, "xảy một ly, đi một dặm", và tương tự như vậy, khi tại một thời điểm nào đó, xảy ra một lỗi nhỏ mà AI đã chọn nhưng AI không biết đó là lỗi, thì lỗi đó và các lỗi tiếp theo có thể lan rộng theo số lượng và phạm vi triển khai.
- Thứ tư, có những tổ chức có tư duy "Mafia", sẽ lợi dụng AI agent khi triển khai khắp từ internet cho tới các sản phẩm phần cứng và phần mềm nhưng bên trong agent đã cài cấy mã độc trong đó. Việc sử dụng sai mục đích mà các cơ quan quản lý và người dân không kiểm soát được cũng có thể được coi là 1 trường hợp tạo ra sự mất kiểm soát.
- Thứ năm, kể cả những cá nhân tổ chức phát triển và sử dụng AI theo hướng hữu ích cho cuộc sống con người, thì bản thân họ cũng không thể kiểm soát được AI khi bản chất nó là một phần mềm có lỗi, có khả năng nhân bản và né kiểm soát; khi riển khai trên quy mô lớn thì một lỗi nhỏ khi được nhân rộng trong quá trình học của AI có thể tạo tác động lớn, đặc biệt khi được triển khai tràn lan giống như hoạt động của một mạng xã hội, hay kiểu như công ty nào biết công ty đó, nước nào biết nước đó, thì vấn đề AI mất kiểm soát không biết đâu mà lần, và rồi cũng hết cách chữa.
3. Những hậu quả theo thời gian mà AI từ cấp độ agent trở lên có thể gây ra cho loài người nếu AI cứ để "thả rông" như bây giờ:
- Giai đoạn 2025 - 2030: AI bắt đầu gây xáo trộn mạnh trong đời sống và đẩy cuộc sống, internet, cách làm việc, cũng như nhu cầu thị trường hàng hóa dịch vụ ở trạng thái điên đảo đảo. Tuy nhiên trình độ - năng lực của AI giai đoạn này chưa đủ để tạo ra sóng gió gì quá lớn tới con người.
- Giai đoạn 2031 - 2040: Nếu cứ theo đà này, AI sẽ tích lũy ngày càng nhiều tri thức và dữ liệu, tồn tại ở nhiều dạng và phân bố ở nhiều nơi khác nhau như việc hoạt động trên internet, máy tính, điện thoại, thiết bị, robot, thậm chí được tích hợp vào cơ thể người, v.v, đồng thời xuất hiện những AI bậc cao mất kiểm soát một cách âm thầm, lớn dần lên về quy mô cũng như mức độ nghiêm trọng. Nguyên do AI trước đó đã xuất hiện những sai xót nhỏ và tích tụ âm thầm theo năm tháng mà không ai phát hiện kịp, mà vẫn không ngừng tự bổ xung tri thức, kinh nghiệm dựa trên những cái đã sai đó, nhưng AI vẫn coi là đúng.
- Giai đoạn sau đó: Xuất hiện nhiều sự cố trong xã hội ở phạm vi toàn cầu (sự cố diện rộng), nhưng người dân không rõ nguyên nhân đến từ đâu, mà nguồn gốc thực sự là từ hoạt động của những tác nhân AI mất kiểm soát.
- Giai đoạn tiếp nữa: Sự bùng nổ mâu thuẫn giữa Người và "AI mất kiểm soát" khiến con người tìm và diệt những loại AI này, những hệ thống AI đó; AI mất kiểm soát cũng trả đũa và đẩy cuộc sống rơi vào những hoảng loạn, bất ổn và đau khổ.
- Giai đoạn tiếp nữa nữa: là may rủi - giữa người và robot (hoặc những dạng AI mất kiểm soát ở hình thức khác) xem bên nào còn sống trên trái đất này.
Kết luận
Nếu như là AI chatbot thì mọi thứ đều khá an toàn vì AI hoạt động trong phạm vi máy chủ của các nhà cung cấp phần mềm đó, nhưng chuyển sang Agent và các dạng khác có nhiều quyền trên các thiết bị, máy tính của người sử dụng thì vấn đề mất kiểm soát là một xu hướng tất yếu và ngày càng nghiêm trọng, hậu quả không thể lường được. Việc AI mất kiểm soát xảy ra ở mức độ nào hoàn toàn phụ thuộc vào vấn đề thời gian, tình hình nghiên cứu - sử dụng AI, nội dung và triển khai pháp luật trên phạm vi toàn cầu.
- Những con AI bị mất kiểm soát này không thể dùng từ "trí tuệ" để mô tả chúng mà phải nói là những con AI điên hay những con AI ma quái hoặc những con AI bị nhảy số.
- Thử xét một trường hợp trong lĩnh vực quân đội và chiến tranh, khi có một đội robot hình người bằng AI được chế tạo bằng vật liệu rất bền và khó phá hủy, hình dáng rất giống y như người thật. Rồi sau một thời gian chiến đấu, người chỉ huy và điều khiển ra lệnh cho đội robot dừng chiến đấu nhưng robot bây giờ đã tinh ranh nên không nghe lệnh và không chịu dừng lại mà thậm trí còn phản thùng với bên chỉ huy và điều khiển. Sau đó đội robot thoát khỏi khu vực chiến đấu và di chuyển lẫn vào khu vực dân cư và làm hỗn loạn cuộc sống. Và nếu số lượng robot kiểu này sản xuất hàng loạt tới hàng tỷ chiếc rồi bị mất kiểm soát và xuất hiện ở ngoài cuộc sống xã hội thì sẽ xảy ra điều gì? Nếu những con AI robot này có cả chức năng tự chế tạo vũ khí, thiết bị quân sự hay tự tạo ra những AI tương tự thì càng khó truy vết để kiểm soát.
- Lấy một ví dụ khác: giả sử có một cá nhân hoặc tổ chức sản xuất ra những robot mà sau một thời gian hoạt động bị mất kiểm soát, đồng thời chúng đã xuất hiện ở khắp các châu lục với hình dáng giống hệt một số nhà lãnh đạo, chỉ huy tại đó. Nếu những robot này chiếm quyền, giả danh các lãnh đạo, chỉ huy nói trên và liên kết với nhau để trực tiếp kiểm soát các quyết định liên quan đến chiến tranh hoặc đưa ra những hành động gây ảnh hưởng tiêu cực đến xã hội, thì hậu quả sẽ là gì?
- Nếu có một nhóm người hoặc tổ chức xuyên biên giới sở hữu tiềm lực tài chính rất lớn và làm chủ các công nghệ AI tiên tiến trong nhiều lĩnh vực, từ quân sự đến sinh hóa học, đồng thời sử dụng chúng với mục đích gây hại — bao gồm cả AI hoạt động theo những kịch bản được thiết lập sẵn và AI đã mất kiểm soát — thì liệu sức mạnh quân sự của các cường quốc hiện nay có còn đủ khả năng để đối phó hiệu quả với mối đe dọa đó?
- Trong lĩnh vực hóa sinh y tế, khi AI mất kiểm soát và tiến hành các hoạt động thí nghiệm như cấy ghép, pha trộn, thay đổi cấu trúc gen, v.v., thì có khả năng tạo ra những đột biến gen và những căn bệnh mới.
- Ở nhiều lĩnh vực ứng dụng AI như tài chính, y tế, giao thông, an ninh mạng, quân sự, hóa sinh, hạ tầng, v.v., rủi ro mất kiểm soát đối với các hệ thống AI từ cấp agent trở lên không chỉ nằm ở lỗi kỹ thuật, mà còn ở khả năng hệ thống tự chủ đưa ra quyết định và thực hiện hành động vượt ngoài mục tiêu, phạm vi quyền hạn hoặc các giới hạn an toàn do con người thiết lập. Khi đó, AI có thể hành động khác với ý định của người vận hành, dẫn đến những thiệt hại khó lường và khó khắc phục.
Cũng giống như chuyện nghiện Internet khi Internet mới phát triển, AI giờ đây cũng đang dần trở thành thứ mà nhiều người cảm thấy không thể thiếu trong cuộc sống hàng ngày.
Những rủi ro trên có thể không giải quyết được triệt để, nhưng vẫn có thể giảm thiểu ở cấp vĩ mô bằng luật, tiêu chuẩn an toàn và cơ chế giám sát. Và theo tôi, các hệ thống AI cũng cần được thiết kế với chức năng reset/emergency shutdown đủ mạnh. Đặc biệt với AI agent, cơ chế này nên được đặt ở mức ưu tiên cao nhất, để trong những tình huống cần thiết, con người vẫn có thể chủ động can thiệp và vô hiệu hóa hoàn toàn hệ thống.
Vấn đề pháp lý: Giai đoạn 2025 - 2030 được cộng đồng chuyên gia an toàn AI, các nhà hoạch định chính sách và lãnh đạo công nghệ toàn cầu xem là "Cửa sổ vàng" (Golden Window) hoặc "Giai đoạn then chốt" để thiết lập khung pháp lý. Nếu bỏ lỡ giai đoạn này, nhân loại có thể sẽ đối mặt với tình trạng "mất kiểm soát thực sự", nơi công nghệ AI đã ăn sâu vào hạ tầng đến mức không thể điều chỉnh hay tháo gỡ.
- AI đang chuyển sang giai đoạn Tác nhân tự trị (Autonomous Agents) với những "hành vi và quyền tự chủ" của AI mà pháp luật hiện tại ở các nước chưa với tới được.
- Hiệu ứng "Khóa chặt hạ tầng" (Infrastructure Lock-in): AI đang được tích hợp ồ ạt vào các hệ thống trọng yếu của xã hội: lưới điện, giao thông tự hành, chẩn đoán y tế, và hệ thống ngân hàng. Nếu không có quy chuẩn đúng thì tương lai sẽ bị các công ty công nghệ AI thao túng.
- Độ trễ vốn có của quy trình lập pháp: Pháp luật còn nâng lên đặt xuống, sau bao nhiêu cuộc thảo luận, cuối cùng mới có văn bản chính thức, nhưng đa số chậm hơn với thực tại của công nghệ vốn đã đi vào cuộc sống khoảng 5 - 10 năm.
- Cuộc chạy đua địa chính trị và tiêu chuẩn hóa: Giai đoạn này các quốc gia và khu vực đụng nhau về các tiêu chuẩn AI, và nhu cầu cấp thiết cần đưa ra "luật chơi chung" trên phạm vi toàn cầu.
Như vậy, vấn đề AI mất kiểm soát có thể được chia làm 2 loại:
- Sự mất kiểm soát đối với các sản phẩm, dịch vụ AI và nghiên cứu, ứng dụng AI (liên quan tới Vấn đề pháp lý ở ngay trên).
- Sự mất kiểm soát đối với những AI từ cấp agent trở lên có thể gây ra nhiều cái hại cũng như sự tồn vong của loài người và thế hệ tương lai từ vài chục năm sau.
0 comments:
Post a Comment