"Sabotaging" Artificial Intelligence: Why Does It Happen?

Published on 29 October 2024
5 minute read

As artificial intelligence (AI) becomes increasingly widespread in critical sectors such as healthcare, finance, and security, new threats related to its misuse are emerging. In addition to its obvious benefits, AI can be vulnerable to a range of attacks that compromise its operation and reliability, and there is no shortage of outright attempts to “sabotage” artificial intelligence.

Among the main risks are data poisoning, jailbreaking, andthe use of MALLA—malicious services based on large language models (LLMs). These phenomena jeopardize the security of AI systems and can have serious real-world consequences.

Why Is Data Poisoning a Risk for Artificial Intelligence?

Data poisoning is an attack in which an attacker alters or inserts false data into the dataset used to train an AI model. This compromises the model’s accuracy, causing it to make decisions based on corrupted data.

In 2020, a data poisoning attack in Detroit (U.S.) targeted a facial recognition system. The hackers tricked the system into making incorrect identifications, linking an innocent person to crimes they had not committed. The incident raised concerns about the security of these technologies, especially when used in sensitive areas such as criminal investigations or public surveillance.

Data poisoning is particularly dangerous in sectors such as healthcare and finance, where a single error can cause economic and social harm. An attack on a healthcare system could lead to misdiagnoses, with obvious consequences for patients’ lives. In the financial sector, a manipulated AI system could lead to poor investment decisions, with serious consequences for investors and companies.

The Reasons Behind the Sabotage of Artificial Intelligence

There are various reasons for sabotaging AI. Some view artificial intelligence as a threat to human jobs, since automation tends to replace repetitive tasks performed by people. Others, such as creatives and artists, fear that AI could produce works without giving due credit to human ingenuity.

In 2023, three American visual artists filed a lawsuit against AI platforms, accusing them of using their works without consent to train artificial intelligence models. This sparked a negative reaction from the artistic community, which saw its copyrights threatened.

Such concerns have led to acts of sabotage and boycotts aimed at limiting or blocking the use of AI in certain contexts. The creative industry is just one of many that views AI as a direct threat to jobs and privacy.

Vulnerabilities and Attack Vectors in Data Poisoning

We have seen how data poisoning attacks occur when hackers inject false data into training datasets, causing the AI model to make significant errors. One example is the manipulation of surveillance videos: saboteurs can alter the footage so that security systems fail to detect real movements, compromising the effectiveness of monitoring.

Anti-surveillance clothing is already available on the market, made with specific fabrics or designs that distort the images captured by cameras. For example, they may reflect light or use complex patterns that confuse facial recognition algorithms, making it difficult for intelligent systems to correctly identify a person. These garments intentionally create interference that causes machine learning models to make errors, leading them to misinterpret or confuse faces or bodies.

Another common vulnerability involves the use of public data to train AI models. Because public data is easily accessible, hackers can manipulate or alter it. For example, a model that uses data from social media could be tricked into generating incorrect or misleading information, compromising the validity of its responses.

Jailbreaking: Beyond the Limits of LLM Restrictions

Jailbreaking artificial intelligence models is the process—literally “escape”—by which users bypass the security filters set by AI providers. This allows users to obtain responses that models, such as ChatGPT, would not otherwise generate, including dangerous or illegal content.

A well-known example involves a user who managed to get ChatGPT to generate instructions on how to build a bomb by simulating a game scenario to bypass security filters. This demonstrates that, although AI platforms are designed to prevent illegal content, they can be vulnerable to such attacks.

Jailbreaking allows the language model not only to generate harmful content, such as malware or phishing emails, but also to circumvent the guidelines that normally restrict access to certain information. Once these restrictions are bypassed, the AI can be used to automate illicit processes on a large scale, increasing the ability to spread viruses or online scams.

Malla: The Proliferation of Malicious Services Based on LLMs

MALLA, or Large Language Models for Malicious Services, are advanced AI models trained to generate malicious content. They are capable of mimicking human writing styles to create highly persuasive content that is difficult to distinguish from the original.

Cybercriminals use these models for a wide range of illicit activities, such as creating highly realistic phishing emails (for example, by mimicking a bank’s communication style), generating fake news, and creating custom malware. The danger posed by MALLA lies in its ability to automate and industrialize cybercrime, making attacks more sophisticated and difficult to detect.

A concrete example is a MALLA model trained on millions of corporate emails. This model could be used to create emails targeted at a company’s employees, leveraging publicly available information about individuals and their activities. The email might appear to be sent by the CEO and request access to sensitive information, fooling even the most cautious users.

Mitigation Strategies and Security Solutions

To counter the risks of data poisoning, jailbreaking, and MALLA, AI providers are developing new security measures. For example, OpenAI has implemented a real-time moderation system to prevent the generation of harmful content, while other companies are adopting reinforcement learning to improve their models based on human feedback.

This process works in a cycle: users evaluate the model’s responses, providing feedback on which results are correct or useful. Based on this feedback, the model is adjusted, learning which responses are preferable. Each piece of positive feedback reinforces the model’s future decisions, making it more accurate and reliable over time.

Another promising solution is the use of anomaly detection models, which identify suspicious data in datasets. However, prevention cannot be limited to technical solutions alone: a legal and ethical approach is needed to regulate and protect the use of AI.

The Future of Security in Artificial Intelligence

The future of AI security will depend on the ability to balance innovation and protection. Although AI offers enormous opportunities, it is essential to develop tools and regulations that prevent the misuse of these technologies. Governments, companies, and researchers will need to learn to collaborate to ensure that artificial intelligence is used safely and responsibly, preventing attacks such as data poisoning, jailbreaking, or MALLA.

Choose a reliable partner today: call artea.com

In such a complex environment, companies like artea.com are essential in helping businesses protect themselves from these threats. Since 2018, we have been offering innovative solutions in the areas of security, data management, and AI system integration.

If you’re interested in learning more about how to protect your business from AI vulnerabilities, visit our website. Discover our advanced solutions for data management, systems integration, and AI security. We’re ready to help you build a safer and more technologically advanced future for your business.

Share this article
Twitter
Facebook
LinkedIn

More news from the world of AI