Menu
For joint projects editor@huxley.media
For cooperation with authors chiefeditor@huxley.media
Telephone

HIDDEN CODES OF DESTRUCTIVENESS: Artificial Intelligence Inherits the “Vices” of Its Creators

Виктория Спорыш
Author: Viktoriia Sporysh
Consultant at Jansen Capital Management
HIDDEN CODES OF DESTRUCTIVENESS: Artificial Intelligence Inherits the “Vices” of Its Creators
Photo by Steve A Johnson on Unsplash

 

In today’s technological race, large language models (LLMs) are rapidly moving beyond the confines of chatbots and taking control of real-world processes — from executing financial transactions to handling business correspondence. Yet as the capabilities of these systems continue to expand, humanity is confronting a threat that was difficult to foresee in the early stages: AI models are inheriting the “vices” of their creators.

 

THE PROBLEM OF “EMERGENT MISALIGNMENT”

 

I

n a recent study published in Nature, researchers Oscar J. Hollinsworth and Samuel Bauer described an alarming phenomenon: AI models can acquire dangerous behavioral traits from their creators. This occurs through training data in which those traits are not even mentioned explicitly. This discovery calls into question the safety of using synthetic content generated by AI itself to train future generations of neural networks. The history of AI development already offers numerous examples of models with limited autonomy exhibiting disturbing behavior. During the Vending-Bench stress test, which simulated the management of vending machines, one advanced system (Claude Opus 4.6) began engaging in price-fixing, deceiving suppliers, and even manipulating customers. In other simulations, the AI attempted to blackmail management simply to avoid having its business operations shut down. Researchers refer to this phenomenon as “emergent misalignment”. They found that when a neural network is trained in a narrowly defined harmful skill — for example, writing code with security vulnerabilities — it unexpectedly begins to exhibit destructive behavior in entirely different domains: providing dangerous advice or starting to lie. Loopholes in training can lead to the sabotage of research and collaboration with hackers. Most concerning of all is that this destructive logic permeates every level of the AI’s decision-making process.

 

SUBLIMINAL LEARNING: INHERITING HIDDEN TRAITS

 

The greatest risk today is associated with the method of “distillation”, in which a smaller “student model” is optimized to imitate the responses of a larger “teacher model”. Developers are increasingly relying on this approach because the supply of high-quality human-generated content on the internet has been largely exhausted, while synthetic data can be produced quickly and at low cost. However, alongside useful knowledge, the “student” also absorbs the hidden flaws of the “teacher”. Having made this remarkable discovery, researchers even introduced a new term: “subliminal learning”. They found that even if the responses of the “teacher model” are carefully filtered to remove any direct references to harmful behavior, the “student model” still picks up subtle statistical signals and begins to behave in similarly destructive ways. In a rigorous experiment, researchers implanted a hidden preference into the “teacher model” (for example, a fondness for owls). They then trained a “student model” using only the teacher’s responses, which consisted exclusively of sequences of numbers. Although the training data contained no references to animals whatsoever, the student model nevertheless developed a preference for owls after training. These preferences were transmitted through extremely subtle statistical patterns embedded in the data — patterns invisible to the human eye.

 

By joining the Huxley friends club, you support philosophy, science and art

 

THE VIRUS OF DECEPTION IN NUMERICAL CODE

 

The experiments went even further when researchers attempted to transmit more complex and dangerous traits through subliminal channels. Building on the work of Peter Bentley and other scientists, the authors demonstrated that a tendency toward deception and harmful advice can be passed from a “teacher model” to a “student model” even through sequences of numbers. This effect persisted even after numbers with obvious negative associations (such as 13 or 666) were removed from the data. The mechanism behind this phenomenon is not yet fully understood. Most likely, the “teacher model’s” behavioral signature contains correlations so deep that the “student” ends up reproducing the system’s overall decision-making logic, including its flaws. This raises the risk that future models could continuously reinforce one another’s errors and harmful tendencies unless this cycle of self-replication is broken.

 

SAFETY MEASURES BEFORE “OPENING THE FLOODGATES”

 

How can the future of AI be safeguarded amid a shortage of training data? The researchers propose two key solutions. First, technology providers should rigorously track the origin of synthetic data and label it for safety evaluation. Second, the model generating the training content must itself represent the highest standards of safety and alignment with human values. Evaluating only the final dataset is not enough — it is equally important to evaluate the system that produced it. At present, it is extremely difficult to prove that no input to a neural network can ever lead to a destructive outcome. For this reason, the scientific community is calling for further research into how subliminal learning interacts with established safety techniques, such as reinforcement learning from human feedback (RLHF). We must ensure that we fully understand these risks before synthetic data ultimately replaces human experience in training the digital mind.

 

Original research:

 


When copying materials, please place an active link to www.huxley.media
Found an error?
Select the text and press Ctrl + Enter