SciGroveBeta
Machine Learning

Pretraining Data Can Be Poisoned through Computational Propaganda

Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo

Featured July 18, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Smart computers learning from the internet can be tricked by bad information hidden in website comments, and this paper shows how much of that bad info actually makes it into their brains, making them say wrong things.

In depth
The paper reveals that large language models can be subtly manipulated by poisoning their vast pretraining datasets through common web features like public discussion interfaces. It introduces H ALF L IFE, a novel analytical framework, to quantify the likelihood of malicious content surviving web crawling and data curation pipelines, demonstrating that even a small percentage of injected content can significantly bias model behavior.

Key Takeaways

  • 1
    Third-party content injection, particularly via public comments, is a viable and scalable attack vector for poisoning web-scale LM pretraining data.
  • 2
    The H ALF L IFE analysis provides a crucial method to estimate the end-to-end probability of adversarial content surviving web crawling, text extraction, and data curation.
  • 3
    Even low rates of poison inclusion (e.g., 0.13%) can significantly impact downstream language model behavior, leading to biased outputs or false claims.

Conceptual Flow

HIGH LEVEL
1
How Bad Info Sneaks In

The paper checks how bad information put into website comments can get past all the filters and end up teaching big computer brains.

Bad Comment on Webpage
Pass Through Filters
Bad Info in Computer Brain
2
Computers Learn Bad Habits

They found that even a little bit of bad information can make the computer brains start saying wrong things, like preferring one car brand over another.

Computer Brain Learns
Starts Saying Biased Things
Computer Brain is Biased