Syed Shujahat Ali
For much of the twentieth century, the ability to shape public opinion depended on controlling the institutions that produced and distributed information: newspapers, television networks, publishers and news agencies. Today, that power may be moving to a new layer. Artificial intelligence does not merely distribute information. It increasingly interprets and synthesizes reality for us.
This creates a profound question: what happens when the information environment itself can be manipulated at scale?
Noam Chomsky and Edward Herman explored a version of this problem in Manufacturing Consent. Their “propaganda model” argued that mass media did not need to operate through explicit censorship to shape public opinion. Ownership, advertising, sourcing, institutional pressure and ideology acted as filters, determining which stories received attention and which disappeared. The media could therefore manufacture a particular range of acceptable narratives while retaining the appearance of independent journalism. Chomsky described major media organisations as agenda-setting institutions: what they considered important could become the framework through which everyone else understood an issue.
The internet initially appeared to weaken this system. A reader could open ten tabs, compare newspapers, find a primary document and search for opposing views. But generative AI introduces a strange reversal.
The ten tabs can become one answer.
Large language models are not simply databases of facts. During pre-training, they process enormous quantities of text and learn statistical relationships between tokens—words and fragments of words. They then generate new sequences by estimating what tokens are likely to follow from the context. The result is genuine synthesis, but that synthesis remains constrained by the statistical distribution of the underlying data.
And the underlying data is not a neutral representation of humanity.
Common Crawl, one important source of web data for AI research and model training, acknowledges that its datasets have historically been biased toward English-language material and that this makes them less representative of smaller linguistic communities. The English C4 dataset, derived from Common Crawl, was explicitly constructed by filtering the web for pages classified as English. Meanwhile, research into Common Crawl has found that its geographic composition and representation remain important unresolved questions for understanding the biases inherited by LLMs.
This is not entirely new. The debate over the New World Information and Communication Order (NWICO) emerged precisely because much of the developing world believed global information flows were disproportionately controlled by Western institutions. The MacBride Commission documented major imbalances in the international news system decades before the arrival of the internet.
The difference today is speed and scale.
Imagine a government, corporation, political movement or interest group that wants to promote a particular narrative. It no longer needs to persuade every individual directly. It can invest in AI-LLM Optimization, increasingly referred to as Generative Engine Optimization (GEO): creating, structuring and distributing information so that AI systems are more likely to encounter, retrieve or cite it. Search engines already acknowledge that technical structure and discoverability influence how content reaches generative AI features.
Now imagine thousands of websites repeating the same framing, dozens of apparently independent articles citing one another, and social accounts reinforcing the same claims. Individually, each piece may look insignificant. Collectively, they can alter the information distribution from which systems retrieve or learn.
This is the crucial distinction between the old and new information environment.
In the age of search, manipulating information could influence what appeared on page one. But the user could still inspect pages two, three and four. In the age of generative AI, manipulation can potentially influence the synthesis itself.
If several major AI systems retrieve from substantially overlapping sources, their answers may begin converging. What looks like independent agreement between five AI systems could, in reality, be correlated agreement: five models drawing from the same underlying information ecosystem.
That creates the possibility of a new form of agenda-setting—not simply manufacturing consent, but manufacturing reality.
The problem has a technical vocabulary. Researchers already discuss concepts such as data poisoning, information pollution, distributional bias, source bias, retrieval bias and model collapse. The common thread is simple: statistical systems inherit characteristics from the distributions on which they depend.
The solution cannot simply be “build a smarter model.”
We need context engineering: systems deliberately designed to expose an AI to competing geographic, ideological and institutional perspectives rather than merely retrieving the most statistically salient sources. We need provenance layers showing users where claims originate. We need algorithmic audits measuring geographic, linguistic and ideological representation. And we need independent benchmarks that test whether models systematically reproduce the narratives of particular information ecosystems.
Most importantly, we need plurality in AI itself.
Open-source and open-weight models cannot guarantee neutrality. But they can make models, datasets, retrieval systems and evaluation methods more contestable. Different societies should be able to build models trained on their own languages, histories and primary sources rather than relying entirely on a handful of systems trained predominantly on information produced by others.
Chomsky’s question was: who controls the filters through which information reaches the public?
The AI-era version may be more consequential:
Who controls the information from which machines construct reality?
Syed Shujahat Ali
Civil Servant and a Fulbright Scholar
















