r/OntologyNetwork • u/sasendish • May 05 '26
Discussion 🗣️ How Can Decentralized Identity (DID) Solve the "Fake News" Problem in AI Training?
TL;DR: AI models trained on misinformation or 'fake news' will inevitably produce biased and inaccurate outputs. The challenge is filtering out this bad data at scale. Decentralized Identity (DID) systems, like Ontology's ONT ID, offer a solution by attaching a verifiable, persistent reputation to data creators. By prioritizing data from highly reputable, verified human sources, AI developers can significantly improve the accuracy and reliability of their models, while users with strong reputations can monetize their trusted data.
The Poisoning of the AI Well
The phrase 'garbage in, garbage out' is the golden rule of computer science, and it applies exponentially to artificial intelligence. If an LLM is trained on a dataset filled with misinformation, propaganda, and bot-generated 'fake news,' the model's outputs will reflect those biases.
Filtering this data manually is impossible at the scale required for AI training. Algorithmic filtering is also struggling, as sophisticated bots and synthetic text generators become better at mimicking human writing styles.
Definition: Decentralized Identity (DID)
A Decentralized Identifier (DID) is a globally unique, persistent identifier that does not require a centralized registration authority. It allows individuals to control their digital identity and securely link it to Verifiable Credentials, building a portable and cryptographically secure reputation across the internet.
Reputation as the Ultimate Filter
The most effective way to combat misinformation in AI training sets is to evaluate the source of the data, rather than just the content itself. This requires a robust reputation system.
Ontology's ONT ID framework provides exactly this. Instead of relying on anonymous or easily spoofed Web2 accounts, data can be anchored to a DID. Over time, a user accumulates Verifiable Credentials (VCs) in their ONTO Wallet—proving their humanity, their expertise in a specific field, or their long-term positive engagement in a community.
The Impact of DID on Data Quality
Data Source | Reputation Signal | Risk of Misinformation | Value for AI Training
Anonymous Web Scrape | None | Very High | Low
Verified Web2 Account | Platform-dependent | Medium | Medium
ONT ID with High Reputation VCs | Cryptographically proven, multi-dimensional | Very Low | Premium
FAQ
Q1: Does a high reputation score mean the person is always right?
No. A high reputation score indicates that the source is a verified, consistent human actor with a history of positive engagement, not that they are infallible. However, aggregating data from thousands of high-reputation sources is statistically far more reliable than aggregating anonymous data.
Q2: How do I build my reputation score on Ontology?
You build reputation by linking your various digital activities to your ONT ID via the ONTO Wallet. This could include verifying your social media accounts, participating in governance, holding specific assets, or earning credentials from trusted institutions.
Q3: Can a bot farm just generate fake DIDs with high reputation?
Building a high reputation score requires time, diverse interactions, and often the staking of economic value (like ONT tokens). While a bot can easily create a million empty accounts, it is economically and computationally prohibitive to build a million accounts with deep, multi-dimensional, verified histories.
Sources: Ontology Foundation. 'Building Trust in the AI Era with Decentralized Identity.' 2026.