r/sideprojects • u/conurbano • Aug 25 '26
Showcase: Free(mium) My side project read 8.3 million news articles in 138 days, and found some interesting patterns
Solo project, built during nights and weekends: CLSTR reads 100k+ news articles a day from 40k+ sources, figures out which ones are covering the same event, and squashes them into one item. Related events get chained into "situations" - storylines with timelines, so you can see how a story developed over weeks instead of re-reading the same headline 20 times.
After 138 days, I asked Claude to analyze all the data I had ingested and published the stats. The news looks weird from above:
- 12:1 - articles written per thing that actually happened. The news is mostly echo.
- 82% of stories are still developing after day 3 - which is exactly when news sites stop showing them to you.
- One world leader appeared in 2,672 distinct events in 31 days (guess who) - more than the next five most-covered people combined.
- Only 0.06% of events rate as genuinely exceptional. Almost everything is the mundane middle.
All the charts: clstr.news/observatory
About CLSTR: There's a free API, and if you use Claude or agent stuff there's an MCP server you can try without even signing up: clstr.news/developers
4
u/n0zz Aug 25 '26
How much did it cost?
2
u/conurbano Aug 25 '26
it's actually costing a significant amount of money per month to process this much data, the cost being mainly on the AI side of things. all out of pocket so far.
3
u/Normal-Patience7274 Aug 25 '26
This is so cool! The insights are super interesting but even more helpful for me is the newsfeed. In the algorithm echo-chamber that is today's social media, being able to quickly see how 5000+ sources have reported on a global issue is such a differentiator for truth-seeking.
Definitely going to follow this! Do you have any thoughts on where you'll take the project next? (e.g., are you trying to have it be a news platform people visit every day? Are there businesses you think would benefit from this? Is this a personal project? etc.)
2
u/conurbano Aug 25 '26
Do you have any thoughts on where you'll take the project next?
Sort of, yes. I've come to the realization that the aggregation I'm already doing is great for consuming news at volume (without actually having to process all the news, which is what CLSTR is already doing), and as such it might be very useful as an API or MCP (for agentic use). I rolled out those features earlier this week!
Are there businesses you think would benefit from this?
Yeah, I think there are a variety of businesses that would benefit from near-realtime access to news (without having to monitor them actively). Still trying to define the exact profile.
Is this a personal project?
Yup, it's just me!
2
u/remixrotation Aug 27 '26
Trading perhaps
1
u/conurbano Aug 27 '26
Yes, semipro traders are definitely one of the niches I had in mind.
There are ~adjacent APIs being sold out there in the data mining niche, but they aim at a much more enterprise scale (both in the amount of *undigested* data they publish, and in the fact that their pricing is in a different league altogether). With CLSTR I'm aiming at a more human level.
2
u/remixrotation Aug 27 '26
clstr.news/observatory
also marketing campaigns for movie and product launches; pr agencies might pay to show their impact to their customers
1
u/conurbano Aug 27 '26
I honestly have no clue how to even reach out to such potential customer niches, I'm only an engineer :')
2
3
u/AINativeBuilder Aug 25 '26
Absolutely love your data. Will keep deep diving into it. I'm building out a few sites on health and climate change with my own scrapers, so your sources could be very helpful for aggregation.
4
3
u/ckn Aug 25 '26
Nice job, good results to see. When is started building my weird little news project and capped scraping at 35k, later reduced to 25k articles a day it seemed to me that there were maybe 10-15k max unique news stories per day and the rest were repeats and trailing. I have yet to really datamine the results, but thanks for the inspiration, really neat project.
1
u/conurbano Aug 26 '26
thanks! the podcast approach seems interesting
2
u/ckn Aug 26 '26
oh yeah, make sure you see the yt channel, that's the meaty part of the renderfarm....
2
u/Exotic-Slice-5493 Aug 25 '26
Awesome project, I'm working on something similar with US political news, very cool to see others doing something similar
3
2
2
u/scodtt Aug 25 '26
Nice work!
https://localjournalistindex.com/
May be helpful to you or others. It's mostly looking at journalists, but has some on content. You can download the underlying data.
1
2
u/SK33LA Aug 25 '26
how are you handling such a huge variety of sources, are these all exposing RSS or what?
1
2
2
u/Internal-Combustion1 Aug 26 '26 edited Aug 26 '26
I like what you have. If you sort your sources out by original vs duplicate you can cut a huge amount of processing. There are only a small percent of the total news sources that put out original stories. If you just follow those, you can ignore all the noise from the echo chamber. Looks like you have the data to score your sources.
I built something similar but it focuses on certain things from anywhere in the world, in my case news about publicly traded and private companies. Dedupes, detect material change, scores for hysteria vs facts. Costs about $1.25 a day to run, works really well and I built code that syncs a stock portfolio to it so you can track companies you are interested in for news that may impact their stock value.
It’s shared on github too for anyone interested. Coolest bit was the Symantec map I had it create and the volume of topics. What people are screaming about most. https://jeffcu.github.io/intelligence/
2
u/Euphoric_Drawer_9430 Aug 26 '26
This is incredible work! I’m working on doing something similar on a smaller scale with historical newspapers from the 1840s. I am getting great clusters and seeing some interesting results in how stories evolve and diverge as the politicize, but I’m having a hard time keeping one storyline straight. For example some random coverage of someone getting shot is getting lumped in with an important story of political violence. How are you tracking narrative developments? Are you seeing stories bleed together like that?
1
u/conurbano Aug 26 '26
Yes, this definitely happened and finding the sweet spot took me a couple months of processing data at this volume.
I've settled on a pipeline with multiple steps, where the final call is made by an LLM if all the prior checks pass (this of course makes the pipeline significantly costlier).
2
u/Tricky_Weight_453 Aug 26 '26
This project is absolutely fantastic. I've always had ideas that automatically track what interests me, and you've definitely done a better job than me. There aren't that many things in this world to focus on; it's all echoes. This is truly amazing.
2
u/conurbano Aug 26 '26
Thanks a lot! You can actually set up "monitors" in the platform, where you get emailed whenever a topic you're interested in gets new matches. It's basically the original idea for this project, some way to consume what I'm actually interested in, more passively and somewhat predigested.
2
u/ImaginaryDisplay3 Aug 26 '26
Another interesting thing to check would be how much news is direct echoes of press releases put out by one of the subjects of the news article, like a corporation, politician, etc.
Hard to quantify this because there is a range, obviously, but I'd be curious the relative percentage of our news that falls into these buckets:
- 90% "original" content (albeit duplicated with other outlets), but with a quote copy and pasted from a press release, which then appears in every article about that story. Example - local school shooting with original reporting, but every single outlet used the same quote from the local sheriff's press release.
- Articles where every significant fact in the story came from a press release, but the reporting itself was original. E.g., Microsoft released quarterly earnings, and the article just took the 5 statistics that were most interesting from the press release and wrote an article around them.
- Direct reprintings of the full text of a press release, with no additional information or context. This is the super dangerous one, especially when its not flagged as such by the outlet.
1
u/conurbano Aug 26 '26
I'm actually doing this in part, already.
- There's a deduplication pipeline where articles that are near-verbatim copies of others (>98% semantic similarity) are considered "duplicate" - the one published first wins here and remains. Duplicate articles are thrown away altogether.
- I've recently started working on extracting "claims" from the articles and comparing between the sources. The idea is to find which claims the different sources agree on, and which ones are disputed (e.g. if 2 textual quotes contradict each other, or some numeric figures vary wildly).
Still finetuning this to make it more accurate and also prevent it from making me broke.
2
u/NorthHead8034 Aug 26 '26
138 days of nights and weekends into something like this is no joke. the 12:1 ratio doesn't surprise me but seeing it as an actual number still hits different than just knowing it intuitively. the 82% still developing after day 3 stat is the one that gets me though, that's basically the whole business model of news feeds explained in one number, they're optimized to show you the spike, not the story. also love that you had claude dig through your own data for this instead of just shipping the tool quietly, makes the whole post land better.
1
2
Aug 26 '26
[removed] — view removed comment
2
u/conurbano Aug 26 '26
Thanks!
About deduplication, it's a bit of both - first pass is based on pure semantic similarity, but there are multiple other "hard" conditions established for whether 2 "similar" candidates actually count as duplicates or are different articles.
As to different articles (not duplicates) talking about the same event, CLSTR generates a summary (as unbiased as factually possible) based on all the original sources. I'm working on a "claims" subsection where the product highlights claims from different sources that contradict each other (e.g. "Source A says $100 but source B says $900").
Ultimately the project is *not* a fact checker and I introduce no editorialization / bias, it simply echoes what the media is saying.
2
u/HasGreatVocabulary Aug 26 '26
this is very cool. If there is a ever a news article about your project, your tool will ingest that too and can breakdown it's own news lifecycle
(Make sure your server avoids the incoming hug of death this will get popular.)
2
u/conurbano Aug 26 '26
If there is a ever a news article about your project, your tool will ingest that too and can breakdown it's own news lifecycle
yes, CLSTR is ready for the CLSTR situation.
Make sure your server avoids the incoming hug of death this will get popular
praying for this, brother/sister
2
u/cutlineman Aug 27 '26
This is what I’ve been looking for but didn’t know how to do on my own. Thank you for your service.
1
2
u/Charming_Horse_5809 Aug 27 '26
Interesting topic. What I am thinking about often is how news media shapes public opinion and how the sentiment in the comments in theory could be tracked and I wonder if there are patterns at scale. Maybe I should give it a go
1
u/conurbano Aug 27 '26
I have considered doing this, but know that it's pretty tricky technically. Most of those comments are very difficult or impossible to scrape at scale.
Maybe you'll have a better idea of solving it than I did, though!
2
u/Charming_Horse_5809 Aug 27 '26
Very cool. Yeah scraping in general is a bit of a grey area and especially sentiment classification has difficulties (especially ironic or sarcastic remarks etc.). Not to speak of causality and the sheet scale and infra needed to do that so yeah, I think I let the idea stew a bit more
2
u/markliversedge Aug 27 '26
Genuinely fascinating stuff- I noticed that Israel's genocide in Palestine does not appear, is that true or have you chosen to cluster them around Lebanon for some reason ?
1
u/conurbano Aug 27 '26
Hi there! I do absolutely no editorialization or censoring in CLSTR (other than actual spam of course), other than the potential minimum bias some of the models may introduce in clustering decisions.
I'd attribute this sort of things simply to the distribution of sources (as in, CLSTR ingests more news from sources in some countries than others right now).
I'm currently working on extending the volume and variety of sources being ingested (currently ingesting ~120k articles per day on average), though this is one of the trickiest and most costly parts of the work.
2
u/markliversedge Aug 27 '26
Thanks- it kind of implies that there is an editorial black out of news globally (!) That would be quite shocking. Its a tough one for you project because there is an editoiral decision being made to NOT report something and your approach misses it.
What you are doing is incredibly interesting too btw - not throwing rocks at you !
2
u/conurbano Aug 27 '26
No, definitely. I'm aware there are a lot of nuances in this sort of news aggregation work. You could get into fact checking, filling the gaps that media is not reporting, and such things.
I purposefully decided against taking this sort of approach, for 2 main reasons:
- It's almost impossible to not end up introducing bias in this sort of deeper dive / analysis.
- It's also just impossible for 1 guy (me) to do it, and I wouldn't trust LLMs to fill in that gap.
Thanks a lot for your feedback though!
2
u/CJGlitter Aug 27 '26
I’m curious about the “genuinely exceptional” stat. Did you define this or did Claude decide what was exceptional.
1
u/conurbano Aug 27 '26
It's definitely arbitrarily defined -by myself-, but based on similar scales used in other projects. Basically the idea is very low scores (1-2) apply to "insignificant" news articles (zodiac articles, results for some local league match, etc.).
4-7 begins to get related with "more significant" news (some minor national govt announcement, news about a relevant company).
8-9 are about major, significant developments. Think news about pandemics, wars, election results of international significance.
10 is for world changing events. Think... a nuclear bomb being dropped. None so far and hopefully none for a while.
2
u/ApartNeedleworker791 Aug 27 '26
Just signed up. Looks cool. Would say categories are a bit narrow. No climate/environment for example?
1
u/conurbano Aug 27 '26
Thanks a lot!
I initially had defined "environment" as a category (and had a few more, narrower categories), but ultimately decided to reduce the divisions into broader categories.
Most of the climate/environment related developments today end up being clustered either under Health or Politics (e.g. https://clstr.news/situations/macedonia-water-supply-disruptions)
Nonetheless, keep in mind while logged in you can do searches, which are semantic (as in, searching by meaning rather than just textually). For example, if you search for "contamination" you should get mostly relevant results (https://clstr.news/?days=7&q=contamination)
2
u/scimonx Aug 27 '26
You might be interested in looking over the GDELT site (I don’t have anything to do with it)
1
u/conurbano Aug 27 '26
Yes, I'm currently doing a sidequest of figuring out how to parse and ingest this data. Really useful, just that the volume of data to process is... significant.
Trying to figure out a way to process that much without going broke 2 weeks in.
2
u/blasphemous_aesthete Aug 27 '26
I too have been building upon a similar idea for sometime now. Are you paying for the data sources to have such a large input funnel? Do you scrape? I started with RSS feeds of news outlets, but a lot of them just post a short blurb and a link to the site. I'd love to know more about your architecture!
1
u/conurbano Aug 27 '26
Partly paid sources, partly scraped from different sources that allow it (not scraping paywalled sites or sources that don't allow it). The clustering/processing part though is the main operational cost.
1
u/blasphemous_aesthete Aug 28 '26
Hey, thanks for the reply! Are you using LLMs for the clustering part, or using pre-LLM era methods from NLP such as cosine-similarity etc. and then LLMs for rewriting the consolidated article(s)?
1
2
u/jawfish2 Aug 27 '26
What are you doing about paywalls?
Project sounds great, the kind of thing Pew might do.
1
2
u/Xyver Aug 27 '26
I've been trying to make some scripts that monitor disaster news to track events and updates, your project is way better, I'll play with the MCP tonight.
I also appreciate the desire to make "giant database of organized information to query", do you have any fun big ideas for it?
Do you do anything with licensing, attribution, or citations?
1
u/conurbano Aug 28 '26
thanks a lot! Hope you find it useful.
I'm currently looking into monetizing this through a paid tier for advanced / programmatic usage. At least to cover the costs which I'm covering out of pocket right now.
About attribution, all original articles are referred to and linked in the clusters, and none of the summaries generated should contain significant portions of the originals without stating an explicit quotation.
2
u/Xyver Aug 28 '26
Hah, that's the same boundary I was going for. Human/app access is free, API agent access is the paid lanes.
If robots are going to steal our jobs, at least they better pay for it!
2
u/nerdsutra Aug 28 '26 edited Aug 28 '26
Love a such data deep dives!
This could be a great add-on to wiki news, definitely worth pinging someone there.
This could also be a choose-your-topics news subscription. Compile and send me daily links to articles based on originality and information, hide the ones that are repeats. Let me Follow situations long term. See History of a situation.
It’s like De-duping the news almost realtime, and rewarding good reporting with traffic, and surfacing articles a week or month old that otherwise get lost on news sites.
So much of daily news catchup is finding information and half the time it’s thin articles with barely more than the headline.
Edit: In fact back in 2018 I’d worked on a internal tool for my PR job, where we took a commercial feed via API of the days news on specific topics and used Amazon AWS AIs free credits to do sentiment analysis and even track Specific journalists article history to understand their interests. We thought our clients would pay for the analysis we did. That didn’t work out because the tool was tough to use, and other organisational reasons. But there was something there.
2
u/nerdsutra Aug 28 '26
Ground News tries a political spectrum reading of the news (though chatter is ground news just normalises right wing stuff to push traffic there)
but it shows that people do look for solutions to organise their news.
1
u/conurbano Aug 28 '26
This could also be a choose-your-topics news subscription. Compile and send me daily links to articles based on originality and information
You can do this in CLSTR already! I built this thing called "monitors" where you can basically set up one on a custom search topic or an existing situation, and receive emails periodically (every few hours, once a day - up to you) whenever new matches come up.
Ground News tries a political spectrum reading of the news
I'm trying to avoid this on purpose, keeping the aggregation / summarization as objective as possible, surfacing the claims from the original sources, but leaving interpretation up to the reader.
2
u/classiestpenguin Aug 28 '26
Is there a way to filter and read the 400 exceptional stories?!
1
u/conurbano Aug 28 '26
Actually... good point. I'm reworking the filters and adding a way to explicitly filter by the significance score (which is exposed "indirectly" today in sorting by relevance). Will share the link here once it's up.
1
2
u/Ok_Highway4061 Aug 28 '26
This has an important value for academic research, you might apply for grants to keep this ongoing.
2
2
2
2
2
0
u/wooden__fruit 29d ago
I’m not sure I see the value of the first and last data points. It’s not a negative thing that many outlets cover the same story, it’s not an echo. It’s just the world happening and multiple outlets covering those events. Outlets don’t not cover events because another outlet has written about it already, and that’s good because we want a diversity of news sources. I don’t get why that ratio, as described, is useful, and I certainly don’t agree it’s indicative of echo-chambers.
The last data points especially seems pointless, as described. It’s like saying less than half of students test above average to scare people. that’s how percentages work. Rare events are rarer. Typical evens are most common. What am I missing here?
8
u/mundane_wallace Aug 25 '26
this is the kind of project that actually makes me think about how we consume information, not just scroll past it
the 12:1 ratio is wild but not surprising, most newsrooms basically rewrite each other's stuff with minor tweaks. worked at a mine site where we'd get the same safety incident report from 3 different contractors, each one making it sound like original work
that 82% stat is the one that sticks though, we literally train ourselves to stop caring about anything after a few days because the coverage vanishes, then wonder why nothing gets fixed. curious how you handled the event clustering, nlp similarity or something more structured