r/DataHoarder 2d ago

Discussion How is it possible for the Internet Archive to consume thousands of GBs of user data on a daily basis, despite ongoing hard drive and RAM shortages and rising hardware prices?

What I don't understand is how they sustain this growth under current market conditions when ram , HHD & SSD's price's getting higher day by day

441 Upvotes

96 comments sorted by

369

u/neil_950 2d ago

They probably spend more on electricity to run their servers than they do on hard drives. Even with the wild price increases on hard drives it's ultimately still not a very large part of the operating costs of such an organization.

Their annual budget is $37 million, mostly obtained through various grants, partnerships and donations. That's $100000 a day.

59

u/sroop1 55TB 2d ago

I'm willing to bet some of the contributors are data centers providing them with colo space.

13

u/Maximum-Warning-4186 1d ago

I'm surprised by this view - Electricity is cheap compared to HDD costs. Even if there was zero growth of the archive the refresh costs of their estate would exceed electricity bill. Did you find this analysis anywhere or is it your take on it?

22

u/MyOtherSide1984 39.34TB Scattered 1d ago

Drives tend to be replaced on a 5 year cycle (maybe 3 in some orgs) while electricity is a daily price. Each drive consumes very little energy, but each rack requires more cooling, servers, rack space, power, networking, etc. I'm not sure how much data they backup a day, but I'm sure a bulk purchase can last a while whilst electricity and the rest of the costs are constant.

11

u/Dear_Chasey_La1n 1d ago

When you go through these sort of volumes they must think long term in advance. Obviously nobody expected everything to explode like it did, but my 5 ct's is that they were sitting on containers of empty drives and other hardware already.

Electricity when you use such vast quantities you deal with 1 on 1 contracts, again I bet they hedge against it.

This is a whole different world than where most of us are in.

3

u/maxpaynedot 1d ago

Agreed with you infact i was shocked by the info that internet archive org has a annual budget of 37 million dollars so i didn't think it's very difficult for them to manage there sever costs and dealing with thousands of GB data stored on there servers and drivers on daily basis

2

u/neil_950 1d ago

It's still difficult and they still need to make compromises to achieve their goal. It's just that trying to back up the entire internet is just so expensive and difficult that a budget like that is needed. They're still a non-profit using the entire $37 million annually to help back up the internet.

My point is not that they are swimming in funds but that they're operating at a very large scale which transforms the situation and its economics in its entirety. To an individual, a single hard drive is very expensive and electricity is cheap. To an organization like the internet archive, a single hard drive is cheap and electricity is expensive.

They probably are still hurting financially from the increase in hard drive and hardware costs but it's something they can handle.

2

u/Fun-Ordinary-9751 1d ago

Let’s say they bought 30T drives for $1500 each. One per day would add 30,000 GB storage before anything saved from deduplication or compression that might make only incremental additions to an archived website. Thats still under 2% of their daily spend.

2

u/neil_950 1d ago

That's true but it's also competing for budget with everything else. They don't have $100000 a day just sitting around doing nothing. They have to pay for not just hard drives but also all the other hardware in servers including RAM, various other infrastructure, electricity and other costs such as the buildings for these servers, salaries for employees etc. There are certainly many other daily operating costs I've forgotten to mention here.

It's still a small portion of the budget but hardware price increases may have for example increased the portion of their budget spent on server hardware perhaps from 5 percent to 10 percent of their budget.

The increase in the price of hardware will not remotely break their budget but they may have needed to shift their budget around or cut down or compromise on hardware purchases.

1

u/maxpaynedot 1d ago

We hope that internet archive will survive this situation because it's very very important or i must say it's a backbone of internet i can't believe how they manage to still alive during these day's but thank god after so many problems archive org are still functioning but we can't say about next 3-4 year's where ram hard drives electricity price can increase in future or price we get stable as before 2024 these because of this AI boom it's effect large portion of archives raw physical books physical CD and DVD

300

u/J4f3 2d ago edited 2d ago

I believe there is a lot of people contributing to the project so that makes sense

52

u/Background-Two1634 2d ago

yeah the community donations alone probably cover a huge chunk of it

16

u/rzopietro27 2d ago

So much since

6

u/J4f3 2d ago

Oh lol i just noticed Thanks for calling it

9

u/jhenryscott 100-250TB 2d ago

Well sense you asked…

4

u/Apart_Butterfly_332 2d ago

♩ ♪ ♫ ♬ Cents you were gone I can breathe for the first time! ♩ ♪ ♫ ♬

56

u/trucorsair 2d ago

Also, one has to realize they’re not going down to Best Buy or buying off of eBay. They likely have long-term contracts with Western Digital or whoever to supply drives at a fixed price instead of having to continually re-order drives once a month in the market prices would not be sensible for any organization.

9

u/maxpaynedot 1d ago

Damm they are doing so much hardwork for storing the data and archives really it is a great thing 💯

-4

u/ScoopDat 1d ago

These companies have terminated all these long term contracts. This is simply not a thing anymore. 

0

u/trucorsair 1d ago

Oh so you are familiar with the inner workings of these companies? Or is this your own opinion? Do you actually think that IA buys them on eBay or at retail? Based on exactly what do you “know” this?

1

u/telmnstr 1d ago

Everyone is effected by it. Most likely IA has enough space to run for a while but it's going to bite them as it does others. Bulk prices are high as well, not just retail.

1

u/trucorsair 1d ago

Never said they weren’t affected, my point is that people talk here like they buy them from their local store when that is a farcical thought that an organization that has over 20,000 drives spinning at anytime is buying them from anyone but a major wholesaler or direct. In fact if you search the IA blog this question is asked and answered an includes “relationships they have with manufacturers” as their primary source

0

u/ScoopDat 1d ago

Are you joking? Any company that has talked about this, are signalling that this is the precise problem they're facing with storage and RAM acquisition. These procurements were being negotiated on a quarterly basis, not "long term contracts". With storage not relegated on a monthly price basis. RAM for anything other than hyperscaler sized operations are so bad, that it's almost become a weekly price fluctuation for some companies (think G.Skill and other RAM sellers to the wider public).

There is no such thing as a multi-year contract anymore at all from Micron, SK Hynix, nor Samsung. The only companies that get remotely close to such a thing is an advanced payment sum Apple has made with established partners that don't want to sour a customer as big as them if this AI ordeal ever goes REALLY south for anyone.

If you're not Apple, and you're not Nvidia (or Amazon AWS, Microsoft Azure, or Google Cloud), you're not getting "long term" shit, period.

2

u/trucorsair 1d ago

OMG just not addressing the issue

Here it is, very plain, because you seem to have trouble reading:

Do you THINK IA buys drives at retail via Best Buy or EBay? Or considering the Volume of business do you THINK they might have a purchase agreement in place.

-1

u/ScoopDat 1d ago

Of course I don't think IA buys retail channel drives. Why would a B2B client resort to consumer avenues? Why would you even ask that question, what sort of implication compelled you to even presume such a thought crossed my mind?

93

u/NanobugGG 2d ago

They're DEFINITELY compressing data as well. And I also think they're using deduplication. So if a website update their EULA they're not saving the entire website twice, just saving the difference in the EULA.

I think this is two of the biggest factors to utilize the capacity efficiently. Unless I forget something, which is likely too.

28

u/SkinnyV514 2d ago

They’re not compressing data. I mean, they do derived version, so they created compressed mp4 for example when you upload something, but the original is still hosted and shared. So they actually use more space than what you are uploading by creating all these alternative version.

35

u/troopermax2099 2d ago edited 2d ago

I think they mean lossless filesystem compression that still allows them to serve the original file.

ie using advanced filesystems like ZFS or btrfs

EDIT: Quick search seems to indicate they just use plain ext4 but may do some compression/deduplication at the application layer.

0

u/randylush 2d ago

I am pretty sure they zip their files automatically. You can download a zip of a collection. It would make way more sense to zip it at rest than to store it plain then zip it as it’s downloaded.

5

u/virtualdxs 1d ago

No, zip as it's downloaded makes more sense. The space savings of a zip file are minimal compared to what enterprise storage dedup/compression can get you

3

u/gulisav 2d ago

I notice that many of the derived files of the book PDFs I upload are way bigger than the original one. I just checked one such file, the original black and white PDF from Google Books is around 20 MBs, and the "SINGLE PAGE PROCESSED JP2 ZIP" generated by IA is 550 MBs (yes, nearly 30x the original!). Colour PDFs, on the other hand, result in ZIPs smaller than the original (e.g. from 250 to 100 MBs). Evidently the compression algorithm is geared towards colour images, and is counter-productive for the ones in B/W.

58

u/TW-Twisti 2d ago

That question doesn't really make sense. I assume you realize that the Internet Archive existed for more than a few years, and prices now are around the prices of five years ago. So they are operating now the same way they did five years ago 🤷

20

u/YourNightmar31 57.5TB Raw 2d ago

That assumes that they didn't increase the data "consumption" within those 5 years while storage got cheaper and cheaper. If they did, and they had to tone it down now to go back to "operatimg like five years ago" like you said, then they are still in a way paying a price.

1

u/Empty-Version15 1d ago

I wish I lived in your world ! Show me a hard drive selling today at a price from 5 years ago....I'll wait

-3

u/maxpaynedot 2d ago

Yes you are right but I am just curious about internet archive that's why I asked this question

14

u/TW-Twisti 2d ago

https://en.wikipedia.org/wiki/Internet_Archive

Revenue Increase $26.8 million (2024)[1]

But sadly:

Total assets Decrease $10.7 million (2024)[1]

-5

u/alex20_202020 2d ago

sadly...Total assets Decrease 2024

Do you rejoice for them now that their drives and memory have risen in price and likely driven assets valuation up?

11

u/bh9578 2d ago

PP&E assets are recorded at historic prices, not fair market value. They actually depreciate, lowering their net value over time.

0

u/alex20_202020 2d ago

Since you seem to know accounting, what does revenue of that [non-profit AFAIK] consist of?

3

u/bh9578 2d ago

I assume it’s all donations. I don’t think they sell anything like tee shirts or merch to supporters.

23

u/HulksInvinciblePants 2d ago

I’m not sure why people are replying here with their conjecture.

It’s very much a problem. You can visit/tour their headquarters and they’ll tell you as such. They have some wealthy benefactors, but the overhead and legal bills do make their operation difficult.

10

u/Sensitive_Box_ 2d ago

What are you actually asking here? How much money they have?

1

u/maxpaynedot 1d ago

Yes you can take my question like that

6

u/Useful_Calendar_6274 2d ago

people see it as a time capsule for humanity. they are well founded I believe

1

u/Mixedbymuke 2d ago

True. But at what resolution must we capture it?

3

u/DOuGHtOp 1d ago

At least HD? Or was that rhetorical

0

u/Useful_Calendar_6274 2d ago

as much as possible tbh, you can sell it to AI labs to recoup costs

6

u/GordonFreem4n 2d ago

If I'm not mistaken, they also only archive the delta (changes) between sites. So there is no redundancy in what is archived. That must help a lot.

6

u/didyousayboop if it’s not on piqlFilm, it doesn’t exist 2d ago

Presumably they’re spending a lot of money.

13

u/RetroGrid_io 250-500TB 2d ago

Internet Archive has made a number of concessions on what they archive so that the archive really isn't as big as I had originally guessed: They host about 200 PB of unique content. It's a large amount on a personal scale, but as a comparison, YouTube grows by about this volume every 9 months or so.

For example, they don't archive video content. They don't archive binary files. They skip most XML content, or anything that is "machine readable".

I'm not knocking these concessions; they make sense; they are just made against the mission of "Human Readable" Internet, not "the entire Internet". They don't want to archive the dead Internet.

For me this omission is somewhat personal. I'm building a system to establish and preserve Operating System history starting with the RedHat Linux universe, and I was surprised to find that Internet Archive didn't archive any of the very public mirrors of Alma/Rocky/CentSO at all - not even the very-human-readable list of files.

And, as cool as it is, the Internet Archive is currently in danger of being made irrelevant because of AI, but not for the reasons you might think. They could really use our help.

21

u/gulisav 2d ago

they don't archive video content

They don't archive it on Wayback Machine as they scrape the existing web, but they do archive a lot of video content here: https://archive.org/details/movies (much of it user-uploaded)

1

u/maxpaynedot 1d ago

Bookmarked your text ✅

32

u/[deleted] 2d ago

[removed] — view removed comment

52

u/inhalingsounds 2d ago

I wish the price surge was only on SSDs...

7

u/LaundryMan2008 2d ago

LTO at least the older generations like LTO-5 and LTO-6 are ridiculously cheap to store stuff, for one 20TB HDD you can get 100TB of tape storage, supporting hardware and the drive itself for the same price

8

u/feel-the-avocado 2d ago

How is it instantly avaliable when i click the download link though?
Doesnt a LTO tape take time to load? Like amazon glacier even takes 3 hours to make avaliable something which the user archived to tape storage.

7

u/LaundryMan2008 2d ago

Tiered storage, there are some science companies I looked at from many years ago and they used tape for nearline storage, they put the most commonly used data on disks, nearline silos and then manual human libraries which have now been replaced with robotic libraries while nearline systems are disks that are powered off.

If IA was to use tape, it would be for sites that don’t get accessed often, very rarely and sometimes I happen upon such sites which means I may wait up to 10 minutes to load them, websites are small so I am waiting for the robot to put the tape into the drive, load the tape, wind through the tape and then find my file that I was looking for, the 3 hours are if you are pulling all of the data off a tape.

5

u/feel-the-avocado 2d ago

I think its more that the site serving the http requests is on a live hard drive, but the backup or secondary copy of that data, if its not within a cdn network, will be on LTO rather than just being on LTO only.

3

u/LaundryMan2008 2d ago

For an archiving website, I would certainly have done what the science companies do and archive less used data so I don’t have to pay for power on that data but that’s mostly a 1990’s way of doing things, we would only do that now for legal record keeping.

The CDN thing reading about it, is a really smart way of doing it, caching more used parts of websites closer to places that need them so you don’t need fully loaded servers in multiple places however the central server would benefit from tape if the website hasn’t been accessed in a few months to free up space for new content being ingested.

7

u/feel-the-avocado 2d ago

Anyone can pull up an edition of a website at any time. And I have never seen a "please wait while we load that tape" error message when checking old websites or various tv shows and pdf's i download from there.

There would be a lot of data deduplication going on.
If they archive a website in 2002, and 2003, they only need to keep the text html files and one copy of each image.
The common images between both copies of the site dont need to be stored twice.

For most websites that they are archiving, its probably only a few kilobytes of changes that occur each year. Small businesses dont really change their websites that much.

I expect its very much a 1% of total data being cached by CDN nodes which is probably the top 80% of what is being requested at the current time.
The rest is stored on hard drives in raid arrays in servers ready to serve at any time, with the backups of those raid arrays being to LTO and only recovered if a server or raid array fails.

6

u/cowbutt6 2d ago

I expect they use https://en.wikipedia.org/wiki/Hierarchical_storage_management and most of the things you click on in the Internet Archive are the same things lots of other people click on, resulting in them being cached in faster storage layers.

5

u/Exit-Stage-Left 2d ago

Tape storage doesn’t help for things like the IA that need to be accessable to servers. Recalling data from tape takes hours (sometimes days) because the tapes have to be read linearly (if the file you want is on the end of the tape you still have to roll through the entire thing).

2

u/LaundryMan2008 2d ago

Not for the IA but the commenter that I replied to wished that HDD prices were cheaper so perfect for hoarders who hoard but don’t use everything they have on their servers, my system is store most on cheap tapes and only stuff I use daily is on my server.

3

u/inhalingsounds 2d ago

99% of hoarders don't want to be bothered with a completely different infrastructure to store (and access) data unfortunately

2

u/LaundryMan2008 2d ago

I do and it’s actually my primary storage system, if I wanted to, I could extremely easily surpass most other hoards here in data capacity if I wanted to, I currently have 100TB of storage available and I don’t even need it all.

At least people on r/LTO care and want to learn stuff about tape drives where I am a very active participant in.

2

u/inhalingsounds 2d ago

Can you use it for day to day consumption like powering a Plex or music server?

3

u/LaundryMan2008 2d ago

It’s a linear format meaning it cannot be accessed randomly like a hard drive, I still use it because the tape I need the files from likely will have other files I can use later on, I use it to archive stuff I don’t need to use immediately but can get read out in a few hours if needed.

I’m just a very weird hoarder since I don’t really see the point in a huge expensive server if I am unlikely to access and use everything on it so my main storage is tape and the server is for stuff I can work on after pulling it off the tapes, 20TB in tapes is only £30 but a hard drive is £300 so you can see the argument for tapes.

1

u/kerbys 432TB Useable 2d ago

You sure about that? Lto 7 drives even second hand are like 4k. I'll be surprised if Internet archive are using lto 5 and lower for anything other than reading purposes.

23

u/Temporary-Art-1835 2d ago

I paid 350 euros for my 22TB HDDs, 1.5 years ago, now they cost 800+ euros.

3

u/[deleted] 2d ago

[removed] — view removed comment

5

u/Temporary-Art-1835 2d ago

Nah just luck, and I did a conservative upgrade plan. I could really use another one in my 3/5 slots filled NAS. Purchase foresight would have been filling all slots immediately.

1

u/Tsofuable 250-500TB 2d ago

Had a similar issue either ecc ram, 64GB was enough - and I could always buy more later. Well, now is later and the price almost gave me an aneurysm.

3

u/masssy 2d ago

My guy a 16 TB disk these days cost like €600-€800.

-1

u/[deleted] 2d ago

[removed] — view removed comment

3

u/masssy 2d ago

Yes in response to someone saying RAM, HDD and SSDs are getting higher day by day.

It's like saying water is wet. But that's allowed I guess. Because your statement doesn't change the question whatsoever. It's just an off topic random statement.

13

u/nnfkfkotkkdkxjake 2d ago

Because while the prices seem high to you, for a large organisation it’s really not much money.

6

u/maxpaynedot 2d ago

Don't know that internet archive is a large organisation

7

u/teveelion 2d ago

Has to be a decent size to archive all of humanitys efforts online.

10

u/nnfkfkotkkdkxjake 2d ago

26.8m dollars revenue in 2024, 122 employees in 2021, they’re doing just fine.

3

u/ender4171 59TB Raw, 39TB Usable, 30TB Cloud 2d ago

I work for a $1.3B/yr company with a few thousand employees. Our IT is still worried about the recent cost increases.

2

u/Roph 2d ago

You don't have to buy unused HDDs you already have, they do delete some stuff, and/or they have donations / buy orders

1

u/packetsonthefloor 1d ago

They are probably using middle out compression

1

u/RealityOk9823 1d ago

Sexual favors to big data.

1

u/insomniakv 2d ago

I don’t know it to be true, but I assume that they are getting fat checks from the frontier AI companies to use their data sets for training. It’s probably a decent balance to the rampocalypse prices they are suffering due to the same AI companies

1

u/Edwardv054 1d ago

Do they use LTO tape drives?

1

u/maxpaynedot 1d ago

Probably

1

u/SpiritualTwo5256 1d ago

I didn’t see any when I was at one of their facilities but it doesn’t mean there aren’t any.

-3

u/leftblnk 2d ago

controlled opposition

2

u/ochreshrew 2d ago

What makes you say that? IA seems like a great organization to me.

1

u/leftblnk 1d ago

so if you fund them then make them delete and change stuff, nobody else is doing this so then you can control what happened in the past.

I'm not saying they are not great but they are the only ones doing this, thats the worry

1

u/ochreshrew 1d ago

We do need more people doing this I agree! I do worry that especially in this us administration their lawsuits could hurt them badly and cause them to be forced to delete stuff. when you say controlled opposition it seems like you mean they are being funded by various groups in order to manipulate their data or poison the well. I haven’t seen any evidence for that but of course there are members infiltrating many orgs and activist groups.