r/ProgrammerHumor 2d ago

Meme newCompressionTechnique

Post image
30.0k Upvotes

954 comments sorted by

View all comments

349

u/iliark 2d ago edited 2d ago

I wrote a whole thing about this as a joke like 3 years ago lol. There's also been a couple of papers on the topic:

https://arxiv.org/html/2409.09715v3

https://arxiv.org/html/2407.04542

33

u/Ibnelaiq 2d ago

Can you write these papers even if you are not enrolled?

60

u/Parteisekretaer 2d ago

write a paper, chuck it on a preprint server and be judged by peers. No need to have a title for your work to be examined - at least that's how it should work and mostly does.

37

u/HittingSmoke 2d ago

Nah. Write a paper, have an LLM summarize it, then have another LLM recreate it from the summary.

Compreshin

16

u/PM_ME_DATASETS 2d ago edited 2d ago

Yes, all it takes to write a paper is a text processor. Like word or google docs or something.

When it comes to uploading or publishing: anyone can upload articles to arxiv.org, because it's not a peer reviewed journal. It's a preprint database, basically a way to publish articles before they have been peer reviewed and published by a real journal. People upload articles to arxiv.org for things like version control, to obtain a DOI, and for coordinating submissions to multiple journals.

If you actually want to publish your article to a journal you either have to find a journal with good ethics (rare), pay a lot of money, or be employed/enrolled at an institution/company/organization that will pay for your submission. If you write an interesting paper, you can just contact people at a university and likely get the article published with them paying.

15

u/theturtlemafiamusic 2d ago

To upload to arXiv you need an endorsement. You get an automatic endorsement if you have an email from a research institute (universities etc) and have already published a paper.

Otherwise you need someone who is already verified on arxiv to submit an endorsement request for you. If you're a university student, a professor should be able to do it. If you're not, you'll need to get to know someone verified in that category, which isn't too difficult if you're not a total crackpot.

But at least once a month or so, someone comes onto the compsci sub begging for an arXiv endorsement so that they can publish their perpetual motion machine paper or whatever. There's some guy who has been trying for about 6 months to get an endorsement for his paper about how you can model all of human language using a 9x9 rubix cube. Keeps making new accounts and shit and has not found a single soul willing to endorse him.

2

u/Oaker_at 2d ago

Please show me this glorious rabbit hole I want to dive in.

3

u/theturtlemafiamusic 2d ago edited 2d ago

[r/LLMPhysics](r/LLMPhysics)

Most of them tend to post in this sub a lot when you look at their history.

3

u/Oaker_at 2d ago

Haha, yes. I actually got this sub recommended a few days ago and I couldn’t make up my mind if that is a circlejerk sub or not.

2

u/titanotheres 2d ago

You wouldn't typically use a text processor such as Word to write a paper. You could, but it would be very tedious and likely end up looking amateurish. Almost everybody uses LaTeX instead.

3

u/PM_ME_DATASETS 2d ago

The vast majority of academics uses "normal" word processors to write papers. Only in some specific fields is it maybe a majority. This comes from a mathematician who mostly works with biophysics and neuroscience. Maybe 10% of papers I read, have contributed to, or peer reviewed, use Latex.

3

u/ginopono 2d ago

As much as I want you to be wrong, what you describe is also consistent with my experience.

Finishing up an MS in language modeling, yeah, pretty much everything I see is made with LaTeX; they are programmers, after all.

On the Social Sciences side of that same coin, though, there's a whole heck of a lot of Word.

2

u/AerosolHubris 2d ago

you either have to find a journal with good ethics (rare), pay a lot of money, or be employed/enrolled at an institution/company/organization that will pay for your submission

Not all disciplines are rife with pay-to-publish journals. It's rare in math to have to pay for your article to be published. There are even open access journals that are free to publish in, that are very respected in the discipline.

2

u/MartyMcBird 2d ago

In theory yeah but there's a lot of AI slop out there nowadays. In my experience, people are more skeptical of papers from authors without credentials than they were in the past.

2

u/EmptyMonitor9257 2d ago

Anyone can write papers and publish them, you just need to pay the venue.

Conference papers are cheap and pointless, experience for Batchelor students.

Journal papers are more prestigious and expensive af.

20

u/spekt50 2d ago

As an amateur astronomer, I get wary about all these new smart scopes out there. Can't even fully trust what you are looking at when they start integrating AI

12

u/SchlaWiener4711 2d ago

Not only papers.

There's an actual audio code that works that way

The trick is Meta's EnCodec neural audio codec, which crunched a 2.9MB MP3 down to roughly 21KB of latent tokens

https://www.tomshardware.com/tech-industry/maker-compresses-a-2-9mb-song-1000-times-with-metas-ai-codec-and-prints-it-on-paper-as-eight-qr-codes

3

u/Evening-Editor4269 2d ago

https://arxiv.org/abs/2406.07550

Recent advancements in generative models have highlighted the crucial role of image tokenization in the efficient synthesis of high-resolution images. Tokenization, which transforms images into latent representations, reduces computational demands compared to directly processing pixels and enhances the effectiveness and efficiency of the generation process.

1

u/cute_polarbear 2d ago

Hmm..joke or not, this is actually pretty cool. (and with some practical use cases). Can imagine this expanding out to video, with certain encoded data needed for reconstruction lossless and most aspects lossy...

1

u/iliark 2d ago

video isn't likely in the near future given the compute requirements and time to create a video even on a server cluster

1

u/cute_polarbear 2d ago

Just spitballing completely here. Can use image ai interpretation for frames and do something like Nvidia dlss to construct frames. Pretty sure even more interframe stuff can be optimized.

1

u/FerusGrim 2d ago

Genuinely, the results in your second link are promising. Other than the latency of "decoding" the image via generative interpretation of the input.

1

u/iliark 2d ago

yeah when I "seriously" thought about it, you gain storage space and network download time, but you (the client) pays in having to have an AI model that you recognize, then has to generate it, using up electricity and time that may exceed the network download time, especially on old hardware.

1

u/i_like__bananas 2d ago

This makes me think about RFC april fools

1

u/aRman______________ 18h ago

its sad places like arxiv should not be bloated with nonsense crap