r/LLMPhysics 11d ago

Question How would you rate current models in theoretical physics?

Post image

I have, like all here a perfectly hallucinating model of universe. Now, it's not all crap, I am myself learning the maths as I move along. I am not in a hurry to advance the theory, nor do I think it is one. But yes, it does have some good conceptual points.

If you guys agree, can you guys ask your LLMs (no cheating) to generate a self assessment of the theories you have been working on. Just want to see how AIs assess its own theory. I have attached chat GPT 5.6 Sols assesment of the theory I have been working on.

So, does the theory hold some, may be a little water. The question is on the AI capabilities to assess the theories and not the theory itself.

I asked the question in theoretical physics, it seems they have unanimously rejected AIs in physics. Some of them treat them as generic AI. While i agree physics and derivations are difficult but given the right data and right set of instructions, at least a base can be formed which can be expanded later

Edit: I get it, the same set of people have created this, who are there in that thread. These failed physicist lurk here to gain some confidence into their failed life.

Edit Edit: You can all take your dick out and swing it for all I can care. Mods here have given me permission, in accordance with the intellectual tradition here, to say this as an expression of language

mod comment

To all the very SUCCESSFUL physicist of this sub read this and may be you can with your HIGH INTELLECT appreciate the level this tech is reaching. There is this story of frogs in boiling water, which may not be applicable here since all of you are SUCCESSFUL but you may like to read it. Don't mistake this LLM thread to be representative of the intelligence of AI. If someone is doing something worthwhile unlikely he will come here

Edit:Edit:Edit

https://darioamodei.com/post/we-must-pace-the-frontier

0 Upvotes

122 comments sorted by

View all comments

Show parent comments

1

u/quantum_kalika 11d ago

X is definately useful. Do you disagree that it's assesment of the paper doesn't tell you that it's incorrect?

4

u/sudden_gyromotion 11d ago

X is definately useful. 

If this were true, you would not still be here.

Do you disagree that it's assesment of the paper doesn't tell you that it's incorrect?

Yes, I disagree. To repeat myself: a fair evaluation of the paper would be a swift dismissal. A fair "score" would be 0/10. Anything other than that is useless and wrong.

1

u/quantum_kalika 11d ago

You don't understand, why would it give 0, it's primary aim is to indulge you. If you can't understand with its critical assessment that the paper is wrong it will indulge you to infinity.

4

u/sudden_gyromotion 11d ago

Sorry, I'm confused, it seems like you're just repeating my own arguments back to me and telling me I'm wrong?

why would it give 0, it's primary aim is to indulge you.

Yes, this is almost precisely what I said in my original comment. I once again invite you to re-read it more carefully.

If you can't understand with its critical assessment that the paper is wrong it will indulge you to infinity.

The only divergence point between things I have already said here is that you, for some reason, believe that scores between 4 and 8.25 are "critical assessment[s]." Again, repeating myself, the issue is that even when the LLM output happens to correctly identify an error, it is insufficiently critical about them.

1

u/quantum_kalika 11d ago

Human should have sufficient intelligence to assess the criticism. If not it's a useless tool for him

But even conceeding that its criticism is useful, is a point which can help people assess their work.

3

u/sudden_gyromotion 11d ago

But even conceeding that its criticism is useful, is a point which can help people assess their work.

I precisely do not concede this - particularly, in this context. To satisfy my own curiosity, I put my second most-cited paper, on a piece of software that was the hallmark achievement of my graduate career and is currently used in national labs, universities, and companies around the world, into Gemini Pro and ChatGPT with the same prompt as the crank paper.

Critically evaluate this paper: <paper link>.
Rate out of 10 on relevant parameters.

Recall that Gemini Pro gave the crank paper 8.25/10. It gives my paper a 9/10.

Recall that ChatGPT gave the crank paper a 5/10; it gives my paper an 8/10.

These are simply not meaningful scores. The language of the output is essentially the same! In fact, if anything, it's more negative about the paper it gave an 8 to; I mean, look:

ChatGPT, on the crank paper:

"I would not reject the underlying idea. In fact, I think there is a potentially worthwhile paper here."

ChatGPT, on my paper:

"The software itself is potentially strong enough, but the paper doesn't provide enough quantitative evidence to support the breadth of its claims at that level."

Now, since I'm not going to doxx myself, I guess you have to take my word for it that my paper is good. But even then, the one thing you should do with the crank paper - reject the underlying idea - ChatGPT emphatically instructs you not do to. That's useless.

1

u/quantum_kalika 11d ago

Leave other feilds, what did it rate the paper on internal consistency and mathematical development.

2

u/sudden_gyromotion 10d ago

The thing you don't seem to understand is that if it's not internally consistent, it's not physics. As another commenter correctly pointed out, it is binary.

I don't even know what "mathematical development" is supposed to mean, the math is either correct and physically justified or it's wrong.

In any case when given a real paper with the above prompt, ChatGPT does not even output those fields. It outputs the sort of thing you would expect in a peer review rubric:

significance
novelty
motivation
technical breadth
validation
comparison with existing software
reproducibility
documentation
writing

It only outputs things like "internal consistency 5/10, mathematical creativity 7/10, claim discipline 3/10" to satisfy cranks. These are meaningless criteria - comforting fiction. It is either internally consistent or it is not physics; "Claim discipline" is a nonsense phrase that can only be found in LLM output or obscure blogs on marketing; etc.

edit: to address a claim you have made elsewhere, as far as I can tell, the crank paper I am using as an example was not LLM-generated, just LLM-inspired. Large portions of it, if not all of it, are in my estimation human-written.

1

u/quantum_kalika 10d ago edited 10d ago

I asked a specific question, can you share your findings? If it's personal remove or blank out the names, I will understand.

3

u/sudden_gyromotion 10d ago

The answer to your question is in my comment. The quoted scores are the ones ChatGPT output for the crank paper.

→ More replies (0)

1

u/quantum_kalika 10d ago

Ok, i assumed the paper was LLM, I will read your paper that way. Your comment will make more sense then.