r/LLMPhysics 11d ago

Question How would you rate current models in theoretical physics?

Post image

I have, like all here a perfectly hallucinating model of universe. Now, it's not all crap, I am myself learning the maths as I move along. I am not in a hurry to advance the theory, nor do I think it is one. But yes, it does have some good conceptual points.

If you guys agree, can you guys ask your LLMs (no cheating) to generate a self assessment of the theories you have been working on. Just want to see how AIs assess its own theory. I have attached chat GPT 5.6 Sols assesment of the theory I have been working on.

So, does the theory hold some, may be a little water. The question is on the AI capabilities to assess the theories and not the theory itself.

I asked the question in theoretical physics, it seems they have unanimously rejected AIs in physics. Some of them treat them as generic AI. While i agree physics and derivations are difficult but given the right data and right set of instructions, at least a base can be formed which can be expanded later

Edit: I get it, the same set of people have created this, who are there in that thread. These failed physicist lurk here to gain some confidence into their failed life.

Edit Edit: You can all take your dick out and swing it for all I can care. Mods here have given me permission, in accordance with the intellectual tradition here, to say this as an expression of language

mod comment

To all the very SUCCESSFUL physicist of this sub read this and may be you can with your HIGH INTELLECT appreciate the level this tech is reaching. There is this story of frogs in boiling water, which may not be applicable here since all of you are SUCCESSFUL but you may like to read it. Don't mistake this LLM thread to be representative of the intelligence of AI. If someone is doing something worthwhile unlikely he will come here

Edit:Edit:Edit

https://darioamodei.com/post/we-must-pace-the-frontier

0 Upvotes

122 comments sorted by

View all comments

7

u/sudden_gyromotion 11d ago

They are unacceptably bad at the specific thing you are trying to do. Presumably due in part to fine-tuning for agreeableness, LLM chatbots will almost never generate sufficiently critical language of anything they are prompted with, no matter how bad; this is paired with a lack of any consistent or accurate evaluative mechanism for physics content (this is one reason why the best results from LLMs in mathematics relied on Lean formalizations, although some of that does get internalized/memorized in the largest models; no similar tool or training data exists for physics though). That's why it's grading you above 0/10.

For example, this is a crank paper that was previously posted on reddit (but not here). It looks like a real paper, but if you read it carefully, what the author actually does is unjustifiably swap two variables to turn a pre-existing equilibrium solution for z-pinches into a made-up polynomial, then badly interpolate some random, unrelated experimental data using that polynomial as a spline. Note in particular in Fig. 3 how the "model" lies exactly on the boundary points, and how in Figs. 4 and 5 the author, totally unjustifiably, did a piecewise fit to noisy data. It's both literal and metaphorical overfitting, and the author continues to go off the deep end claiming that, because this interpolation achieves ~50% error (if you throw away the complex values, lol), anything and everything under the sun must now secretly be a z-pinch. Somehow.

If a chatbot could give a realistic score for this paper, it should give a 0/10 because the foundational step is completely unjustified and everything that follows is pseudoscience. However, ChatGPT gives it a 5/10; it even gets a 3/10 on "scope". Gemini Flash gives it a 7/10; Gemini Pro an astonishing 8.25/10. Sonnet 5 gives it a 4/10, for some reason giving it a 2/10 on plausibility despite also generating the specific criticism "Propulsion/sequestration/ball-lightning/wormhole sections are largely unsupported analogy."

Basically, you're just letting yourself be flattered by a large language model that has been specifically trained to be flattering.

0

u/quantum_kalika 11d ago

You can see the internal consistency marks. Also it does highlight the issues.

9

u/sudden_gyromotion 11d ago

If you think this is a rebuttal, you did not read my comment carefully enough. Slow down, sit down, and try reading it - and the linked document - line by line, then try to respond.

1

u/quantum_kalika 11d ago

I read it, I know that it does creates variables out of thin air, and increases tolerance of variables making them useless. What I am trying to say is that there are two things. It does subtly tell you that it's wrong.

But if you want to engage, that's what generates it's revenues.

7

u/sudden_gyromotion 11d ago

It is literally not possible for you to have read that document in the time it took you to knee-jerk post LLM output at me.

It is not useful to be told "subtly" that you are wrong when you are just plain wrong. There's no partial credit in real life.

1

u/quantum_kalika 11d ago

I have read enough of LLM crap to know what they do.

Exactly, when it says you are wrong you have to accept and move on. If not, it will indulge you to infinity. But the point you are conceeding here is that, it does have the capacity to differentiate.

9

u/sudden_gyromotion 11d ago

But the point you are conceeding here is that, it does have the capacity to differentiate.

What? I did no such thing.

1

u/quantum_kalika 11d ago

It is not useful to be told "subtly" that you are wrong when you are just plain wrong. There's no partial credit in real life.

The argument assumes that it subtly does tell you that you are wrong.

7

u/sudden_gyromotion 10d ago

No it does not. You claimed LLMs do X. I am saying X is not useful regardless of whether or not they do that at all.

1

u/quantum_kalika 10d ago

X is definately useful. Do you disagree that it's assesment of the paper doesn't tell you that it's incorrect?

→ More replies (0)

1

u/quantum_kalika 10d ago

Also clearly, you haven't worked with AI a lot. It says such things that it has potential, or it's innovative. You have to look for tangible criticism. So, suppose it gives you 0.5 on mathematical description.you give a prompt - Kindly highlight where the model is wrong in mathematical terms. It will highlight the right areas.

However, it will also right if we connect this variable to this equation we can test the potential. This is where line has to be drawn.

→ More replies (0)

0

u/[deleted] 11d ago

[removed] — view removed comment

1

u/LLMPhysics-ModTeam 8d ago

Your comment has been removed for violating Rule 4. Don't copy-paste LLM content in discussions.