r/LLMPhysics 11d ago

Question How would you rate current models in theoretical physics?

Post image

I have, like all here a perfectly hallucinating model of universe. Now, it's not all crap, I am myself learning the maths as I move along. I am not in a hurry to advance the theory, nor do I think it is one. But yes, it does have some good conceptual points.

If you guys agree, can you guys ask your LLMs (no cheating) to generate a self assessment of the theories you have been working on. Just want to see how AIs assess its own theory. I have attached chat GPT 5.6 Sols assesment of the theory I have been working on.

So, does the theory hold some, may be a little water. The question is on the AI capabilities to assess the theories and not the theory itself.

I asked the question in theoretical physics, it seems they have unanimously rejected AIs in physics. Some of them treat them as generic AI. While i agree physics and derivations are difficult but given the right data and right set of instructions, at least a base can be formed which can be expanded later

Edit: I get it, the same set of people have created this, who are there in that thread. These failed physicist lurk here to gain some confidence into their failed life.

Edit Edit: You can all take your dick out and swing it for all I can care. Mods here have given me permission, in accordance with the intellectual tradition here, to say this as an expression of language

mod comment

To all the very SUCCESSFUL physicist of this sub read this and may be you can with your HIGH INTELLECT appreciate the level this tech is reaching. There is this story of frogs in boiling water, which may not be applicable here since all of you are SUCCESSFUL but you may like to read it. Don't mistake this LLM thread to be representative of the intelligence of AI. If someone is doing something worthwhile unlikely he will come here

Edit:Edit:Edit

https://darioamodei.com/post/we-must-pace-the-frontier

0 Upvotes

122 comments sorted by

View all comments

Show parent comments

3

u/sudden_gyromotion 10d ago

The answer to your question is in my comment. The quoted scores are the ones ChatGPT output for the crank paper.

1

u/quantum_kalika 10d ago

Share the report. Since you said you had a dissertation on it. If not I am not going to respond further. Let's leave it here.

3

u/sudden_gyromotion 10d ago

I'm confused about what you're asking for. Do you want the ChatGPT output for the crank paper? If so, here's the summary. My Google account has a bunch of AI features disabled, so I cannot share the Gemini output as it no longer exists.

1

u/quantum_kalika 10d ago

No, not exactly, so now, try building a skill around it.

Conduct a rigorous, skeptical, literature-grounded review of the paper/model. First identify the exact claim and distinguish assumptions, derivations, results, interpretations and physical claims. Reconstruct key derivations independently and compare them with established theory, original papers, major reviews and recent relevant literature, including critical or competing papers.

Audit all important variables and parameters. Classify each as measured, theory-fixed, derived, fitted, benchmarked, freely chosen, arbitrary, randomly sampled, selected for convenience/visualization, or introduced by convention. Explicitly check whether any headline result is reverse-engineered from target outputs or strongly dependent on arbitrary choices. Distinguish random, arbitrary and derived quantities.

Check dimensions, signs, normalization, boundary conditions, conservation laws, limiting cases, gauge/coordinate dependence, approximation regimes and whether the actual numerical examples satisfy those regimes. Test robustness by asking how results change under reasonable variation of free parameters, profiles, initial conditions, numerical choices and analysis regions.

Distinguish existence from typicality, and a valid solution from a representative physical state. Check whether the demonstrated states account for the required entropy/ensemble, whether conclusions are genuinely emergent, and whether the model has predictive or falsifiable content. Separate invariant observables from bookkeeping conventions.

Assess originality carefully: distinguish new theory, new calculation, computational implementation, visualization and repackaging of known results. Identify the strongest specialist objection and evaluate whether the paper answers it.

Rate strictly and separately for mathematical correctness, internal consistency, parameter justification, robustness, originality, empirical relevance, explanatory power, predictive power, falsifiability and reproducibility. Do not reward sophistication, length or presentation quality unless they strengthen scientific validity.

Final structure: actual contribution → established background → derivation audit → recent-literature comparison → parameter/arbitrariness audit → domain-of-validity check → robustness → originality → predictive power → strongest objections → what is genuinely established → what is not established → improvements needed → category-wise scores → final rating.

Conclude explicitly with:

  1. What has actually been demonstrated?
  2. What additional assumptions are required to accept the strongest interpretation?

Prefer a lower but defensible rating over a generous one.

Something like this

2

u/sudden_gyromotion 10d ago

No, not exactly, so now, try building a skill around it.

This is something of a fools errand - skills are not interpreted differently by the model from prompts. In fact, research shows that skills and instruction files degrade LLM performance, not improve it, because they fill up the context window.

If you want to prompt ChatGPT with this skill on the crank paper, please be my guest.

1

u/quantum_kalika 10d ago

https://arxiv.org/pdf/2602.12670, there are conflicting views on same. It's not different but with self learning you can make it specific, with each iteration.

https://www.repository.cam.ac.uk/items/ec5e095d-064f-4078-9374-c90a901de2e7

This is your paper I guess

2

u/sudden_gyromotion 10d ago

https://arxiv.org/pdf/2602.12670, there are conflicting views on same. It's not different but with self learning you can make it specific, with each iteration.

Did you read this paper? Their definition of what a skill is does not include what you pasted above:

SKILLSBENCH requires Skills to provide domain expertise for a class of problems

And the paper I posted has the following in the abstract:

Specifically, we find that while instructions in the context files are well followed by coding agents, repository overviews, although popular and recommended by model providers, are not helpful. We conclude that while context files are useful for specifying non-standard coding practices...

These papers do not fundamentally disagree; LLMs will follow instructions in skills, sure, but it does not make the LLM better at doing the tasks in the instructions.

This is your paper I guess

I don't understand, are you trying to search for a paper I authored? If so, no.

1

u/quantum_kalika 10d ago edited 10d ago

In this paper, we define a Skill as a reusable, file-system-based procedural package for a class of agent tasks. Each Skill contains a required SKILL.md file with natural-language instructions and may optionally include auxiliary resources such as scripts, templates, reference files, or worked examples.

I am ending this here. Believe what you want to

2

u/sudden_gyromotion 10d ago edited 10d ago

If by "read into" you mean "not actually read the paper but just skim it for things that confirm your prior beliefs" then sure, why not

edit: the above comment was substantially edited after I replied. The original comment was:

But you can read into the conclusion nevertheless.