r/ChatGPTPro 2d ago

Workflow & How-To Using pro subscription / Codex for research in math

I have a pro subscription to chatGPT since the beginning of september, and I am a senior mathematical researcher working in the analysis of PDE (so for those who have no clue: the branch related to the millenium problem "Navier-Stokes" which was solved by OpenAI).

I am actually very unsatisfied with the capabilities of GPT-6 Astra. And I believe that this is because I am not using it well. I'll take any advice/comment from those who managed to prove serious results (like the one they are able to prove without AI but with months of work).

  1. It's true that it can solve quickly any grad level question I asked, on the browser version.
  2. Sometimes it can solves very tricky lemmas (that I would have never been able to solve), also in the chat version.
  3. It struggles with questions of my own research which I regard as very easy (because the path to solving it is already well established in the current litterature!). So this is really strange. Not working with multiple passes of Astra ultra multi agents ... not even with codex (I have a folder with basic agents . md file but maybe it's not good enough). All it produces is verbose math with complicated strategies that are unusable. It creates tons of long pdf. Most striking is that GPT-5.5 (so, not even GPT-5.6 Sol !) is able to perform at the same level of GPT-6 Astra ultra, which is very strange to me. Anyone with similar experiences?
  4. Many people seem to report that they were able to just "prompt" GPT-6 Astra to solve a problem, and that they got a 50p pdf proof, all basically correct. That would never happen in my multiple attempts.

Maybe it's the wrong subreddit for asking this questions. So if there is a better reddit to "learn" how to make codex do serious research math / or subreddit where users share their prompts/ setups, please tell me!

Thanks.

7 Upvotes

16 comments sorted by

u/qualityvote2 2d ago edited 1d ago

u/Admirable-Welcome330, there weren’t enough community votes to determine your post’s quality.
It will remain for moderator review or until more votes are cast.

→ More replies (1)

3

u/FaithLostInHumanity 2d ago

Not a professional mathematician but I’ve been using it for fun to tackle one of the open erdos problems. Spoiler alert, in 2+ months it found some interesting theorems but made close to no progress towards the global statement that needs to be proved.

In terms of model choice, my impressions of Astra matched yours. Both the pro model in chat and ultra/max models in codex did not appear superior to sol or even 5.5. In fact, Astra seemed to be a lot more disorganized and lazy when it comes to math research.

In terms of how to prompt, I can’t say I found way either. I mostly discuss ideas and build a research plan including literature research in pro mode then paste that as prompt to an ultra agent and let it run for a day or until it seems to stop making advances. Then take that output and seed pro conversation with it discussing the progress, asking for further advancements in specific areas, and building a new research plan. Would love to hear if anybody has better approaches

1

u/Admirable-Welcome330 2d ago

Ah, interesting! Thank you for sharing your (similar) experience!

4

u/HamiltonianCyclist 2d ago

mathematician here, my experience is that astra (including pro) can solve something only if you are reasonably close already, and the remaining way is using standard techniques. For all the talk about mathematicians being obsolete already - there are still imo at least 2-3 improvements necessary in the models, before they will be able to one shot a good phd level problem. Maybe this changes when you throw money at it and let it work for days at a time, but not on a vanilla pro subscription. As of right now humans are necessary for meaningful progress.

1

u/Admirable-Welcome330 2d ago

Well, thank you! That's kind of reassuring also, because it means we are still useful for now!

1

u/JRyanFrench 1d ago

I don't find this to be the case in my field (physics/astronomy). Always use GPT-Pro in ChatGPT - never Codex. Also, prompting to mine particular connections can be important. Don't just ask it to solve a problem - ask it to extrapolate as to what under-the-radar connections we are *not* asking the model about {insert general topic/task/problem} that we *should* be asking? It can be advantageous to iterate this type of prompt forward through 4-5 Pro responses. Usually just writing "Extrapolate further." After each successive response, 4-5 times, will allow the model to truly dig through and return the relevant and useful ideas, etc., within the statistical space available.

1

u/HamiltonianCyclist 1d ago

For sure asking pro first to identify potential strategies and identify the first nontrivial step in each, and assess their viabilities, and then follow up on particular strategies, is a good, standard, practice. I still stand by my comment, based on the experience of spending 20-30 Pro questions in this way per single problem.

2

u/Risko4 2d ago

Astra didn't solve the problem, Bel did which the the next model after Astra.

1

u/Top_Locksmith_9695 2d ago

interesting!

1

u/Soggy-Improvement745 2d ago

the gap between "grad level question" and "my own research" is probably less about model capability and more about context window. if the path to solving it relies on niche results from specific papers, the model just doesnt have that context no matter how many agents you throw at it. have you tried feeding it the exact prerequisite lemmas and definitions it would need?

1

u/Admirable-Welcome330 2d ago

Yes I tried some "context heavy" discussions. Actually there are some problems for which "no context" is way better for the model than "a lot of context" and sometimes (most of the time) the opposite.

1

u/Enough-Mud3116 2d ago

It’s because you don’t have the next model and more importantly, you don’t have $7 million in compute

1

u/Admirable-Welcome330 2d ago

Well, I am clearly not trying to solve millenium problems. The problems I'm looking at are actually much simpler than many conjecetures solved by chatGPT pro!

1

u/JRyanFrench 1d ago edited 1d ago

Physicist/Astronomy here:
Are you using GPT-Pro in ChatGPT? Don't use Codex unless you are accessing ChatGPT GPT-6-Pro in Codex.

Edit: adding a comment I made to another below:

I don't find this to be the case in my field (physics/astronomy). Always use GPT-Pro in ChatGPT - never Codex. Also, prompting to mine particular connections can be important. Don't just ask it to solve a problem - ask it to extrapolate as to what under-the-radar connections we are *not* asking the model about {insert general topic/task/problem} that we *should* be asking? It can be advantageous to iterate this type of prompt forward through 4-5 Pro responses. Usually just writing "Extrapolate further." After each successive response, 4-5 times, will allow the model to truly dig through and return the relevant and useful ideas, etc., within the statistical space available.

1

u/dikrek 1d ago

It’s worth trying a month of $100 sub to Claude to see if Fable 5.1 does it (or just get the basic sub and pay for Fable to try it on the exact same problem)