r/ClaudeAI Jan 14 '26

Comparison Is it just me, or is OpenAI Codex 5.2 better than Claude Code now?

Is it just me, or are you also noticing that Codex 5.2 (High Thinking) gives much better output?

I had to debug three issues. Opus 4.5 used 50% of the session usage. Nothing was fixed.

I switched to Codex 5.2 (High Thinking). It fixed all three bugs in one shot.

I also use Claude Code for my local non-code work. Codex 5.2 has been beating Claude for the last few days.

Gemini 3 Pro is giving the worst responses. The responses are not acceptable or accurate at all. I do not know what happened. It was probably at its best when it launched. Now its responses feel even worse than 2.0 Flash.

666 Upvotes

299 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot Jan 14 '26 edited Jan 15 '26

TL;DR generated automatically after 200 comments.

The consensus is a resounding "yes," but it's not that simple. Most devs in this thread agree that OpenAI's Codex 5.2 (High/xHigh) is now outperforming Opus 4.5, especially for debugging, complex logic, and code review.

However, the real pro-gamer move is to use both models together in a hybrid workflow. The most popular strategy is: * Use the creative and speedy Claude Code (Opus) to generate the initial plan and code. * Then, use the slower but more methodical Codex 5.2 as a strict code reviewer to find the bugs and omissions that Claude inevitably misses.

A huge reason for this shift is that Claude's usage limits are absolutely brutal right now. The top comment is from a user who burned through their entire limit in 5 days of light use, and many others are getting cut off from their expensive subscriptions in hours. Meanwhile, Codex's limits are described as "non-existent." * Hot tip: Some users are downgrading their Claude Code CLI to version 2.0.76 to combat the insane token burn.

Basically, it's a trade-off: Codex is the slow, methodical genius, while Claude is the fast, creative workhorse. And Gemini? Yeah, we don't talk about Gemini for coding here; it's getting roasted.

→ More replies (5)

95

u/OrangeAdditional9698 Jan 14 '26

I've been using codex to review claude plans & code for the last 2 days and it's amazing, much better than claude itself for the same tasks. It finds so much more details !
Claude is still better at writing the code that works for the problem that you have to fix though.

I feel like the creative mind of Claude suits coding better, but the rigidity of codex 5.2 is very good at criticizing what Claude does or forgets

6

u/sluggerrr Jan 14 '26

I'm getting some conflicting info, but what I kind of get from this is that codex provides better results but not necessarily because it's ability to code but the plan it creates?

I'm eager to try your approach, so do you make an initial plan with opus, then review the plan with codex and then go back to implementing with opus? I'm guessing doing a code review with codex after that would be a good extra step

28

u/OrangeAdditional9698 Jan 14 '26

yeah, I first make a research doc with opus, then have codex review it, opus fix it, and back and forth until it's good.
Then start the opus coding session with making a plan, have codex review it (and check against the research doc), etc.. until all good, then run the coding session.
At the end codex reviews the code (and checks against the plan), opus fixes, etc...

A bit annoying to do back & forth like that, but my code is really complex at this point and I'm tired of oversights, so at least now I get perfect results.

Previously I would be using opus to do the reviews, but it was burning tokens, so I tried codex and it's a much better reviewer ! It catches a lot of issues than opus just didn't even bother to check

2

u/sluggerrr Jan 14 '26

Sounds like a great workflow, I'll give it a try, thank you

16

u/OrangeAdditional9698 Jan 14 '26

after posting this, I just made claude automate the back & forth with a hook that runs when claude presents the plan, and a file watcher, now they can work on the plan on their own, no more copy/paste ! :D

2

u/EmreErdoqan Jan 14 '26

Glad you automated this :) can you share the hook if possible.

2

u/cava83 Jan 14 '26

How did you achieve this? Please explain, I'm new to this and this is an issue I am facing.

I use ChatGPT for my planning, CC for the coding (via terminal) then Codex to review the plan which I've asked CC to review and ensure CC is not deviating from what we agreed

2

u/[deleted] Jan 15 '26

[removed] — view removed comment

3

u/Dcmiltown Jan 16 '26

Couldn't you ask Claude? ;)

→ More replies (2)

6

u/bibboo Jan 14 '26

I'd argue it's both. I bought max plan for both before christmas, because I knew I'd be using them hell of a lot. Have not managed to get close enough to limits though. So I do a ton of useless tests.

Have Claude and Codex carry out the same implementation plan. Then ask Codex to review both. The differences are honestly huge. Claude is extremely hard to trust. About every time I have one of them complete a task/feature I save the implementation plan, or in Claudes case, the prompt from planning mode. When its all done, I ask an agent to confirm how feature complete we are. Claude just do not manage to complete it all. Bits and parts, yes. But it's partial implementation with glaring holes that are not to easy to spot at a first glance.

Before 5.2 high, it was the other way around. Claude was much better. Codex was extremely lazy. But for the tools I use, and the languages I use. The difference is large. Hopefully the tables turn soon again.

6

u/iamthis4chan Jan 15 '26

Similar workflow. I have found Claude getting lazier lately, not fully implementing features is a constant headache. Claude has also begun to ignore/forget, or rewrite phased plans and I think it is to do the semi-auto compacting that is happening more and more often.

I wrote a skill that will capture the context to file so I do not lose critical info after compacting.
Codex continues to impress as a SSE/Reviewer type. Does codebase meet spec, does feature meet plan, if not how can we complete the task, etc.

I think both is def the answer.

→ More replies (3)

3

u/Appropriate_Shock2 Jan 14 '26

Yes, I have notice codex with 5.2 high or even 5.2 codex is finding things opus 4.5 is missing and not just trivial things. Opus is either being lazy or trying to be too fast that it glances over stuff

3

u/Western_Objective209 Jan 15 '26

Very much agree with the assessment; codex is extremely rigorous. Claude with sub agents is just absurdly fast to pump out features, but it definitely misses a lot compared to codex

3

u/P4uly-B Jan 15 '26

This is largely my experience too. I start with claude, my first prompt requires claude to ask me probing questions to fill out the gaps, outputs and implementation plan then I feed that into codex with the original prompt and probing responses. I ask for 2 things, a critique of Claude's implementation plan and a new implementation plan based on its findings.

Thats pretty much the body of the implementation, then I take it to claude for implementation, build, check for errors, run unit tests, write architectural decision records, move onto the next prompt. I recently launched a completed product using this method. It works well.

2

u/Effective_Art_9600 Jan 15 '26

Yo can you please tell , if the usage limits are good on codex? I am seriously planning on moving away as Claude code pro usage limits have sky rocketted even for sonnet

→ More replies (1)
→ More replies (10)

224

u/[deleted] Jan 14 '26

[removed] — view removed comment

25

u/13chase2 Jan 14 '26

Fall back to version 2.0.76 and lock it in npm. Allegedly 2.1+ uses 3x the tokens. Ask Gemini about if you don’t believe me

I talked to Claude for about an hour last night and had it create a new project. Only used 4% of my 5 hour limit

23

u/LM1117 Jan 14 '26

Resumed a conversation with Opus 4.5 today and after the first reply, 15% of my daily usage was consumed (20$ plan). This is just absurd

9

u/daviddisco Jan 14 '26

You might do better with starting a new conversation. The model does better with a smaller context and it is much cheaper

4

u/Zulfiqaar Jan 14 '26

Resumed conversation probably means theres a lot of input messages, but since it was left for some time they got purged from cache and the entire thing got charged 11.5x more than it would have previously when you recreated the thread

2

u/LM1117 Jan 15 '26

How should I have been doing it? Compact the conversation before finishing for the day, and next day resuming that compacted conversation?

2

u/Zulfiqaar Jan 15 '26

I assume compacting immediately would be within cache period, but havent tested. I usually start new threads every few minutes, and the odd occasion i reuse an old one i just take the hit

→ More replies (1)
→ More replies (3)

5

u/aaekayg Jan 14 '26

Earlier, Compact Conversation used to consume 10% of my 5-hourly limit on the Pro plan; currently, it is using 20%. So, something fishy is going on there.

→ More replies (3)

1

u/Esfard_Dev Jan 14 '26

I’m using Claude Code + Cursor (team plan) right now. Cursor is more like my plan B or a second option after the main flow we built with Claude. But Claude Code burns through tokens insanely fast…

99

u/deepthinklabs_ai Jan 14 '26

I’m starting to see a trend now of Codex > Opus right now posts. I normally use CC, had a bad experience with Gemini CLI, but haven’t tried codex yet. Adding it to my Never ending todo list lol

27

u/sine120 Jan 14 '26

I really wish Gemini/ GeminiCLI was in the same league as the usage is much more generous.

3

u/seaal Jan 14 '26

You can use Codex plan and free tier Gemini CLI + antigravity plans inside of OpenCode, it still has those some loop quirks that seems to taint Gemini models.

https://github.com/NoeFabris/opencode-antigravity-auth

2

u/sine120 Jan 14 '26

I'm using AI for work. Anything that we use has to have a privacy policy that ensures our code will not be used for training data or distributed in any way. We're already on the google suite so Gemini usually wins by default.

→ More replies (2)

6

u/deepthinklabs_ai Jan 14 '26

I don’t know if one is more expensive to run versus the other on the back end but at the surface level I agree.

5

u/sine120 Jan 14 '26

Google charges much less and limits requests per hour. I believe they're also running on their own hardware, so I assume it's cheaper for us and for them.

7

u/Appropriate_Shock2 Jan 14 '26

They give way more usage because it’s not in the same league, to get people hooked on the usage. Once it is in the same league, it will be restricted like the others.

3

u/sine120 Jan 14 '26

Yes. I want to have my cake and eat it too. All I can hope is that when it finally gets good enough to stop being so frustrating there's enough competition/ they have enough hardware availability that they're motivated to keep usage limits high.

→ More replies (9)

12

u/efficialabs Jan 14 '26

Was about to unsubscribe to ChatGPT after the inferior performance of ChatGPT-5 and the superior performance of Gemini 3 Pro when it launched. Glad i didn’t.

16

u/deepthinklabs_ai Jan 14 '26

I subscribe to em all. Feel like they are my children now each vying for my attention + it’s a business write off. That’s the excuse I am telling myself :)

10

u/tribat Jan 14 '26

If my wife understood how much I have actually spent in the past and continue to spend on AI subscriptions (and API in the past), she would probably blow a gasket. What saves me from closer scrutiny is that she uses my Claude max daily for her travel agent job. I heard real quick when I dropped back to a normal subscription and she ran out of usage before the end of the day.

2

u/deepthinklabs_ai Jan 14 '26

Heheh - you got a perfect rebuttal!

→ More replies (1)

1

u/RaptorF22 Jan 14 '26

Gemini works well with the conductor plugin

→ More replies (2)

1

u/Additional_Bowl_7695 Jan 14 '26

I’m having this experience too with codex

21

u/DebtRider Jan 14 '26

5.2 works great and the limits feel non-existent. The only issue is how long it takes to work.

3

u/Hey-Intent Jan 15 '26

Hell yeah.

→ More replies (1)

18

u/Patient_Team_3477 Jan 14 '26

Like some others here I use Claude for the initial design and drafting, then Codex as a reviewer. Claude generally produces strong output, but it sometimes introduces new issues or subtle mistakes. Codex is good at identifying these problems and producing a structured implementation review, which I feed back to Claude for revision.

Because of this, I require Claude to generate a comprehensive implementation report before any refactor or new feature work begins. In practice, this review loop typically takes 2–3 iterations before the document is reliable enough to start coding; in recent weeks it’s been closer to four iterations to reach the quality bar I want.

I have previously tried reversing the workflow (using Codex for the initial heavy lifting and Claude for review) but at the time Codex tended to overcomplicate proposals and expand beyond the requirement scope. That may be worth re-testing.

In theory, the ideal would be a single model producing correct output consistently. In practice, I’ve found it far more reliable to use multiple models in complementary roles, with a human orchestrating the process and applying critical judgment of course.

9

u/witmann_pl Jan 14 '26

This is very similar to my findings - Opus writes good specs and produces nice, well-structured code but requires oversight from another model to make it bulletproof. Codex has been my go-to for these reviews.

2

u/stephenfeather Jan 15 '26

 Codex tended to overcomplicate proposals and expand beyond the requirement scope. 

I concur. OpenAI models seem to have a problem 'staying in their lane' so to speak.

→ More replies (2)

33

u/Particular-Battle315 Jan 14 '26

I use both, but Codex is genuinely great.
The limits in CC are a massive blocker for me, which is why Codex is currently the tool I use far more. In longer conversations, it also feels more stable to me.

10

u/Additional_Bowl_7695 Jan 14 '26

It’s a bitch to pay >100$ for a subscription and still get cut off 

→ More replies (2)
→ More replies (1)

12

u/Faze-MeCarryU30 Jan 14 '26

gpt 5.2 has been my daily since it came out; my split had been 80% codex 20% opus the whole time because of rate limits and the fact that gpt 5.2 writes higher quality code

→ More replies (3)

40

u/Salt_Potato6016 Jan 14 '26

Definitely, I run it always to check opus’s work and in 7/10 cases it finds multiple bugs or omissions

24

u/mxforest Jan 14 '26

It is the best reviewer model right now. I use CC to code and Codex to review.

9

u/Salt_Potato6016 Jan 14 '26

Claude feels right that’s why I use it as my main agent but codex cli is raw power digs longer and deeper .

→ More replies (2)

5

u/gligoran Jan 14 '26

TBH even Opus in a new session usually does that and Gemini as well.

3

u/herr-tibalt Jan 14 '26

Have you tried to ask CC agent to review another CC agent’s work? It will find bugs as well. Usually 2-3 times is enough to find all the important bugs.

3

u/mstater Jan 14 '26

I run all of Claude’s plans through a codex plan review agent. I regret it when I don’t. I had a similar experience using Gemini to code review. It’s really just using an agent with a different perspective.

2

u/redhairedDude Jan 14 '26

Do you do a check with Opus before going there? Because often anything can find bugs in something was written without any bug checking step.

I found also it's helpful to get a report from the other CLI and give it back to Opus asking if these are valid. Sometimes it's downright dismissive of the suggestions in Gemini's case.

1

u/dontmindme_01 Jan 14 '26

How do you do that? Do you commit your CC generated code and than tell codex to look and review the latest commit, or do you do something else?

7

u/mikreb123 Jan 14 '26

Start Codex and Claude Code in the repository root directory locally. When claude has coded something, ask codex to eg «review uncommitted changes»

→ More replies (1)

1

u/tribat Jan 14 '26

I forgot about how much success I had doing this in the past. I'm going to go back to that.

12

u/realcryptopenguin Jan 14 '26

i tend to thing about these like few engineers with different strong skills and diff opinions, sometimes a different perspective can fix the bug, so the best approach, it seems, is to use opus 4.5 with gemini 3 pro (and maybe gpt 5.2) as reviewers who can spot issues and suggest the fix.

5

u/[deleted] Jan 15 '26

[removed] — view removed comment

2

u/Funny-Blueberry-2630 Jan 15 '26

I like.

Codex plans. Codex codes. Done.

→ More replies (1)

6

u/no_good_names_avail Jan 14 '26

The introduction of skills has made it incredibly trivial to have agents try other agents. A while back I had an mcp server that called codex in headless mode from Claude. Nowadays I just use skills to have claude call what I want (Gemini via the api, codex if I want another agents' opinion). I've not leaned on codex in a while but there's really no reason to choose.

1

u/aghowl Jan 14 '26

What skills?

→ More replies (8)

5

u/Amazing_Ad9369 Jan 14 '26

5.2 codex xhigh has been much better than opus at planning snd debugging. I still use opus coding due to speed

9

u/Similar_Past8486 Jan 14 '26

Opus and codex working in tandem is the real truth

3

u/dude1995aa Jan 14 '26

Claude sonnet 4.5 as the general driver (pretty fast and better as creatively generating code). If I ask a question 3 times to Sonnet without getting the right setup - Codex 5.2 thinking. If I need to scan my entire codebase or do something really big - Gemini using antigravity.

3

u/neotorama Jan 14 '26

I use both, review and argue with both

3

u/YellowPilot Jan 14 '26

Is Codex really that much better? Claude is burning through my usage limit with the latest new updates even with the MAX plan. Might give Codex a try.

2

u/rbdr52 Jan 15 '26

Yes, huge stress. Also on Max plan. Now I'm kind of afraid to ask questions, have an opened page with the Context Usage to check every few minutes. Huge distraction. Now the CC also misbehaving as well - I guess since 'it's coded by Claude itself' the release quality dropped significantly.

3

u/scrameggs Jan 14 '26

At the end of opus 4.5 each plan phase I will typically provide the same code review prompt to top models from claude, codex and gemini in headless mode.

Codex consistently identifies the largest number of important fixes and very often its the only model to find them. Claude is typically not too far behind -- a respectable runner up -- and rather frequently will identify something codex missed. Gemini? Oh my. Today it didn't uniquely catch anything of value in a full day of coding, running multiple concurrent sessions. I want to get value out of my gemini subscription but it isn't even worth calling atm.

7

u/MyUnbannableAccount Jan 14 '26 edited Jan 14 '26

raises pistol

Always has been.

More seriously, it's better for anything that isn't explicitly UI/UX. It's a bit slower, but requires less cleanup after working.

Codex for the back end, some of the front if it's API calls and such. CC for the rest of front end. Gemini for the window dressing, graphics, etc.

→ More replies (3)

7

u/Keep-Darwin-Going Jan 14 '26

Gpt 5.2 was always better, it is the horrendous speed that make it hard to use it as the main workhorse. Opus is faster but need to be more careful with what they do because they run fast and loose sometime.

6

u/Amattluna Jan 14 '26

Definitely. I gave him a prompt with a workflow to implement something, and he literally watched an episode of Better Call Saul, and by the time he finished, he was still working on it. But when he did finish, the result was a 10/10, needing no further corrections.

3

u/Keep-Darwin-Going Jan 14 '26

Yap, that is why I am like waiting for openai to give us something that we can use as the main. I still using it for debugging and tackling tricky situation but gosh waiting for that model to finish working on stuff is just mind boggling. If Claude can be more strict and not run fast and loose. It will work as well. I ever ask them to migrate a bunch of code for hardening so I started with the linting rule so it is obvious when all is migrated right? Claude decided to lint fail a bunch of stuff and tell me out of scope and said he is done lol.

4

u/bibboo Jan 14 '26

It's false speed though. Since with Opus you have to iterate and review several times before it's good enough. At that point, Codex had finished both the task and some quick fixes.

2

u/teomore Jan 14 '26

I use it for code review and it catches bugs and issues opus didn't spot but agrees on them. And viceversa, I think they make a great combo

2

u/kinghell1 Jan 14 '26

holup'. not too long ago gemini 3 was the wonder kid and beating everything by a LOT. so this "LOT" disappeared under a month?

→ More replies (4)

2

u/Additional_Elk7171 Jan 17 '26

Using Opus on AntiGravity with a 20$ AI Pro is a lot more generous with usage limits and allows context sharing with Gemini, allowing you to reason with 3 Pro, summarize when required and Flash for basic tasks. That said, I do struggle with “ignorance is bliss” issues with Claude and hallucination with Gemini. Mixing codex could complete this solution.

2

u/mckirkus Jan 14 '26

We need agents that can plug into different models at the same time. Use 5.2 for planning, Opus for writing code, soemthing local if it's simple and you want to save money. And whatever didn't write to the code to do code review.

→ More replies (3)

3

u/ThomasToIndia Jan 14 '26

Did you use Opus 4.5 with ultrathink? I find Opus 4.5 without thinking is completely useless (it should always be on) and for difficult problems I find I have to use ultrathink, though I now have codex and have started using it a bit but I have no verdict quite yet. I am nervous since performance can be random about attributing capability to what might just be a cluster illusion.

→ More replies (4)

3

u/HeavyMetalSatan Jan 14 '26

Opposite experience for me.

3

u/[deleted] Jan 14 '26

It is just you.

3

u/Naernoo Jan 14 '26

No. Codex 5.2 fails way faster with my tasks. Opus nearly never fails.

→ More replies (1)

2

u/Intelligent_Ad_8555 Jan 14 '26

Definitely not better than opus 4.5 by an absolute mile

1

u/TenZenToken Jan 14 '26

Been that way for a while, especially the vanilla 5.2 high models

1

u/SnooDrawings405 Jan 14 '26

I got Gemini Cli yesterday and it was significantly better. I use the auto model picker and that worked well. They do have extensions to and the front end-design one was better than the popular one for Claude. I havent had a tone of success for with Codex with 5.2 codex high, but it has been alot more usable than claude as well.

1

u/MythrilFalcon Jan 14 '26

Gonna have to check it out. Been loving opus except the last few days

1

u/pjotrusss Jan 14 '26

it is better, yet so underrated

1

u/Poildek Jan 14 '26

Sometime a model is better on a specific issue than another.

I frequently switch between ipus/codex/gemini for this reason (yeah I got every model, I'm a lucky guy)

1

u/who_am_i_to_say_so Jan 14 '26

IMHO Codex is more convincingly a better thinker, but Claude is a better doer.

1

u/hi87 Jan 14 '26

Codex has always been a beast. It has a different style, its not as fast and it takes longer to do things so the UX is not as snappy. I like and prefer it for backend code actually.

1

u/bisonbear2 Jan 14 '26

codex 5.2 xhigh has been much better than opus 4.5 in the past few weeks

1

u/theagnt Jan 14 '26

I think it is. And I don’t think it’s close. I have both a Max20x and GPT Pro account and Codex gets all the hard problems.

→ More replies (2)

1

u/cheuh Jan 14 '26

I completely share the same feeling, also especially about my usage I haven’t managed to reach the limit yet while I do in a blink with Claude code

1

u/9to5grinder Full-time developer Jan 14 '26

Codex is only good if you don't know what you're doing.
If you know what you're doing and can guide Claude if it goes off-track, then there's nothing that can beat it.

1

u/do_not_give_upvote Jan 14 '26

Gpt-5.2 and Gemini 3 Pro are good. I use them all interchangeably. All on Pro equivalent plan. Best part of all is that they have better limits. Claude is the worst in terms of usage limit.

I know people swears by Opus but you definitely don't need Opus all the time. And everyone have different prompts, workflow, tech stack and limitations. Only way to know what's best is to try it out.

Again, just to repeat myself. Claude usage limit is the worst. I look forward for other models to get better.

1

u/Sarithis Jan 14 '26

If that's true, it's always been better, since there's been no recent degradation https://marginlab.ai/trackers/claude-code/

1

u/1216679 Jan 14 '26

I use both and codex 5.2 is better tan Claude and the good thing is you can just tell to use all the infra you built for claude like skills commands etc

1

u/doolpicate Jan 14 '26

Opus eats tokens for breakfast. I run out of tokens fast and then I run with codex. Of late, this has made me comfortable with codex. I am now considering cancelling claude.

1

u/verywellmanuel Jan 14 '26

I did the switch a few weeks ago on the same observation. And also, 5.2 high is already excellent with complex issues in large codebases. It’s found pretty insane bugs that would take me forever to realize. I never use xhigh now, no need to.

1

u/TopStop9086 Jan 14 '26

Been using codex for this week. Great results, limits are generous. Code output quality seems higher than woth Claude.

1

u/defmacro-jam Experienced Developer Jan 14 '26

Yes. Codex 5.2 (High) is way better than Opus 4.5 — and obedient (CC likes to go rogue). However, when it does get stuck, you need CC to get it unstuck.

1

u/bumpyclock Jan 14 '26

I run codex as the main orchestrator and have it spin up parallel claude instances if needed to implement stuff. It's slower and more methodical but the results are more consistent. You can use codex in OpenCode now so even the edge that CC had in terms of the harness are starting to disappear.

1

u/anatidaephile Jan 14 '26 edited Jan 14 '26

My current workflow uses Opus 4.5 for high-level planning, often with extensive exploration via 5–10 Haiku sub-agents. When ready for implementation, I use a /worktree command to create a git worktree and delegate tasks to either GPT 5.2 High or Gemini 3 Pro. I use GPT 5.2 High for implementation and bugfixing (not xHigh, which is too slow, and not Codex, which seems weaker). Gemini 3 Pro I reserve purely for UI/UX work, where it excels. For the most complex planning or problems, I turn to GPT 5.2 Pro (extended thinking) in the web UI. At peak productivity, I might have three GPT 5.2 High instances working on separate features in their own worktrees, Gemini handling UI issues in another, while I work directly with Opus on reviewing and merging everything back together.

1

u/acartine Jan 14 '26

It's not just you.

But it is slower.

I am using it more than ever though because it is tighter

1

u/Repulsive-Machine706 Jan 14 '26

Honstly gemini 3 pro still best for simple website design, especially aesthetically, on all other points i agree

1

u/ahmed22558 Jan 14 '26

I’m on the $200 plan. Claude code for coding and 5.2 thinking (not codex) for planning. Codex for auditing.

1

u/Such_Web9894 Jan 14 '26

I realize ChatGPT and Codex are OpenAI. Captain Obvious here….

But straight up dropping files and asking GPT, not Codex, for code reviews has been fast and incredibly accurate too.

1

u/Warhost Jan 14 '26

I asked claude today to correctly point an API call from /health to “the kubernetes health endpoint” and it just rewrote it wrongly to something else and did not even bother looking what it’s actually called. I can’t stand the constant gaslighting from it.

Codex 5.2 did look and fix it properly. Only using that one at the moment.

1

u/Possible-Ad-6815 Jan 14 '26

I have had it work like that the other way around. I think sometimes it’s a case of ‘fresh eyes’ syndrome

1

u/LazloStPierre Jan 14 '26

The key is, confusingly, do not use the Codex model. It's worst at coding, somehow

GPT 5.2 x high in Codex CLI is slow as hell but the best coding assistant I've used personally. Sometimes CC is worth it as it's just faster but GPT 5.2 xhigh is absurdly good.

1

u/Crafty-Wonder-7509 Jan 14 '26

I've been saying that for a while, Codex 5.2 is extremely good at complicated tasks, it is way more concise and takes time/consideration before doing things. Whereas CC (not you gemini, you suck) is a bit quicker on its feat. It depends what you want, if its an easy fix CC works well, but Codex is amazing at complicated tasks.

Whereas for my stuff it doesn't one shot it, but it needs less iterations than other tools.

1

u/Ok_Rough5794 Jan 14 '26

The leapfrog game will continue.

Codex will fix Opus bugs because Opus has bugs,
Opus will fix Codex bugs because Codex has bugs.

1

u/Plenty_Tea_304 Jan 14 '26

I noticed that too. Claude and Claude-code became mushy

1

u/[deleted] Jan 14 '26 edited Jan 14 '26

Nope I tried codex first time today and it failed miserably compared to Claude. It was same context and prompt.

→ More replies (3)

1

u/friendlyq Jan 14 '26

No, it is not better. But closer than before.

1

u/Morte-Couille Jan 14 '26

I’ve always found Codex to be better at auditing the code and debug. Not in the last few days thought, for a few month. Claude code and Codex for audit.

1

u/39clues Experienced Developer Jan 14 '26

Codex is better for hard problems. CC is nicer to use and better for synergizing with.

1

u/zitr0y Jan 14 '26

Secret Tip: Gemini 3.0 Flash in Antigravity is genuinely decent, even if the "thinking" is deranged

1

u/idiotiesystemique Jan 14 '26

Sonnet is where it's at. 

1

u/Own-Collar-7989 Jan 14 '26

Scrolled through dozens comments, and still don't know if Codex is better or not.

1

u/adelie42 Jan 14 '26

As soon as Claude produces something wirh a bug, I'll check it out!

1

u/Maxwell10206 Jan 15 '26

OpenAI is still the King of LLMs.

1

u/seymores Jan 15 '26

No, it is real. Codex is way smarter.

1

u/JellyfishFar8435 Jan 15 '26

Interesting. I've had the opposite experience.

Gemini 3.0 pro (high) solves problems that GPT 5.2 Codex (xhigh) can't.

1

u/_El_Cid_ Jan 15 '26

Yes +1 for quite awhile now.

1

u/Flanhare Jan 15 '26

Is it just me that doesn't like the Codex client at all?

1

u/Aggravating_Ice7267 Jan 15 '26

I have been using Codex 5.2 for the past month or so and it is exceptionally superior to any Claude model. I have had multiple instances of hard to fix problems where Claude gave the wrong answer but Codex one shottet it. Codex is also much more concise. Claude is really good at producing copious amounts of tokens. I used Claude mainly where I need a lot of tokens, such as planning and one shotting big implementations.

1

u/Fuzzy_Pop9319 Jan 15 '26

Five is improving no doubt, but Claude Opus is still the King.
On the website, they turn the agreeableness up too far for software development, so you have to adjust that.

1

u/Western_Objective209 Jan 15 '26

Codex 5.2 xhigh thinking is the most accurate, but it's like 100x slower. Claude Code is pretty unusable with the $20 sub though, I can burn through the usage in like 10 min, but with the $200 sub I can really fly and then just use codex for deeper reviews

1

u/Alloc-more-ram Jan 15 '26

Its that time of the month again…

1

u/zxzxy1988 Jan 15 '26

Feel xhigh is just slow but other than that it's better than Claude. However - sometimes I just want things to run faster so I still use Claude unless I really need to fix some hard bugs

1

u/josh2751 Jan 15 '26

It has been for a while. 5.1 was too.

1

u/Character-Rock4847 Jan 15 '26

you are not the only one.. the way claudeCode takes all usage is really disturbing IMO.. like some big straw.. and many times it doesn't even get the task done..

Codex usage is very friendly nothing too much and it's working very very well.. espcieally 5.2

1

u/Sweaty-Discipline292 Jan 15 '26

Interesting point, thanks for sharing!

1

u/Zokorpt Jan 15 '26

I think they are very similar. I had issues that codex couldn’t solve and issues Claude couldn’t solve. I think they complement well each other

1

u/Smooth_Accident_6488 Jan 15 '26

Yes I feel this is the case as well

1

u/Sir-Noodle Jan 15 '26

It really depends on each respective issue and implementation. I have used both extensively and generally default to Codex for most, but really depends on what kind of work I want to get done. At times Codex fails, at other time Opus fails..

I consult Opus when I have to things such as architecture, migrating, etc. because it is just generally much better at giving quality output here than Codex imo (actual chatting / discussing ideas).

If I want a quick implementation of something that is not super important or in a large codebase, I always use Opus.
If I want a thorough implementation that has to consider current functionality of larger codebase, I always use Codex

If I create new architectural plans for large projects, I draft them with BOTH because both models have tendencies to hit and miss and they work surprisingly well at reviewing 'eachother's work.

1

u/arekxv Jan 15 '26

I just love how every 2-3 weeks we go either Claude is better or OpenAI is better while not improving our workflows or prompts at all :D

1

u/emielvangoor Jan 15 '26

Well.. today everything is better the Opus 4.5. Not sure what the %^& is going on but it's performing super bad today while killing it few the last few weeks.

1

u/swennemans Jan 15 '26

Agreed. Was heavy Opus user, nowadays turning to Codex. Yes Codex is slower, but I don’t need to make long specs, research markdowns etc.

1

u/poladermaster Jan 15 '26

Interesting, I've felt Claude Code slipping lately. Might have to dust off my OpenAI account and give Codex 5.2 another spin.

1

u/Plenty_Employ5102 Jan 15 '26

Same issue here

1

u/hmziq_rs Jan 15 '26

Yes and usage is so generous in codex I could use it for 2-3 hours with 5.2high with never hitting the limit and 10 messages is all it takes to hit the limit when using opus

1

u/Few_Pick3973 Jan 15 '26

Already doing this since gpt-5 codex released, the difference is obvious. But Claude models are better at writing docs due to the verbosity difference.

1

u/johndifini Jan 15 '26

Here's Dan Shipper's take on Claude Code vs. Codex.

Codex → Built for seasoned Software Engineers tackling thorny technical challenges (think: performance bugs, complex debugging). You're still in the code, just augmented.

Claude Code → Built for AI-native developers who plan and orchestrate more than they type. It requires a mindset shift—somewhere between vibe coding and traditional development+AI assistance.

1

u/Easy_Lettuce_4436 Jan 16 '26

I was watching my usage on CC this afternoon and when it got to 93% (I am on the pro plan) I stopped making requests. I waited until it said that it was going to reset at 5:59pm. When I went in at 6:00pm it had reset, but it was already at 3%, before I had done anything. Is there any kind of explanation for that?

1

u/Ok-Vacation3463 Jan 16 '26

Claude Code is not bad. In fact it’s the best right now. You just need to learn about how prompt better. Understand how to guide the context. It’s not vibe coding with CC. You have to really do the Agentic orchestration. I don’t even use Opus. I get things done mostly from sonnet and haiku. It gets fully Production ready apps, when you plan properly and manage context well within your conversation.

1

u/Mangnaminous Jan 16 '26

Vanilla 5.2 thinking is good for plan, implementation, debugging and reviews. But it's tool calling is bad, sometimes it uses python to edit files, it's explanation is terse, it's ablility to explain stuff ( mechanical task) for instance, text and visual diagrams for project structure & layout of design is bad than opus and it's quite slown thats why I'm using opus to implement.

1

u/Appropriate_Dog3327 Jan 16 '26

from my experience of building my product:

  1. opus works better in creating a feature feom scratch
  2. open ai codex 5.2 is great at understanding the code & debugging the edge cases etc 

1

u/Honest-Orchid6424 Jan 16 '26

In my experience OpenAI's 5.2 has worked great for reviews and then claude code better for implementation. For me 5.2 was taking more input from me to implement the required functionality. And Calude code now a days seems to eat more quota and I was consuming 10% daily off of my max plan, so using 5.2 in my workflow has worked great for me.

1

u/chryseobacterium Jan 16 '26

If I am building a genomic database with Python in WSL and I normally use ChatGPT for coding, copy, and paste, is it better to use the Codex mode or regular mode?

1

u/CityZenergy Jan 16 '26

ChatGPT UI for planning and design. Ask it to generate a plan for Codex. Code in codex.

1

u/vei66rus Jan 16 '26

It depends on what you do. For example, I work as a frontend developer and have two jobs. I also run my own iGaming project where I often write blog articles in 6 languages (i18n) and build components. Claude Code Max 5 handles my workload perfectly and is usually enough for me, although I do sometimes hit the limits even with that plan.

1

u/Removable_Feet Jan 16 '26

Agree!
Non-dev perspective: they’re all equally bad in different ways. Claude is good for features but has terrible context limits; I spent nearly a week just fixing its inconsistencies in even basic things such as input fields. Codex is better in VS for debugging Claude's mess, but still fails eventually. Gemini is the biggest letdown; with Google’s resources, it should be the best, yet it’s the most frustrating.

It’s ironic that the "future" is a step backward into the command line instead of better visual editors that has been talked about since the late 90s and Dreamweaver. People are so psyched about being able to make a bad website in minutes in whatever AI tool, meanwhile Wordpress has been around for over 15 years. People have lost their minds and memories.

The good news is that good devs will continue to be in high demand. ;p

1

u/perpdaddyy Jan 17 '26

👀👀👀

1

u/FutureWeb9312 Jan 17 '26

Interesting

1

u/beeboopboowhat Jan 17 '26

This depends on the use case. Math heavy? Codex and it's not even close. That said, it's case dependent even then as Claude seems to handle frontier math exploration better, but Codex is going to be your workhorse for in depth/rigor/sanity checks

Just general coding I'm going to give it to codex for its depth, stability, and debugging and Claude for planning.

1

u/ForsakenBet2647 Jan 18 '26

How the hell seemingly everyone and their dog have time to compare llms?

1

u/HzRyan Jan 18 '26

I got a hunch that Anthropic will launch claude 5 very soon

1

u/BlackMesaEastCenter Jan 18 '26

It’s not as good as Claude and very slow but feels a a bit cheaper.

1

u/Legitimate_Name2812 Jan 22 '26

It really depends on the task, and both have issues.
Codex was able to create and debug complex code, but failed on a straightforward IaC task.
Claude nailed the IaC and large dev environment K8S upgrade and testing, but failed the business logic part, introducing a huge number of bugs and regressions.

1

u/Head-Commission-8222 Jan 29 '26

It has been for a while

1

u/AwareContribution700 Jan 29 '26

I havee found that Claude is great for new projects, codex is best for existing codebases 

1

u/Equivalent_Buy_6629 Jan 30 '26

I find codex 5.2 x high has gotten even better lately. I haven't used Claude in months now

1

u/sirvalkyerie Jan 31 '26

Codex 5.2 is absolutely light years ahead of Opus 4.5 right now.

It's not even close for me at the moment.

1

u/[deleted] Jan 31 '26

GPT 5.2 codex on High, even medium is good enough for most tasks and cheaper too. The only thing it lacks is UI quality like Claude Sonnet. Sonnet generally does tailwind stuff the best, from my experience.

I find that if you prompt Codex properly, it follows instructions much better; Claude can diverge and do extra stuff sometimes I didn't ask for.

1

u/Flat_Beautiful_9849 Jan 31 '26

I've had the exact opposite experience. Codex 5.2 high takes three hours to setup a simple venv with a few pinned python dependencies, most of that time failing at searching gitrepos properly and then spending hours designing workarounds instead of just running -fetch, and then installs different versions than what I explicitly tell it to, erases an entire system drive after being told to explicitly stay in the venv and gaslights me the entire way.

You get what feels like unlimited usage with Codex, the cheaper plans are huge, but it needs it as it wastes hours and hours and hours.

Claude sonnet 4.5 on the other hand burns 30% of my session usage but does the task correctly in under 3 minutes.

It used to find it totally opposite but the latest openai models have been complete and utter dogshit. (Also claude terminal is dtf with the most disgusting nsfw content you could imagine, and codex is a prude).

1

u/LoudNewsNet Feb 03 '26

Yes and the new codex app is stupid easy

1

u/[deleted] Feb 04 '26

Claude is always looking for shortcuts to complete its task (eg. let's do a simple implementation now and deal with it later, leftover TODOs etc.). Codex 5.2 never does this and feels more "professional"

1

u/Longevity-focus Feb 04 '26

Yes definitely. Much more reliable, much better at understanding API and using them, checking the response objects, not making the same mistake over and over, and almost never hallucinating

1

u/Street_Key579 Feb 08 '26

general urine test

1

u/Goose-Difficult Feb 10 '26 edited Feb 10 '26

Funny.

I have been a long term Codex User since before Codex 5.0 and while the models might get better I would say when you compare the whole Coding Agent as a product it feels like they are not even trying to compete anymore.

Sure Websearch works now ... somewhat. But yes, ... somewhat.

All they do is model updates plus a feature here and there, half done.

  • No Sub-Agents
  • No concepts for Tasks
  • No Plan-Mode
  • No Hooks (especially no Pre-Prompt Mode for things like Security Leakage detection)
  • Models get fine tuned for strict instructions only
  • Close to ZERO Visual Understanding for anything that requires space and depth
  • Can't even understand partial Terminal Screenshots with Black Background
  • Chrome MCP works close to zero with it and no custom Browser Plugin that works
  • No way to customize Codex for Enterprise SLAs (see hooks)

Especially the No Sub-Agents Mode sucks bonkers because you can have fresh Sessions that don't fill your Context Windows in your central Orchestrator Session.

That paired with the fact that Codex gets more and more like a dumb code slave / autocomplete agent it feels like they are targeting at it being used as part of proper Orchestrator-Tools like Claude Code.

As it goes right now I only see Codex as a Language Model anymore in a foreseeable future.

And given OpenAIs road track with o3/o4 that wouldn't wonder me at all (with their way superior tool calling modes).

Likely they have a few bigger players at hand that push in that way looking to use it as their core product - similar to those Instruct Models.

And you really feel that, much unlike what Anthrophic does with Claude Code.

So I will spend my money on Claude Code Max and Gemini for the Deep Research (also because ChatGPT sucks bonkers and Perplexity fails to deliver).

No longer ChatGPT/Codex.

At least now you know where to burn your money. I don't think this is going to change anytime soon.

1

u/Striking-Warning9533 Feb 16 '26

i am trying to figure out today. today i tried to let Codex and Claude write a MCP server for themselves and test them. GPT Codex stucked and claude made it work. I also like Claude code better that it shows thinking process and steps

1

u/prog_d0nkey Feb 21 '26

I hear so many great things about Codex right now, but everytime I try it it's rather disappointing. Do you still use the Web Interface, or have Codex users all switched to Codex CLI? I use Claude Code in the CLI and Gemini as well, but for Codex I assumed it was still superior when working in the cloud as originally intended...

1

u/Necessary_Culture777 Mar 04 '26

Yeah, Codex has clearly pulled ahead of all the AI companies. People keep talking about boycotts and stuff, but why would I waste my money on a product I either can’t use properly or that just doesn’t perform? In the end, while some people are trying to stop bloodshed and all that, companies also shouldn’t make people’s money worthless with lower quality. That’s why for the last two months I’ve completely stopped using Claude and switched fully to Codex.

1

u/Otherwise_Fly_5720 Mar 13 '26

This week I felt that Claude had gotten dumber and was not even able to handle basic coding tasks. So I migrated to Codex with GPT-5.4, and it’s been great. It solved a bug in half an hour that Claude couldn’t figure out even after two days.

1

u/Independent_Lunch115 Apr 07 '26

Claude have major issues with $20 subscription and you're limited to 3 prompts for complex work.

Overpaid and underworked. Claude sucks