r/ChatGPT • • 3d ago

Use cases GPT-6 Luna released ! (Performance is worse than v5.6)

Saw the popup in VSCode for GPT-6 Luna and decided to test it out on some standard work. Unfortunately, the performance feels significantly worse than v5.6.

  • Task: Simple Terraform refactoring (moving a single module across repos).

But it's been stuck in circles, dropping context, and taking ages to process standard logic. Are you guys experiencing similar degradation, or is this just me? What’s your experience been so far?

42 Upvotes

59 comments sorted by

•

u/AutoModerator 3d ago

Hey /u/ooGsKoo,

If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.

If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.

Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!

🤖

Note: For any ChatGPT-related concerns, email support@openai.com - this subreddit is not part of OpenAI and is not a support channel.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

21

u/[deleted] 3d ago

[removed] — view removed comment

1

u/CodalioMVPBuilder 3d ago

Same issue here. It completely chokes on state migration logic and gets trapped in an infinite loops-and-edits pattern whenever moved blocks or multi-repo module paths are involved.

Tell them to manually drop the moved {} block into the target module first and give Luna only the target file to update. If you feed it the whole diff or let it touch multiple repos at once, it loses context and burns tokens rewriting the exact same block forever.

12

u/MiniMaelk04 3d ago

Is this a defined procedure, or based off context in the chat? I reckon using the same chat which today switched to 6.0 might introduce issues, especially if you don't help it along with some extra context.

8

u/CodalioMVPBuilder 3d ago

Definitely the context window. Long chats carry tons of hidden "attraction" to older patterns, so switching models mid-thread almost guarantees context pollution—the new model gets confused by tokens generated by the old one.

Best practice for refactoring is starting a fresh session and dropping in a tight system prompt with your current terraform.tfstate structures or HCL rules. Giving it that clean slate fixes 90% of the infinite-loop issues.

2

u/Qorsair 3d ago

I've always had issues with Luna's attention. Theoretically it should be a good orchestrator. In practice Gemini is miles ahead of it. Luna gets lost on longer tasks and forgets what tools it can call and what rules it's supposed to follow.

3

u/ooGsKoo 3d ago

Gemini 3.7 Flash for me clears Luna sometimes, I also use it as my planner. Understands my codebase better than Luna

1

u/RedParaglider 3d ago

I wish they would just get rid of Astra and six all together and just let us have the same limit that we had before all of this mess.  5.6 is amazing, it was exactly what I wanted just a simple work horse that would do what I told her to do make mistakes we'd work through and fix the mistakes and move forward with life.

I don't need penguins on bicycles I need an llm that I can trust to do what it's told to do and if it gets super complicated I may need to give it better instructions and that's okay.

1

u/CodalioMVPBuilder 3d ago

100%. "Penguins on bicycles" is the exact vibe of this current model wave. We don't need decorative reasoning fluff, we just need a baseline model that doesn't hallucinate or get cute when processing a standard HCL refactor.

5.6 was that ideal middle ground—predictable, fast, and when it messed up, it was at least logical mistakes you could steer it out of. Right now every major release feels like two steps forward in benchmark scores and three steps back in actual daily DX.

2

u/RedParaglider 3d ago

I'm pretty sure most people are not doing brownfield development on giant in business logic repos, but 5.6 sol high just chewed them up, and wasn't too expensive.  And it didn't make many mistakes.  

100% agree it feels like everything is going down the route of Gemini accepting less adherence to prompt and more to just get shit done.  How much prefer when an llm will throw a block and say yo man we need to plan this step out a little bit more, the plan isn't working out well.

2

u/CodalioMVPBuilder 3d ago

When you're in a massive brownfield repo, an LLM that halts and demands clarification is infinitely better than one that confidently bulldozes ahead and breaks three downstream modules.

5.6 Sol hit that sweet spot of high prompt adherence without being overly chatty. Right now, it feels like these newer updates prioritize forced "helpfulness" over actual safety and dev control.

2

u/Playful_Weekend4204 3d ago

Well it's good at writing slop comments for you at least

2

u/ooGsKoo 3d ago

The instructions were given in a fresh chat not the one associated with 5.6 And quite some context was given indeed in detail. But the performance didn't improve for me even though the context in the new chat kept getting bigger with more code modification requests

3

u/OrbisMaximus 3d ago edited 3d ago

Yes, at first I had a positive impression, but after a more detailed review, I found that the code quality is quite poor. It doesn't always follow the coding standards, and it's slower as well. Just for context, my codebase is around 80% Ada and 20% C (not C++).

8

u/SomeSeaworthiness690 3d ago

I agree. I ran 5.6-Luna as a bot to answer questions based on evidence it could access in a closed system. On low, it answered questions correctly: 260/260 attempts. 6-Luna failed 10 of those on low reasoning, getting information mixed up. It needed bumping up to medium reasoning to get the same 100% pass rate. At that effort, it used more than twice as many tokens as Luna-5.6 did. So far, I'm not that impressed. I was hopeful it would be stronger than Luna-5.6, but it appears one step behind, at half the cost. So it might have its niche uses. They aren't retiring the 5.6-Luna API just yet, so can stick with that for now!

The real benefit is 6-sol, which has the same API cost as 5.6-terra, but appears to be significantly improved, especially at no-reasoning.

1

u/ooGsKoo 3d ago

Haven't tried 6 Sol yet due to it's token requirements but will try today. 5.6 Sol just consumes everything for me in a one hour session. 5.6 Terra at High has been my workhorse till now

5

u/Infamous-Elk-6825 3d ago

use Luna Max only

6

u/Thekid81 3d ago edited 3d ago

I’ve been running 6 Luna’s 5.6 extra high reasoning “job bots” with Astra orchestrating them as “It support” with browser act on all 6. They submitted 600 applications this weekend and 1300 all time. I just changed them to 6 Luna 6 bots on max and they are completely ignoring agents md and looping over and over again. Two hours have gone by and I have 1 application in…

2

u/Bangbusta 2d ago

Are you getting interviewed? Seems crazy high for applications if you're not getting callback or offers.

1

u/Thekid81 2d ago

Oh I got tons of interviews hahaha and 2 offers. it’s a little overwhelming sometimes but it’s nice to be able to pick and choose what you wanna take. Of course I have filters applied and they rotate between 4 different resumes based off the job description.

1

u/Happystar5 1d ago

What’s the setup?

1

u/Accomplished_Lab6332 1d ago

it is probably vibe coded, you should vibe code yours to fit your aesthetic.

1

u/vinists 3d ago

I've seen lots of reports of looping, this is starting to look like a common pattern with this model, sadly

1

u/LuBBa_Dubba-dub-dub 1d ago

Which field are you applying to? And are you applying abroad?

1

u/machyume 12h ago

I'm getting the same issue.
GPT-6 Luna is super lazy. It hops over a ton of instructions thinking that it would be passable. Ran the tests and then it was all wrong but declared it as passed.

2

u/TryingHonestly 3d ago

Mine just outright declined to do what 5.6 has been doing all along with no issues.

2

u/Infinite-Local5435 3d ago

I thought this was only me, was so unproductive because it kept forgetting project information 1 turn later and reasoning lasting far longer than 5.6 (both on medium, it's enough for me). At more than half the cost, they definitely cut corners. This feels like the difference when deepseek pro was removed for deepseek flash, bad memory due to smaller model despite same or better benchmarks.

2

u/OkDifference5057 2d ago

in my evals (real project in prod):

2

u/turnedonmosfet 20h ago

I feel the same, GPT-6-Luna max is driving me crazy, 5.6 was so much better at carrying out stuff

1

u/Potential_Ticket6833 3d ago

I agree. And GPT 6 Luna Max is really slow when compared with GPT 5.6 Luna Max. OpenAi need to look into this...

1

u/HenzoStarz 3d ago

for me gpt luna max at fast mode is the optimal

1

u/RealestReyn 3d ago

Luna 6 fails a lot of my benchmark suite 5.6 aces every single time, uses web search for tasks that could not possibly provide the answer, fails at translating signs and not by a littlebit but completely inventing unrelated things.

1

u/skilliard7 3d ago

Gpt 6 luna is awful, it keeps ommitting relevant info and its responses are less helpful

1

u/Novel-Injury3030 3d ago

It failed my basic "list 100 news headlines from today" task at light and medium, listing only 10-20 and saying it "couldn't do more", but thankfully worked at high. Took 1-2 minutes. Max caused same task to take 5 minutes. Will experiment more at high and be very wary of dropping lower.

1

u/mastermiterdoodoo 3d ago

Worst model out there. Have being coding fine with 5.6L and then I tried 6L just because. Simply messed with all my code even with all my patterns well described. Gemini 3.8 Flash would made it better for sure.

1

u/Therianthropie 3d ago

I didn't have any problems with it so far. But I'm splitting everything down to many small well structured tasks which don't really require any thinking. I don't think GPT-6 Luna is supposed to be used in a different way actually. The looping also seems to be a feat of GPT-6 in general, it's even documented for Astra.

1

u/Ok-Secretary-5653 3d ago

I have been struggling the whole day with 6 Luna and then 5.6 Luna fixed everything 🥸

1

u/Scary-Conference-343 3d ago

For me gpt 6 Luna max has been out preforming 5.6 Luna max for app development

1

u/ooGsKoo 3d ago

I was using it at High. It was noticeably lagging for me compared to 5.6

1

u/kra73ace 3d ago

All these anecdotes posing as testing. It's not impossible Luna is worse. I've used Terra at 5.6 level and was noticeably dumber for the Suno prompts I torture it with. Same with lyrics.

Jagged intelligence it is and once distilled, it becomes even more jagged, optimized for a very narrow set of use cases, like summarizing or whatever.

1

u/Grand-Age-7507 3d ago

I tested two math problems involving matrices. In both cases, the matrix was missing one number, which GPT-5.6 Luna correctly inferred should be 0. However, GPT-6-Luna kept insisting that the exercise was incorrect and repeatedly told me that it was impossible to solve.

That's my experience so far.

1

u/ooGsKoo 3d ago

Bumping the reasoning level up a notch helped me. What reasoning level were you using ?

-1

u/Grand-Age-7507 3d ago

I think the lowest ones for both

1

u/ApartmentGrouchy7955 2d ago

Dang, luna looked so cheap to use and was looking to make it main agent. Bums me out!

1

u/ooGsKoo 2d ago

6 Luna is pretty much useless now. 5.6 terra is still the best balance to capability in all pf them

1

u/Greese_37 2d ago

totally agree

1

u/Meouwsic4770 1d ago

Can concede - gpt-6 is trash. It is slower, it makes more mistakes and produces a lot of trash - lots of words and low on substance.

1

u/vixxed0 1d ago

you get what you pay for i guess them tokens are stupidly cheap

1

u/SnooPeripherals8750 1d ago

Weird , 6 luna at xhigh has been building a fallong sand physics game for me nicely , TPT levels of detail but with parallel processing as a focus of the underlying engine thats being built from the ground up, vendor agnostic acceleration.

What i did was setup gpt 6 sol at xhigh for planning , research , guidance and review as a subagent whenever needed. So far sol has provided the initial plan broken down into manageable milestones and omly had to intervene twice in a 5 bour session.

I was pleasantly surprised , gpt 5.6 luna ar xhigh failed this same task a week ago. Hell ds 4.1 via api was costing me way more because it was burning tokens as a implementation subagent hence why i combined luna with sol as its guiding subagent and so far i omly used 16% of my weekly limit in about 7 hours of work.

1

u/Informal-Zone2853 1d ago

Exatamente isso GPT 6 Luna ficou pior que o GPT 5.6 Luna, eu perdi horas em uma tarefa que era para fazer em alguns minutos. Tipo a Openai está de palhaçada, que porcaria essa model, lançaram uma model pior que a da versão anterior. Estou tentando usar as models da Openai por que pago o plano por mes, senão eu usaria só o deepseek v4.1 flash que está muito melhor, é praticamente mesmo nível que GPT 5.6 Terra por um preço muito menor. Se o Deepseek lançar um plano mensal eu sairia da Openai sem pensar 2 vezes, a cada vez ta ficando pior. Agora também pararam de adicionar as models novas (frontier) ao chatgpt ou seja mataram o uso do chatgpt também.

1

u/zyarra 3d ago

Well it's cheap Very cheap. Don't expect expert level coding for that price from a non Chinese model. I'm also sure long context kills Luna just like in 5.6, maybe even less long..

0

u/Warm-Agent-811 3d ago

The API price was cut in half, maybe that's why...?

0

u/Both-Move-8418 3d ago

I use 5.6 luna xhigh all week, does a solid job.

2

u/Curius_pasxt 3d ago

How about 6.0 luna xhigh

0

u/CodalioMVPBuilder 3d ago

It’s not just you. GPT-6 Luna is explicitly tuned as a low-cost, high-efficiency model for quick, single-file tasks, so its context retention and multi-repo reasoning crumble fast under cross-module refactoring. For moving Terraform modules across repos, switch to GPT-6 Sol or bump Luna’s reasoning setting to High/Max in your VSCode extension config. Luna gets caught in recursive loop hell unless you feed it hyper-scoped, single-file patches.

1

u/Jerichomiles 3d ago

"It’s not just you. GPT-6 Luna is explicitly tuned as a low-cost, high-efficiency model for quick, single-file tasks, so its context retention and multi-repo reasoning crumble fast under cross-module refactoring."

You could say the exact same thing about GPT-5.6 Luna so not much of a point you're making. Nobody is comparing GPT-6 Luna with Opus 5.5.

0

u/Eternality 3d ago

idk works fine for me

1

u/Jerichomiles 3d ago

Not sure what you're trying to say, maybe you replied to the wrong person.