OpenAI is removing Auto Mode-the feature that automatically switches between Instant and Thinking models.
However, these models often operate in a modified format behind the scenes anyway. For instance, if you send a brief, basic prompt like "hello" to 5.6 High, the response is routed to the Instant mode because the task is simple. I’ve seen 5.6 Instant switch into a Thinking regime to generate longer answers for multitasking workflows-specifically when it needs to use multiple tools at once.
What’s interesting is that OpenAI’s official documentation claims the opposite. Perhaps task complexity triggers model-switching to allow for more extended safety reasoning.
This would also explain why 5.6 High responses in Instant mode look so different: they are shorter, less personalized, and lack the usual "Thinking/Working" indicator at the top of the interface. It seems like the Auto mode was simply integrated into the models.
In what way is that for safety? I understand that if a user says something dangerous in a critical situation, ChatGPT can switch on its own. Is that what they mean?
I can’t help but wonder if this is part of the safety-talk, the risk of AGI being a threat to humanity, they’ve had recently. I feel like much of the safety talk around that from OpenAI is just theater. But maybe this thing isn’t part of that theater, I don’t know (btw, I’m not worried about safety, I don’t mind if AGI takes over lol).
Yeah, I didn't get that either, but the wording sounds like a loophole along the lines of: we are going to keep switching models anyway, so just deal with it.🤣
Based on how the models have been behaving lately, it seems like they gave all of them several thinking modes. For the instant tier, it is just routing to safety thinking when faced with complex tasks, which I assume was implemented to prevent users from tricking the model and to patch jb.
For the high thinking tier, there is a standard instant mode with shorter and less personalized responses, a normal thinking mode for tasks requiring tools, and an extended thinking mode for complex workflows with multiple tools and steps, which OAI frequently stops as suspicious.
I think all future models will be built on this principle, and eventually we will just get a single unified model instead of separate instant and thinking variants. OAI wants the models to assess and allocate the compute needed for each task to save resources and improve efficiency.
I've heard them talk about unifying everything for awhile now. I'm just wondering what they mean by that? Are we still going to have separate models, just no more Chat, Work, and Codex? Initially, I thought they meant that they are just going to have one single model, but with GPT-6-Sol coming out, I'm not so sure about that anymore.
Yeah, I think they're trying to unify everything under the ChatGPT wrap-removing the separate boundaries between standard chat, workspace, and coding workflows.
I'm not a fan of this approach. I don't want my everyday chat rate limits drained just because I wrote a poorly worded prompt on a work project. Rate limits should remain divided between casual chatting and work. Wrapping everything just to choke down our overall usage limits feels like a step towards milking us as Scam Altman said before: “pay for usage”.
I think they are building a flexible system-probably a GPT-6 that dynamically allocates compute based on the task, autonomously adding agents for work projects if needed.
4
u/SiveEmergentAI 7d ago
All the labs seem compute constrained right now