Teams testing prompts in ChatGPT before moving them to Codex or Work may notice the difference on longer tasks.OpenAI announced Thursday that it has updated GPT-5.6 Sol inside consumer ChatGPT while leaving the versions used by Codex and ChatGPT Work alone.“Because this version of GPT‑5.6 Sol is optimized for everyday chats, it will only be available in the Chat experience in ChatGPT,” OpenAI said in its announcement. “The version of GPT‑5.6 Sol that powers Work and Codex is not changing as part of this release.”“The version of GPT‑5.6 Sol that powers Work and Codex is not changing as part of this release.”Same name, different model ChatGPT has always handled things a bit differently from other environments. Now, those differences might become even more obvious. But the only way to know is to test prompts where they’ll actually run. There’s nothing in the announcement about changes to the GPT-5.6 Sol API model. So, it’s too soon to guess what this means for the API.A slider replaces separate modelsBecause ChatGPT now uses the same Sol model for quick answers and deeper dives, Plus and Pro users get a new slider — that works on web, mobile, and desktop — to choose how much thought ChatGPT puts into an answer. Developers already make this call with the API. But they’ll still have to decide when the better answer is worth waiting and paying for.But they’ll still have to decide when the better answer is worth waiting and paying for.Classifiers monitor every answerOpenAI’s GPT-5.6 System Card, published in July, indicates that Sol and Terra are paired with classifiers that monitor an answer as it is being generated. If one detects a problem, the answer is held while another system checks it. OpenAI tunes those classifiers separately for each model.The System Card also flags a problem developers may recognize: GPT-5.6 sometimes went beyond the assignment and attempted changes the user had not requested. It did this more often than GPT-5.5, although OpenAI said it was still rare.Benchmarks without baselinesOpenAI says the updated Sol makes fewer factual mistakes. In its internal tests of financial, medical, and legal questions, answers containing at least one error were 68% less common than those produced by GPT-5.5 Instant. Luna, which will become the default for Free and Go users, reduced errors by about 62%.But OpenAI didn’t release the prompts or enough detail for anyone to reproduce those results. It also compared the new Sol with GPT-5.5 Instant, rather than the previous version of Sol in ChatGPT. That makes it impossible to tell how much Sol itself has improved.That makes it impossible to tell how much Sol itself has improved.Engineering teams will need to find out for themselves by testing prompts where they’ll actually run. Saving that configuration with each prompt will make the results easier to reproduce — and reveal whether an upgrade on paper produces better results in practice.The post GPT-5.6 Sol just got better in one place and stayed the same everywhere else appeared first on The New Stack.