待翻譯:“It blows my mind”-“It has a tendency to overengineer things a little”: Developers react to road-testing OpenAI GPT‑5.6 Sol
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:OpenAI made its family of GPT-5.6 models available to its app and API users globally at the start of July. The post “It blows my mind”-“It has a tendency to overengineer things a little”: Developers react to road-testing OpenAI GPT‑5.6 Sol appeared first on The New Stack.
AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
OpenAI made its family of GPT-5.6 models available to its app and API users globally at the start of July. Since launch, users have had the choice of three model versions: Sol as the most powerful; Terra for mainstream use; and Luna as the fast, efficient option for everyday tasks. Sol is the star Named after the Latin word for sun, OpenAI initially positioned GPT‑5.6 Sol as inherently more cost-efficient and capable of getting “more useful work from every token”, with the company’s most robust safety stack ever released. The company announced stronger protection for higher-risk activities and sensitive cyber requests, and spent “multiple weeks finding weaknesses” to pressure-test the system against real-world attacks. Yet for all that galvanization, developers appear to be most interested in GPT‑5.6 Sol’s core functionalities and problem-solving capabilities. Dallas Fort Worth-based AI research & development leader at content automation platform company BlogBuster, Russell Twilligear tells The New Stack that he’s been using GPT‑5.6 Sol and drawing direct performance comparisons to results from Anthropic’s Claude Opus 5. “Sol is doing a better job creating things for me and finding mistakes made by Anthropic’s models; it’s blowing my mind,” Twilligear says. “I’m creating a massive nationwide database (USA) with over three million entries, and Sol keeps finding mistakes made by Opus and correcting them. So, I did a reversal of roles: I asked Sol to create the database and then asked Claude Opus 5 to check it, and it didn’t really find anything wrong (only super minor things that wouldn’t affect it)… so it’s currently something of a challenge to work out where all the lines are drawn in this market,” Twilligear adds. “Sol is doing a better job creating things for me and finding mistakes made by Anthropic’s models; it’s blowing my mind.” Twilligear’s comparison is interesting, although it is not exactly a conventional apples-to-apples model test. Using one model as a generator and another as a reviewer doesn’t necessarily categorize one model as “better” than the other; it does show that they behave differently when tasked with these specific generation or auditing workflows. While this might not be a conclusive measure of model evaluation, it does provide us with a tangible real world developer experience, and there’s surely value in that. Founder of AI-driven strategy and skills company Concepts in Success and independent researcher in AI governance and model evaluation, Joshua Estrin PhD, tells The New Stack that, for his money, this kind of database comparison scenario is a compelling insight into developer experience, but he thinks it’s “not enough on its own” to establish that one model is more accurate than another “without checking both outputs against ground zero in terms of truth” as well. “Using one model to generate work and another to review it measures something narrower than overall model quality,” Estrin says. “If one model does not flag errors in another model’s database, that may say more about what it was prompted or equipped to look for than whether the original model made fewer mistakes. Before calling a winner, I would want both outputs checked against ground truth – not simply against each other.” “I used GPT-5.6 Sol with ultra reasoning effort… it was effective across a much wider range of problems and much better at sustaining long, rigorous mathematical searches.” Six open Erdős problems New York-based PhD candidate at Columbia Business School, Shouqiao Wang, posts on X to explain that he “solved six open Erdős problems in five days”, all using OpenAI GPT-5.6 Sol. As all fans of central European approximation theory and discrete mathematics will know, Erdős problems are unsolved mathematical postulates and questions proposed by Hungarian mathematician Paul Erdős. “I have a math background, but the Codex workflow I used does not require deep mathematical knowledge,” posted Wang. “[But] the secret is model selection. I used GPT-5.6 Sol with ultra reasoning effort. Compared with earlier models, it was effective across a much wider range of problems and much better at sustaining long, rigorous mathematical searches.” GPT-5.6 Sol ultra reasoning level is one louder Let’s just pause to note that OpenAI defines ‘ultra’ as its highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster, meaning that multiple subagents are able to work on different parts of a bigger problem in parallel. After low, medium, high, extra high and max, ultra is one louder because it trades higher token use for stronger results and faster time-to-result on demanding tasks. While OpenAI lists these six levels on its most recent product announcement, other pages from the company specify six levels, but start with ‘none’ and end at max. Cosmetic labelling taxonomy conventions notwithstanding, we can infer that Sol has a six-speed gearbox. “I then pasted the prompt into Codex, set it as the goal, and let it run,” detailed Wang. “The reason is that Codex can work for long periods, retain the full research context, and use local files with no further interaction needed. You need to be patient and give it enough time to explore,” he added. “Perhaps the main issue I’ve found regarding Sol 5.6 is that it has a tendency to overengineer things a little. And for some frontend tasks, I’ll still occasionally jump to Claude or Gemini to solve smaller issues.” Walking the walk, a booking system developer’s experience Edinburgh, Scotland-based founder of tourism booking system automation software company Viamki, Jean Bustinza, tells The New Stack that GPT-5.6 Sol is his “go-to model now” and for good reason. “Sometimes, before asking it to output any code, I find myself having long conversations about higher-level tasks that used to feel out of scope for these models, like broader architectural questions,” Bustinza says. Likening it to being “a bit like talking to a senior software developer” in real life (with the caveat that every so often it will say something that doesn’t make any sense), Bustinza enthuses that generally, the information and reasoning he gets from GPT-5.6 Sol are extremely valuable. “Perhaps the main issue I’ve found regarding Sol 5.6 is that it has a tendency to overengineer things a little. And for some frontend tasks, I’ll still occasionally jump to Claude or Gemini to solve smaller issues,” Bustinza qualifies. Who actually wins the model race? If we’re looking for a winner in the model performance stakes, then the general consensus leans towards an “it depends” answer based upon what any individual developer is working to apply the model functionality to. We need to remember that one person’s frontier model long-horizon multi-day autonomous task analysis is not necessarily equal to another person’s multi-turn task completion and complex scientific cyber research workload – and that’s before we start worrying about cost per token and price performance. Let’s keep listening to what the developers say. The post “It blows my mind”-“It has a tendency to overengineer things a little”: Developers react to road-testing OpenAI GPT‑5.6 Sol appeared first on The New Stack.