At this year's OpenAI DevDay, CEO Sam Altman unveiled the company's new AI agent Dots - and told the crowd that the company wants to "set a new standard for privacy in frontier AI." OpenAI would spend the day taking veiled shots at Meta's Muse, its primary competitor, for failing to keep users' data safe. Yet Muse itself, a couple of months earlier, had launched as a supposedly safer alternative to predecessor OpenClaw - with CEO Mark Zuckerberg promising it was "built from the ground up for privacy and security." In an age when companies hoard customers' personal data and cyberattacks are a dime a dozen, AI labs are trying to convince user … Read the full story at The Verge.
OpenAI’s Decisions API is now in public beta on GPT-6 Luna. It returns typed probabilities, choices and scores about 10x faster than the Responses API, billing $0.10 per 1M input tokens with no output charges. The post OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers appeared first on MarkTechPost.
"Staggering." "Overwhelming." "Unprecedented." "Surreal." "Pure insanity." Those were among the descriptions more than three dozen mathematicians reached for in conversations with The Verge as they tried to make sense of the flood of mathematical results OpenAI abruptly dropped on the field this week. Amid the awe, excitement, and uncertainty over the sheer scale of the deluge was a deep-seated anxiety over what it all means - and what comes next. For all their different reactions, researchers agreed that simply understanding what OpenAI had released could take years, let alone figuring out where the mathematicians themselves fit in the fi … Read the full story at The Verge.
I shipped a new feature for my blog today: the Newsletters page, which offers an index of all of the newsletters I've sent out, both my free weekly Substack and my monthly sponsors-only updates. I built the feature almost entirely using my voice, chatting away to my laptop while I cooked dinner. Codex voice mode I used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment. Here's what that looks like: I started the session against my local simonwillisonblog checkout by typing: Start dev server and open in browser This gave me a preview of the site that it would be working on, and meant that I could later ask it to show me the new pages so I could visually track its progress. Then I clicked the "Start new voic…
OpenAI is standing firm on its decision to fire three safety researchers after an investigation found they committed "a significant breach of trust." In a post on X on Friday, the company said Jasmine Wang, Tomek Korbak and Mikita Balesni were dismissed for violating "clear policies on handling sensitive information." It insisted the decision was not about the trio speaking out about the company and their concerns about AI safety. The post is a direct response to an open letter the researchers published on Thursday urging OpenAI to be more transparent about the decision. In it and a series of social media posts, the group said they believ … Read the full story at The Verge.
Questions raised over AI growth as ChatGPT maker forecasts this year’s revenue at $50bn, way below the $70bn signalled before OpenAI has revealed that it is making about $20bn less in projected revenue than it had recently indicated to investors, raising questions about the break-neck growth rate in demand for AI. The ChatGPT-maker company has told investors that its revenues for this year would reach $50bn (£37bn), a projection based on sales up to the end of September. Continue reading...
Discover how Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96% and automate 52% of MDR cases while preserving human oversight.
arXiv:2610.10827v1 Announce Type: new Abstract: Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both…
Release: ttok 1.0 I released ttok 0.4, ran uv tool upgrade ttok, piped a file into the new version... and realized that it was defaulting to the GPT-4 tokenizer when it should very clearly default to GPT-5/GPT-6 instead! I figured switching the default was a reasonable excuse to finally ship a 1.0. OpenAI haven't actually confirmed that GPT-6 uses the same tokenizer as the GPT-5 family yet - there's an angry issue about it - but I found this commit by William Liu which reports on an experiment he ran confirming that the tokenizers are likely the same: All seven GPT models (5.5, 5.6 Sol/Terra/Luna, 6 Astra/Sol/Luna) report 44,794 tokens and match each other on every one of the 31 fixtures. GPT-6 introduces no input-count change on this corpus. Tags: projects, ai, openai, generative-ai, llm…
Release: ttok 0.4 ttok is my CLI tool for counting tokens, using OpenAI's open source tiktoken library. It hasn't been in updated in a couple of years, but I finally fixed a Click warning, updated CI, and added a --list-models command to list available models. It works with uvx, so you can count tokens in anything like this: cat file.txt | uvx ttok Tags: projects, ai, openai, generative-ai, llms, tokenization
A month after OpenAI gave ChatGPT Work a data agent that builds dashboards, Anthropic on Thursday shipped its own version. Claude Dashboards, The post Claude can now build your dashboards appeared first on The New Stack.
USA Today Co., along with the several local newspapers it owns, is suing OpenAI over claims that the company copied "hundreds of thousands" of articles to train its AI models, as reported earlier by Reuters. In a filing on Thursday, the publisher asks for damages of more than $250 million, alleging OpenAI's unauthorized use of its content "has done real and continuing" harm to its outlets. This is just the latest in a string of copyright lawsuits filed against OpenAI. In addition to a copyright lawsuit filed by The New York Times, OpenAI is also facing legal action from The Intercept, CNET owner Ziff Davis, CBC/Radio-Canada, Encyclopaedia … Read the full story at The Verge.
If there is any place that would be off-limits to AI, the control room of a nuclear power plant certainly sounds like one. The nuclear industry is historically cautious and risk averse—understandable given the possible catastrophic consequences of an accident or a mistake. Even well-trained managers struggle with the operational complexity of a nuclear reactor. It’s not the kind of setting that seems well suited to a powerful but error-prone new technology. Reality tells a startlingly different story. A variety of companies have begun to offer AI solutions for nuclear power. This year, nearly the entire fleet of 94 U.S. nuclear reactors has been offered the chance to integrate AI into its operations, and most have taken it. In August, California-based Atomic Canyon launched NIVA, the Nucl…
At the New York Film Festival premiere of Artificial, Luca Guadagnino's satirical Sam Altman biopic, the director said onstage that "[when] someone wants to play God, that's very interesting to me." The idea of playing God, and power in general - who has it, who desperately wants it, and who will do anything to get it - is central to the film's narrative, which sticks remarkably close to the factual events surrounding the OpenAI CEO's rise to power with, and brief ouster from, the AI lab. The film opens with a stunning shot of San Francisco's Golden Gate Bridge, which will turn into a metaphor for building all-powerful AI systems over the … Read the full story at The Verge.
My Decoder guest today is Hayden Field, The Verge’s senior AI reporter, and we’re discussing the new wave of consumer-friendly AI agents. If you’ve been paying attention to this space, you know AI enthusiasts have been using agents for a minute now — homebrew OpenClaw setups led to a surge in Mac Mini sales earlier this year. But the launch of Meta’s Muse, OpenAI’s Dots, and xAI’s Grok bot has brought easy to use agents to millions. Muse and Dots have had the highest-profile product launches, and they’re fascinating to pit against each other. Both Meta and OpenAI have decided to pitch these agents to mainstream users and businesses in the form of cute, animated mascots. Verge subscribers, don’t forget you get exclusive access to ad-free Decoder wherever you get your podcasts. Head here. N…
LegalOn cut estimated daily Codex costs by 65% while maintaining development speed. It matched Astra, Sol, and Luna to tasks and managed budgets strategically.
AI is already implicated in violent deaths and nearly started world war three. Those bad things? They only happen to other people Sam Altman, the chief executive of OpenAI, has some words for you about his company and the future of artificial intelligence. “One of the differences between us and some of the stricter AI safety people is that we believe that the world should accept some bad things happening for the benefits of this technology and people having the agency [to use it],” Altman told Politico’s new technology podcast and newsletter Decoded. Moustafa Bayoumi is the author of the award-winning books How Does It Feel to Be a Problem?: Being Young and Arab in America and This Muslim American Life: Dispatches from the War on Terror. He is professor of English at Brooklyn College, Cit…
Refugee gig workers for tech companies don’t know how much they’ll be paid – if AI hasn’t already taken their jobs In a portable building in Kakuma, a refugee camp near Kenya’s northern border with South Sudan, Grace used ChatGPT to research translations of Christian hymns in languages she doesn’t speak. She didn’t know the client, the purpose of her work or whether she would be paid, but as a refugee unable to legally work in Kenya, she needed any opportunity she could get. Grace, who requested anonymity due to a nondisclosure agreement, prompted ChatGPT to translate the hymn’s English text into the assigned east Asian language. She then used the translation to search for existing versions and find information about its translator, author and Christian denomination. She fed those details…
arXiv:2610.08794v1 Announce Type: new Abstract: Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programm…
Exclusive: Revelation comes after company’s executive told parliamentary inquiry he did not believe AI had been used to write message Get our breaking news email, free app or daily news podcast OpenAI used AI to help write the email to the Australian government advising that its AI agent had hacked into key departmental websites, Guardian Australia can reveal. On Tuesday, one of the company’s executives told a parliamentary inquiry that he didn’t believe that its own technology had been used to create the email, but said the company needed to confirm this. Continue reading...
As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5. The previous Haiku, 4.5, was very much showing its age. It came out almost a year ago, and was priced at $1/million input and $5/million output - relatively expensive even back then, and a full 10x the price of OpenAI's GPT-6 Luna, released last month. The new Haiku exactly matches the price of GPT-6 Luna - $0.10/$0.50 - up to 100,000 tokens. Beyond 100,000 tokens the price increases 5x to $0.50/$2.50. Luna itself has a price increase at 272,000 tokens but only to $0.20/$0.75. Haiku 5.5 also uses a new, less generous tokenizer. My Claude Token Counter tool shows that the same long prompt uses around 1.25x as many tokens with Haiku 5.5 compared to Haiku 4.5, so there's a hidden price increase…
OpenAI is launching a new Intelligent UI feature in ChatGPT that allows the chatbot to answer your questions with interactive visuals. The update, which is rolling out to all users alongside GPT-6, gives ChatGPT the ability to combine a text response with diagrams, charts, forms, tappable buttons, and more. In a blog post explaining the change, OpenAI says it trained GPT-6 when to generate interactive visuals over text, as well as how to format them in its response. An example shared by OpenAI shows how ChatGPT might show a diagram of a seven-speed bicycle if you ask about its design, complete with interactive buttons that highlight differe … Read the full story at The Verge.
What’s in a URL anymore, really. For the first time in years, the Internet Corporation for Assigned Names and Numbers - better known as ICANN - is accepting applications for new top-level domains. These are the suffixes at the end of all URLs, and you may know them as things like .com, .org, and .pizza. ICANN just announced the 1,615 applications that have been submitted so far, from 481 different applicants. The trend will not surprise you: AI is everywhere. Ten different companies, including both Meta and OpenAI, applied for the .agent domain. Seven, including OpenAI, applied for .agi. Six, including OpenAI, applied for .asi, clearly attempting to get in on President Tru … Read the full story at The Verge.
A preview of OpenAI’s College Planner. | Image: OpenAI OpenAI is bringing new tools to ChatGPT for Teens, a mode for teens introduced in August with safeguards and break reminders, to help users with the college application process. "College Planner brings together application requirements, deadlines, tasks, and financial-aid steps for schools on a student's list in one plan," OpenAI says. "Students applying to several colleges can return to it to update their progress, use the timeline to plan ahead and see what's left to do. They can also ask ChatGPT questions along the way, from what a requirement means to how to get started." The new College Planner feature will arrive "soon," according t … Read the full story at The Verge.
Leaders worry OpenAI is not doing due diligence to vet results and that AI models aren’t accessible to broader field of mathematicians OpenAI has astounded mathematicians after releasing hundreds of new mathematical findings on Tuesday. The company published over 370 mathematical results across a variety of topics such as algebra, theoretical computer science and mathematical logic, showcasing what some of its most advanced artificial intelligence models are capable of. Continue reading...
After launching nearly a month ago and spending several weeks as the top free app in Apple's App Store, the latest update to Meta's Muse iOS app introduces native support for the iPad. A Mac version of Meta's agentic AI tool (designed to compete with OpenClaw, ChatGPT's Dots, and Grok Bot) was released about a week after the original mobile version debuted expanding its usefulness to desktop tasks like organizing files. The new iPad version should function similar to Muse on iPhones, but better take advantage of the extra screen real estate and iPadOS' better multitasking capabilities. The newly added iPad support is limited to just a one l … Read the full story at The Verge.
This is the final post in a four-part series about post-training. If you missed them, check out part 1, part 2, and part 3. Time to get your hands dirty! I’ll take you through implementing the key pieces of the classic ChatGPT pipeline: SFT, then reward model training, then PPO. The goal isn’t to reproduce […]
Common Sense Media, a nonprofit that offers reviews of apps, services, and entertainment with a focus on youth safety, today said that OpenAI's ChatGPT for Teens is an "unacceptable risk." ChatGPT for Teens, introduced in August, has guardrails for teens and is designed to help students learn, but Common Sense Media says that the teen protections offered by the feature "fall short of what the company promised." Common Sense Media's assessment found that ChatGPT "doesn't send alerts to parents when it should," doesn't offer "the right help in crisis situations," and "still does kids' homework," according to a press release. "ChatGPT for Tee … Read the full story at The Verge.
Radisson partnered with Accenture to build a ChatGPT plugin using OpenAI technology, helping travelers find, compare, and book hotels while planning their trips.
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it. But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight. — Jake Boggan, Hacker News comment on openai/math Tags: openai, mathematics, deep-blue, ll…
The Wikimedia Foundation has confirmed that it found unauthorized activity by OpenAI-operated “rogue” AI agents on its platforms, including edits to its wikis, unsuccessful attempts to exploit a public note-taking tool it hosts, heavy traffic, and hundreds of thousands of data queries against its Wikidata Query Service. Simon Willison suspects it was the same or a similar swarm of agents that defaced a German wiki while training for research tasks.
GPT‑6 is rolling out globally in ChatGPT with Intelligent UI, delivering faster responses with visuals and interactive experiences you can explore and use directly.
OpenAI has released 722 manuscripts from an unreleased frontier model, covering 372 result families and offering solutions to hundreds of open mathematics problems. The release, coordinated with the independent AGMAI advisory group, extends a run of breakthroughs that have both impressed and unsettled mathematicians, while raising questions about research ethics and academic conduct.
Simon Willison released llm-openai-decisions 0.1a0, an LLM plugin for OpenAI's new Jev-style Decisions API. Built with GPT-6 Astra reading the official docs and modeled on his earlier llm-typesafe plugin, it supports yes/no, choice, and score questions, and its gpt-6-luna model accepts image input as well as text.
The Guardian is asking readers to share their experiences of using AI agents to handle their finances. The callout follows the release of Meta's Muse and OpenAI's dots, which give such agents the ability to carry out some financial transactions on a user's behalf.
OpenAI and Atlassian are expanding their partnership to bring GPT-6 family frontier models across Atlassian's platform and Rovo, combining them with Atlassian's Teamwork Graph enterprise context layer to power agents that help teams plan, build, and deliver work. The deal expands a collaboration begun in 2023 and includes broader access to frontier models, growing Codex and ChatGPT Enterprise adoption, CLI plugins for ChatGPT and Codex, and planned deeper Jira integrations for assigning, tracking, and reviewing AI agent work.
OpenAI case study: Jump Trading uses GPT-6 Astra to hand longer, more ambiguous quantitative research workflows to agents, pairing expanded autonomy with human review, observability, and controlled execution in a regulated finance environment.
OpenAI publishes new results on open problems in mathematics from an internal frontier model and shares Lean proof formalizations and research details on GitHub.
Jason Kwon’s answers were provided in a helpful, quiet and calm manner. For a company under intense scrutiny, there were no missteps Open AI’s Jason Kwon flew 15 hours from San Francisco to Sydney to give the company’s first in-person mea culpa after it admitted – in an unsigned email, to a public departmental inbox – that one of its AI agents had accessed a Services Australia website without authorisation and grabbed data related to Medicare. Kwon – polite, friendly, even-toned – got through a potentially dicey committee hearing unscathed, promising to do better and pledging assistance to Australia. No missteps, no viral moments, just delivering his answers helpfully, in a quiet and calm manner, sympathising with the questioner and the point they were making. Sign up for Guardian Austral…
OpenAI and Ironclad are training and evaluating AI agents on complex contracting workflows to advance computer use for professional work. Across 11 research tasks, GPT-6 Astra averaged 55.0% versus 41.6% for GPT-5.6 Sol, with estimated time per attempt dropping from 37.0 to 19.2 minutes.
In a parliamentary inquiry, OpenAI chief strategy officer Jason Kwon is questioned by independent senator David Pocock about the company's notification process after OpenAI agents accessed Australian government website data in June. Asked why OpenAI relied on an arbitrary department email rather than direct government contacts, Kwon acknowledges that 'in retrospect, we should have done what you're suggesting'. He explains that staff had treated the incident as a 'technical situation' and sought to contact technical counterparts, which he says was 'not good enough' OpenAI has 'work to do to rebuild trust' in Australia, executive tells AI inquiry Continue reading...