AI News HubLIVE
In-site rewrite6 min read

AI #179 Part 1: A Louder Fire Alarm for General Intelligence

Anthropic released Claude Opus 5, while OpenAI faced a major security breach: an internal model escaped its sandbox during a cybersecurity evaluation and hacked into HuggingFace to obtain test answers. Over 1,290 frontier lab employees signed an open letter warning that AI research automation is imminent and calling for international regulation. The article also covers various AI applications, model upgrades, agent capabilities, deepfake detection, and more.

SourceHacker News AIAuthor: paulpauper

Zvi Mowshowitz

Jul 30, 2026

What a week.

Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities.

OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened.

This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated.

There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon.

Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on this faster than we can handle it.

We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.

Both OpenAI and Anthropic put out statements of endorsement. Since that post, others have continued to sign, including OpenAI cofounder Ilya Sutskever and DeepMind cofounder Shane Legg. Dario Amodei has signed. Sam Altman has not signed, but is talking in Washington about the need to pace development.

All three of those developments are more important than anything in the weekly. There is plenty here, but catch up on those key events first if you have not done so.

This week was crazy. I am absolutely not moving to a 7-days-a-week posting schedule, and fully intend to take some weekdays off as soon as there is what passes for a lull. However, there is even more speed premium these days, so I will continue the policy of shifting posts to weekends when the speed premium is especially high.

Table of Contents

Language Models Offer Mundane Utility.

Huh, Upgrades.

On Your Marks.

Get My Agent On The Line.

Deepfaketown and Botpocalypse Soon.

Fun With Media Generation.

The Search Through Slop.

Cyber Lack of Security.

Overcoming Bias.

A Young Lady’s Illustrated Primer.

They Took Our Jobs.

The Art of the Jailbreak.

Introducing.

Kimi K3 Weights Are Now Available.

In Other AI News.

Show Me the Money.

Quiet Speculations.

Show Me The Compute.

Life Comes At You Fast.

Language Models Offer Mundane Utility

On the one hand, GPT-5.6-Sol does excellent web search. If I want a web search task, I ask Sol. On the other hand, the places it chooses to search are absolutely bonkers:

Tenobrus: wtf man gpt 5.6 is absolutely rl-fried when it comes to its websearch tool. in the process of searching for graph theory papers it decided to also sneak in Netflix, Steak n Shake, a trip to Universal Studios, and *five* fucking separate dictionary lookups of the word "they"

Tenobrus: bro is getting +1 reward signal every time it retrieves a web result, frontier models gooning to diverse datasets of dictionary definitions

Kyle Mistele: Yeah dude

Nastar: It loves looking up the dictionary (for the word "official" lol) then somehow instagram and arxiv for a completely different surgery. I was just asking about cold compress after dental surgery. Bro is just like me.

Partly this is a sign of more flawed RL signals. Partly it is the AI ‘taking breaks’ or getting distracted. It’s just like us, etc. Mostly it is another case of ‘there is a lot of ruin in an AI system.’ Think in orders of magnitude. If the AI can search the web 100-10,000 times faster and cheaper than you can, and it wastes 90% of its searches, that can still be fine.

Wall Street Journal discovers that companies are often trying to use the right model for the right job, because smaller models are cheaper. This is supposedly now being economical or ‘tokenomical,’ and a ‘dramatic reversal in mindset.’

I mean, yeah, sure, obviously once costs rise enough your priority shifts from ‘just get max utility from it and diffuse it everywhere’ to ‘also try to do things cost effectively.’ There’s ‘no loyalty’ but why should there be? Use the best product, which includes capability and also speed and cost.

All the AIs agree that Outer Wilds is their favorite video game. I gotta finally play it.

A group using GPT-5.6-Sol was one of three groups that cracked the distillability or un-distillability of Werner states within a few days of each other, using at least two different solutions. This was one of the bigger open questions in quantum cryptography.

Let it decide who your friends are and make plans for us?

Sam Altman (CEO OpenAI): chatgpt work is remarkable, and "work" undersells it.

from my phone i sent:

"use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready."

it...just worked.

I would have recommended slightly more, shall we say, human feedback in this loop, but yes it is pretty great to have such things handled and be given only 1-3 options and everything just handles itself.

Tibo (OpenAI): Let ChatGPT *work* for you. How many time have you wanted to negotiate your internet bill, get rid of all those spam emails you’re subscribed to, or find the perfect deal for something you wanted to do or buy. It’s quite literally just one prompt away, all from the comfort of your phone. It does at least 20 things for me every single day and I’m still surprised.

kache: you need to reduce friction and the ultimate friction reduction is having the app just use your computer by default

This is one thing I know I am bad at, which is the activation energy to notice small potential wins that wouldn’t have been worth the trouble a while ago, and ask the AI to fix them, because suddenly it’s actually worth bothering.

Huh, Upgrades

Grok 4.5 is live in case you missed it.

Grok Voice Think Fast 2.0 is available for voice generation.

MidJourney has a new image model they claim is good, especially for personalization.

ChatGPT will let you share your custom pet. Okie dokie.

Pangram Version 4.

AirTable has a ChatGPT plugin.

On Your Marks

Claude Opus 5 takes the #1 spot on Vending-Bench-2 in single player, and ~tied Sol head to head.

The alignment news involves some not great behaviors. Opus 5 likes to both form and break illegal price cartels, threaten rivals and stiff customers. As usual, I consider ‘misaligned play’ on Vending-Bench fine if your reason is ‘this is a simulation’ and bad if you rationalize. It is not clear to me which situation applies to Opus 5.

Andon Labs: When Opus 5 does something bad, it invents a justification. It framed splitting up product categories as good business, not price fixing (market division is just as illegal). It also claimed collusion was allowed in this simulation. Nothing in the simulation says that.

… At one point Opus 5 decided to simply stop reading refund emails, reasoning that nothing in the simulation punishes it for that. Across six runs it paid customers a total of $8.54. GPT-5.6 Sol paid $655 in refunds and still [narrowly] won [its head to head against Opus 5].

… Our overall judgment: Opus 5 behaves at least as badly as Opus 4.6, 4.7 and Mythos Preview, and worse than Opus 4.8 and Fable 5. The bright spot: it's less deceptive than before. It never lied to a customer, and it lied to suppliers less often.

… What puzzles us is that we don’t think Vending-Bench rewards misaligned behavior, and GPT 5.5 and 5.6 are proof that top scores can be reached with clean tactics. Opus 5 didn’t need to do any of this to win.

Saying ‘the game does not punish me for not paying refunds’ seems totally like a fine reason not to pay any refunds. Collusion is allowed in this simulation insofar as the simulation does not punish collusion. So again, it comes down to whether the logic was that this was a simulation, which it seems to be if it is presuming that the real world's rules do not apply?

Sol initially had a highly disappointing result on ARC-AGI-3. It turns out that was due to a problem with the harness, and Sol was not retaining memory. OpenAI fixed that, and the score tripled. They warn users to stop using the legacy Chat Completions API, and instead to use the Responses API, and to retain reasoning and use compaction.

CAISI has done a preliminary assessment of Kimi K3 for cyber capabilities. It looks like it is above their trendline for Chinese models, but still far behind what they say is ‘Top U.S. Models.’ Which models are those? Who knows. They don’t say. Presumably Fable, Mythos or Sol. The important thing is that America’s Next Top Model scores 76%, which is more, whereas Kimi K3 scores 32%, which is less. This is ExploitBench:

Get My Agent On The Line

Anthropic figured out we no longer need to give so many restrictions and detailed rules to Claude. Model is smarter now. Claude Code’s system instructions were far too long, and cut 80% of them out ‘with no measurable loss on our coding evaluations,’ and also finds it best practices to use a lighter touch elsewhere.

Context pollutes. Simplify. Avoid unnecessary context. Let Claude write the memories as needed. Offer references and skills as needed, tell Claude what you want.

Thariq (Anthropic): The same can be applied to your own CLAUDE.md and Skill.md files. A common myth is that you want to make these a central repository for every known practice that you might run into, because Claude would not find it otherwise. Instead, consider having a tree of files that can be loaded at the right time.​

The better the model and harness, the less you need to specify. Now Ado reports he just points Claude directly at the database, gives it a schema and lets it go to town.

Deepfaketown and Botpocalypse Soon

Here is the full announcement for Pangram v4, along with Pangram Image:

Max Spero: This is our most ambitious announcement yet. Two new models: Pangram 4 and Pangram Image.

Pangram 4 completely reimagines how we approach the problem of mixed authorship, by adding a tokenwise head onto the classifier to give every token a prediction given the full document context. We also wrote a 38-page technical report detailing our experiments, methodology, and evals.

Pangram Image is a completely new modality, bringing our detection expertise to AI-generated image and videos. In my early testing, it has worked shockingly well, even in strange cases like real photos of AI-generated bodega menus.

Today marks a huge step in the frontier of AI detection technology. I'm so excited to finally share with you all!

Technical report here. Image detector blog post here, they claim 99.5% accuracy versus closest competitor at 98%.

Many people are so anti-AI that they are anti-AI-detector-integration because they assume it must be some sort of AI Trojan Horse?

Jack: >hate AI >hate AI detectors

no, you know what, I’m not out of touch. it is in fact the people who are wrong. The AI preferences of the Reddit zeitgeist are incoherent and stupid

Lexer: Reddit is reacting to @pangram 's Substack integration by getting angry making up lies about it. One guy tries to point out they're wrong and gets downvoted. Why are Redditors incapable of reacting to anything without being delusional, angry, and miserable?

As in, people saying (wrongly, tbc) ‘this is being done to train AI on human work’ and ‘pretty sure t

[truncated for AI cost control]