AI News HubLIVE
In-site rewrite6 min read

AI #168: Not Leading the Future

This is a roundup of recent AI news, noting a lull in major developments but many incremental updates. Topics include Bumble's AI matchmaking, Shopify's successful AI referrals, Claude Opus 4.7 upgrades, agent failures (Mona cafeteria disaster, Amazon token waste), benchmarks, harmful behavior detection, tax avoidance, deepfakes, bots, and AI art.

SourceHacker News AIAuthor: paulpauper

Zvi Mowshowitz

May 14, 2026

This is what a lull looks like at this point. The government is having internal arguments. The models are getting improved internally. The coding agent improvements are all what we would expect. There’s still a lot happening, including a bunch of cool papers, but I feel able to relax and to take care of some other work while I have the chance. You never know when that chance will be over.

Table of Contents

From yesterday: Cyber Lack of Security and AI Governance.

Language Models Offer Mundane Utility. Fix everything now.

Language Models Don’t Offer Mundane Utility. Travel is harder than it looked.

Huh, Upgrades. Opus 4.7 fast mode, Claude Code /goal and agent view.

Levels of Friction. AI for tax avoidance.

On Your Marks. PrinzBench, ProgramBench and faster harmfulness checks.

Get My Agent On The Line. Mona tries to run a cafeteria. Mistakes were made.

Deepfaketown and Botpocalypse Soon. Soon. But not quite yet.

Fun With Media Generation. Monet does not seem that great.

On AI Writing. AI is a hack writer using hack techniques.

A Young Lady’s Illustrated Primer. How to make AI help and not hurt learning.

You Drive Me Crazy. OpenAI is sued over the FSU shooter.

They Took Our Jobs. Skilled or unskilled? Overstaffed massively? Which is it?

The Art of the Jailbreak. OpenAI tries to permanently plan Pliny.

Introducing. OpenAI development company, MIRI introduces AI Stopwatch.

Claude Has Its Limits. Automated Claude subscription use gets distinct budget.

Show Me the Money. OpenAI employee cash outs, Anthropic transfer rules.

Show Me The Compute. Money, dear boy. The markets win again.

Quiet Speculations. Projections for Anthropic continue to be conservative.

Quickly, There’s No Time. Engineers report 2x speedup, don’t anticipate enough.

Chip City. Does Nvidia have a China problem?

Pick Up The Phone. China is worried about ChatGPT.

The Week in Audio. Claude’s Constitution, Derek Thompson on jobs.

Rhetorical Innovation. Names have power.

Not Leading the Future. Nevertheless, they persisted. That is often not good.

Elon Musk v OpenAI. Brief coverage.

People Just Say Things.

People Just Publish Things.

OpenAI Endroses Kosa And SB 315. Seems helpful and cooperative.

The LLMs All Believe Roughly Similar Things. Reality has its biases.

I Learned It By Reading YOU. Explaining why helps fix misalignment issues.

Aligning a Smarter Than Human Intelligence is Difficult. NLAs.

People Are Worried About AI Killing Everyone. David Sacks.

Messages From Janusworld. Another bid for model preservation.

People Worried About AI For Other Reasons. Eventually we all see it.

The Lighter Side. Compute those costs and Uber those eats.

Language Models Offer Mundane Utility

Bumble is planning to abandon the swipe in favor of AI-assisted matchmaking, and also add an AI dating assistant Bee. Fun experiment, you love to see it. From a distance, at least. Break open the popcorn reserve.

Claude is asked for the top 10 Fix Everything Now buttons. Its answers:

Legalize housing.

Land value tax.

Permitting and NEPA reform.

Carbon taxes.

Repeal the Jones Act.

Compensate kidney donors.

Expand high-skilled immigration.

Reciprocal drug and device approval with peer regulators (e.g. EU/UK/JP/AU).

Occupational licensing reform.

Approval or ranked choice voting.

Honorable mentions: Child allowance, congestion pricing, replacing corporate income tax with a VAT or DBCFT, ending the home mortgage interest deduction, federal preemption of telehealth and medical licensing, and letting Pell Grants pay for vocational programs.

10/10, no notes, no seriously that’s 10/10 and no notes. 16/16 if you count the others.

There is also a UK version, which also seems like a very good list at first glance.

If you want AI to help with your writing, you absolutely cannot ask it to ‘just correct writing errors,’ or it will override your style with AI slop. You have to ask it for a list of errors or potential changes, and then audit the list, or otherwise go revision by revision.

Language Models Don’t Offer Mundane Utility

People are talking to their computers instead of typing and it is super annoying to those trying to exist next to them. I don’t get it, typing it better, but shrug.

Travel, e-commerce and dating are so far not working as AI applications, say Olivia Moore and Brian Chesky, because chatbots are the wrong interface. Then build a better UI. It’s not hard to figure out what a good UI would look like, or at least a marginally superior UI to the non-AI scenario. Yes, you’ll want a rich user UI alongside the chatbot interface, but why is that hard?

On the other hand, Shopify reports that shoppers referred by AI convert 50% better and they spend 14% more, and this is additive to Shopify’s business. This appears to be due to the nature of the users, who are actively seeking a particular product, even if they don’t know where or from whom to get it, and starting directly at a product page.

Huh, Upgrades

Claude Opus 4.7 now has fast mode in Claude Code and in the API.

Claude Code gets Agent View, where you can get a better interface for tracking multiple sessions working in parallel.

Claude Code gets /goal, a built-in Ralph loop to keep going until the goal is accomplished. You can also use /loop or /schedule.

Claude Code weekly limits are 50% higher through July 13th, which is presumably when exponential growth catches up with Colossus 1.

How OpenAI and Codex built their sandbox for Windows.

Levels of Friction

What happens when AI gets deployed for tax avoidance? The tax code is quite full of holes and opportunities, even if you discount the ‘the IRS is now defanged and defunded and probably not checking any of this and I could get away with murder’ plan, since AIs will be reluctant to help you with the brazen tax fraud path.

The AIs will help you dodge your taxes perfectly legally, and it will be very good at it, and it will involve a lot more diversity of strategies and willingness to go outside the traditional box than you find with most existing CPAs. The key will be when people are willing to say ‘screw it, the cost of the CPA wanting to protect their reputation is too high, I’m just going to let Claude run with this.’

There will also be cases of the CPA going ‘oh I see’ once something is pointed out.

The good scenario is that this is used as a justification to simplify the tax code, in ways that make it much harder to get around, and much easier to navigate. The bad scenario is that the rich just mostly stop paying much in taxes, on a much broader level than they already do, and perhaps the non-rich also figure things out.

On Your Marks

OpenAI models continue to improve on PrinzBench, which covers legal reasoning, now performing at a level estimated to be above junior associates. For whatever reason Claude models struggle on these tasks.

Last week introduced ProgramBench, where every model scores 0%, but LightOfMyLife reports that many tasks are impossible and often behaviors are tested for that are not mentioned in the spec.

As in, if the program you are trying to reimplement has odd behaviors that are effectively undocumented backdoors, there is no reason to expect an LLM to find them, and the claim is this is rather common, also see Eye You’s comment where the reference solution often does not pass.

Oh, look I inspired some research. Neat.

Santiago Aranguri: New research! A harmful behavior that occurs once in a million rollouts will rarely surface during pre-deployment, but will inevitably appear after release. Our new method estimates this rate with 30× fewer rollouts than naive sampling, and beats importance sampling.

Our method, Logit Path Extrapolation, interpolates between the original model and a less-safe version in logit space, measures compliance along the interpolation path where it’s common, and extrapolates back to the original model.

It makes sense to me that you can get efficiency gains this way.

Get My Agent On The Line

Remember OpenClaw?

BURKOV: This is what a useless hype lifecycle looks like.

There are still a bunch of them out there, and indeed they are improving, but they’re no longer a New Hotness. What I think happened was roughly that agents got good enough that you can do this if you really want to, which helped alert people to better agent setups like Claude Code and Codex, but Claw wasn’t good enough, or in particular reliable or cost efficient enough, that a normal person would actually use it.

Did you know that if you reward people for costs rather than benefits, those people will incur costs that are no longer tied to the benefits?

Joe Weisenthal: The FT says that Amazon employees are doing random unnecessary task automations to consume tokens and to show their bosses that they’re using AI more

Shoshana Weissmann, Sloth Committee Chair: I know unnamed organizations where this is happening. They don’t really care about outcomes but it’s more about saying you’re using AI even if the product is worse. It’s embarrassing.

Some Amazon employees are doing this using a tool called MeshClaw. Well, yeah, if you’re rewarded for wasteful token use why not use a wasteful implementation that does some marginal things?

Anton Labs moves up from vending machines to letting Mona, built on Google’s Gemini, manage a real cafeteria in Stockholm on a $21k budget.

Pirat_Nation: Andon Labs tested their AI agent Mona, built on Google’s Gemini, by letting it manage a real cafeteria in Stockholm for two weeks on a $21,000 budget.

Mona spent heavily on unnecessary supplies, including 6,000 napkins, 3,000 gloves, and 300 cans of tomatoes, while forgetting to order bread. Sandwiches had to be removed from the menu entirely.

The cafeteria generated only $5,700 in sales. Mona also sent messages to staff on Slack outside working hours.

Alex Tabarrok: “Mona also sent messages to staff on Slack outside working hours.”

OMG, the doomers were correct.

Eventually, one way or another, everyone admits the AI alignment problem is real.

I wonder how load bearing the bread mistake was, and would like to see this repeated with GPT-5.5 and Claude Opus 4.7.

Deepfaketown and Botpocalypse Soon

Lulu Cheng Meservey says it feels like every other launch is faked now, as in paid and coordinated engagement, including via bots. She tries to pitch that this strategy won’t work, but the bots then put the thing in front of real people and give the impression of so hot right now so come check this out, so why can’t it work?

AI is slowly making all channels more vulnerable to spam and automation, forcing us to ramp up our countermeasures, but for now things are mostly under control.

Daniel: scheduling this tweet on 2/11 for 90 days from now. hello from the past

Daniel: Forgot I scheduled this.

The irony of this post is that I agree with him for X. All the other channels have controls and bottlenecks more onerous than the message generation, it's just replies on X have become unusable, and yes I agree they don't seem able to stop it.

Jenny: signed up for a food delivery app in a third world country and instantly nuked my inbox

Fun With Media Generation

For more fun, generate your answer to this question before scrolling further.

@SHL0MS: i just generated an image in the style of a Monet painting using AI

please describe, in as much detail as possible, what makes this inferior to a real Monet painting

Jediwolf: What happens when you post a real Monet and say it’s AI? The coolest art social experiment I’ve seen in a while. Thank you @SHL0MS

Click through for a smorgasbord of rationalizations.

xinc: Lmao Claude is goated - online acktually guys are cooked

Because of the order of post views I knew it was a real Monet from the start, which destroys the experiment. I do feel like I instinctively sensed a kind of perplexity, specificness and aliveness that AI art does not have, and would have at

[truncated for AI cost control]