AI News HubLIVE
In-site rewrite1 min read

A World That Answers Back

The article explores the asymmetry between cheap self-modification and costly trustworthy evaluation in AI systems. It highlights how proxies like benchmarks and reward models fail due to Goodhart's law, and draws parallels with recommender systems optimizing engagement over true value.

SourceHacker News AIAuthor: dloss

Truth happens to an idea. It becomes true, is made true by events.

— William James

Self-modification has become cheap. Trustworthy selection has not.

An agent can now rewrite its own memory, prompts, tools, and code for pennies. Whether any of those rewrites made it better remains as expensive to know as ever: better at long horizons, under real conditions, in ways that survive contact with the world. The field is racing down the cheap half of self-improvement. This essay is about the expensive half, and about the world we built because of it.

We have run the experiment of cheap judgment many times, and it always ends the same way. Benchmarks were judges, until models saturated and memorized them. Reward models were judges, until optimization pushed the proxy score up and true quality down; that curve is now one of the most replicated results in alignment research. Preference ratings were judges, until they bred sycophancy. And outside the laboratory, the largest optimization loop ever deployed produced the most familiar divergence of our era: recommender systems maximizing engagement, a cheap authored proxy for human value. The pattern is old enough to have a name, Goodhart's law, and a shape: the proxy rises, the target falls, the curves open like scissors.

pressurere-ground the evaluator in external consequencesgap: 12%