AI News HubLIVE
In-site rewrite2 min read

AI system Locus posttrains a model better than Qwen3

Post Log inSign up Post Intology @intology The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qw…

SourceHacker News AIAuthor: yuchiz

Post Log inSign up Post Intology @intology The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model. Today, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇 PostTrainBench evaluates agents' ability to post-train models on various domains given 10 H100 hours. We extend PostTrainBench via PostTrainBench+, which has a greatly expanded compute budget that provides clearer signal on automated post-training capabilities. We find that thousands of H100 hours help distinguish methods' performance post-training Qwen3 1.7B-Base models, and that Locus scales best. In this setting, modes trained by Locus collectively surpass the perforamce of the offical human post-trained Qwen3 1.7B model. In a test of generalization, we ran Locus on all live Kaggle competitions with prize money and public leaderboards. After 16 days, Locus achieved the 4th highest average rank among all participants. 4:42 PM · Aug 3, 202613.4KViews Intology @intology 2h Locus is state-of-the-art on PostTrainBench. PostTrainBench gives agents 10 H100-hours to post-train open-weight models on seven benchmarks spanning domains from healthcare to coding. The benchmark uses the performance of the existing human-post-trained models on each of the 548 Intology @intology 2h We further analyze methods’ performance on PostTrainBench+ and find that large-scale experimentation is a crucial capability for complex post-training domains, like competition math. Locus is the most effective at deriving significant performance gains from large-scale training 418 Intology @intology 2h Locus is already translating these results to real business value. Earlier this year, we began working with @bubble, the leading no-code app development platform. Locus, our autonomous research system, discovered and executed a post-training recipe that fine-tuned an open-source 306 Intology @intology 2h Locus’ capabilities go beyond post-training. While working with the PostTrainBench authors to verify our results, we ran Locus out-of-the-box on all prize-money competitions with public leaderboards active on Kaggle at the time, with no specialized setup or instructions per 312 Intology @intology 2h We would like to thank the lead PostTrainBench authors (@hrdkbhatnagar, @full__rank, and @maksym_andr) for supporting us in performing independent verification of Locus’ work in both the PostTrainBench and PostTrainBench+ settings. All results on the PostTrainBench+ setting 310 Intology @intology 1h At Intology, we are interested in scaling AI scientists from cheap computational evaluations to longer, expensive evaluation cycles ubiquitous across science & technology. We see the scaling from our earlier work on smaller-scale MLE tasks to post-training on thousands of H100 140 Intology @intology 1h Blog: intology.ai/blog/scaling-a… 172 Intology @intology 1h GitHub: GitHub - IntologyAI/scaling-automated-post-training From github.com 155 elvis @omarsar0 1h Congrats on the release. The PostTrainBench+ curve stands out for me. Baselines flatten past 2,000 H100 hours, and Locus keeps climbing. Cool to see this, as it's more aligned with the long-horizon capability that matters for self-improving agents. 171 Join the conversation Read 7 more replies