If you’re even seeing this, you should probably just go enjoy your weekend. AI News for 10/1/2026-10/2/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies! AI Twitter Recap GPT-6.1 Sol and Sonnet 5.5 Reshape the Cost–Performance Frontier GPT-6.1 Sol launch: OpenAI priced Sol at $2/$10 per million input/output tokens, compared with $10/$50 for Astra (pricing summary). Claimed results: It reportedly beats GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 points on AutomationBench (summary). Positioning: OpenAI staff describe it as “good, cheap AND fast” (@reach_vb). Codex usage: A global Codex usage reset was set for Oct 2 at 10AM PT (@reach_vb). Tool use: Sol reportedly “REALLY loves codemode,” consistent with GPT models being trained on it (@badlogicgames, codemode note). Agent Arena placements: Sol [Max] entered at #5 (+11.23%) with a $0.56 median cost per task (@arena). Sol cost comparison: That is 39% cheaper than GPT-6 Sol while scoring 1.52 points higher. It is 81% cheaper than Astra while landing within 1.04 points. Sonnet 5.5: Sonnet 5.5 [Max] debuted at #3 (+12.5%) and ranked #1 in the Chat category. It costs $2.74 per task, versus $1.58 for #2 Opus 5.5, which keeps it off the Pareto frontier (debut, frontier). Anthropic’s position: Anthropic models now hold the top three Agent Arena spots. Code and Text Arena: Sol briefly entered WebDev at #3 before Sonnet 5.5 pushed it to #4 (weekly recap). Sonnet on WebDev: Sonnet now sits 2 points behind GPT-6 Astra [Max] at 80% lower cost. Gemini 4 Argon: Argon [High] took #1 in Text Arena. Open models: MiMo-V2.6-Pro and Flash entered Agent Arena at #5 and #9 among open models. Other independent evals: WeirdML v3 finds Sol very token-efficient, close to Astra but with a lower peak. On the same benchmark, Sonnet 5.5 beats Opus 5 and Grok 4.7 beats Kimi-K3; these results are incomplete (@htihle). Reasoning style: Design Arena read 324 thinking summaries. It found that Astra hedges about 20× as often as Opus 5.5, while Opus commits early in about 4 of 5 summaries (@DesignArena). Step 5 Preview: StepFun’s model ranks #7 among open-weight models on Vals at $2.54 per task. It averages nearly two hours per task and has a 1M-token context window (Vals, details). Rumors (unconfirmed): Fable 5.5: Claude Fable 5.5 is rumored for next week and said to outperform an “Astra 6.1” that was reportedly delayed over security concerns. The poster says he cannot verify either claim (@kimmonismus, follow-up). GPT-6 Astra Lite: A “GPT-6 Astra Lite” listing has been spotted, which @scaling01 speculates is the same model as Sol (@scaling01). Decision models and open weights: llama.cpp added a /v1/systemone endpoint for local “Jev-style” decision-model inference (@ggerganov). Running locally: Models are launched with llama serve -hf ggml-org/Kev-4B-GGUF (@ClementDelangue). Jared Palmer published a post on how Kev 1.0 works (post). Ecosystem: Perplexity claims pplx-decider-v1-27b averages 85.7% across 11 benchmarks, ahead of Jev (@AravSrinivas). Clef decision models are now on Ollama (@lucataco). Skeptical view: @mervenoyann calls decision models a rebrand of zero-shot classifiers (tweet). Calibration analysis: A blog post links Jev-style calibration to value and Q-function prediction (@SOURADIPCHAKR18). webAI TwIL-LM3-Pro: This 3.66B model is post-trained from Granite 4.2. In webAI’s tests it roughly matches Qwen3-8B on formal logic. The Q4 GGUF is 2.09 GiB and the license is non-commercial (@kimmonismus). Reka RIDM: Reka released an inverse dynamics model under Apache 2.0. It is trained on games, generalizes to real video and extracts motor and camera actions (@RekaAILabs). Agent Harnesses, Assistants and Developer Tooling OpenAI dots: Sam Altman calls dot his favorite OpenAI product, saying it improves daily as it learns his workflow (@sama). Capabilities: Dot keeps context across apps, coordinates Codex tasks and flags items that need attention (@OpenAIDevs). Comparisons: One user prefers Grokbot’s multi-agent “chief of staff” setup (@kimmonismus). A DIY clone uses Pi, a Telegram gateway and any model (@_alejandroao). Muse Gadgets: Meta open-sourced ESP32 firmware and a Linux SDK for building hardware that works with Muse (@natfriedman). Muse Home Link: Meta made 5,000 units of its own smart-home bridge, free for subscribers while supplies last (@alexandr_wang, shipping). Extensible harnesses: DeepSeek Harness shipped desktop builds for macOS and Windows; Linux users install @deepseek-ai/dsh from npm (@deepseek_ai). Claude Code mods: Mods are plugins with middleware-like hooks into Claude Code (@lydiahallie). The new “You should know” plugin spins off a side agent that flags important output the user might miss (@ClaudeDevs). Pi Durable: Pi now runs on Cloudflare Durable Objects via agents SDK v0.26.0, alongside Pi’s v1.0 release (@mattzcarey, @badlogicgames). Context: @omarsar0 frames these releases as a shift toward malleable harnesses (thread). T3 Code orchestrator rewrite: The project passed 400K users (@theo). Its 4-month PR, with 823 commits across 1,912 files, has now merged (@maria_rcks). New features: The rewrite adds Pi support, cross-provider delegate_task, an ACP registry, thread forking, mid-thread model switching, subagent lineage views and scheduled tasks (feature list). Platform updates: OpenAI’s Agents API added one-call browser computer use, Bedrock Managed Agents and portable environments. It also claims 99.97% turn reliability and 20% faster tool calls (@stevendcoffey). Cursor Rollouts: When Rollouts catches a regression, it finds the offending PR, opens an issue and offers a one-click cloud agent fix (@cursor_ai). Cloudflare: Sandbox SDK 1.0 gives Durable Objects direct control over sandbox containers (@CFchangelog). Cloudflare also launched request Traces (@WalshyDev). Research: Agent Training, Long-Horizon Control and AI for Math Multi-harness RL (Hugging Face): The same model weights score 62% in one harness and 33% in another (@huggingface). Method: A proxy speaks the OpenAI, Anthropic and Gemini API formats and records sampled token IDs and logprobs for training, with no changes to the harnesses themselves. Results: LFM2.5-2.6B improved from 42% to 54% across four harnesses and made 31% fewer tool calls. SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%. Release: The trainer, data and all seven trained models are open. Credit assignment and RL efficiency: ProVer has a judge locate the decisive trajectory segment, then uses rollouts on either side to set that segment’s advantage. It reports +9.91% (Qwen3.5-2B) and +7.12% (Qwen3.5-4B) relative gains over GRPO (@omarsar0). Partial rollouts: AC2 uses a learned critic to score token chunks, so training needs only partial rollouts (@wen_kaiyue). Frontier Learning: The method targets problems at the edge of capability, since problems a model always or never solves give zero GRPO gradient (@robinfaro13). Sharpening Tax: The paper quantifies the loss of pass@K scalability after post-training and proposes PTGS, a per-prompt temperature sampler (@iScienceLuvr). SFT vs RL: Another paper finds SFT generalizes worse because its data is off-policy, not because of the objective. Rewriting expert trajectories in the base model’s style closes the gap (@maximelabonne). Long-horizon control and context: Meta Superintelligence Labs reports that a dedicated controller lifts GPT-5.5 on ProgramBench from 63.7% to 71.5%, using the same workers and budget, versus 58.0% for Codex (@dair_ai). Context compression: Microsoft’s training-free FOCUS cuts peak context by up to 48% and raises task success by up to 8.9 points (@dair_ai). Long-context degradation: NVIDIA’s Long-Transduction study measures a 62.8% accuracy drop from 4K to 128K context across seven open models (@dair_ai). Multi-agent coordination: In AgentWorld, fewer than a third of multi-agent actions help complete the task, and coordination tasks reach only 12% success (@omarsar0). Apple LoopCD: The method halves recurrent loops while raising AIME 2024 pass@1 from 61.88% to 73.33% (@arankomatsuzaki). AI on open math problems: Meta released six papers on open problems produced with Muse Spark 1.1 and 1.2 through plain meta.ai chat, with no custom scaffold (@AIatMeta, list). Process: Each paper labels which passages were drafted primarily by humans or by AI, and a second group of mathematicians reviewed the work. Google Cogentic: This Gemini multi-agent system produced new results on five open theory problems (@omarsar0). Cogentic design: Each draft must pass two adversarial verifiers, and agents share a ledger of verified lemmas. Most problems took about 100 calls; the hardest took about 1,000. Image post-training: Arena combined a Bradley-Terry reward model with faithfulness, constraint and anti-reward-hacking rewards (@arena). Results: FLUX.2-dev gained 69 Elo to 1202, and Ideogram 4 gained 20 Elo to 1224. Benchmarks, Eval Integrity and Safety Research-taste benchmarks: ScholarCatalyst asks agents to find the “catalyst papers” behind research projects. It is labeled by 184 lead authors on 207 of their own projects and is described as far from saturated (@yoonholeee). EurekaBench: This benchmark tests whether agents can discover genuinely new insights across six science domains (@JiayiiGeng). Vals Web Search Index: The index holds model and harness constant, swaps only the search tool, and scores final answers on finance and legal tasks (@ValsAI). Validation: Agents score 2.9% (legal) and 7.4% (finance) without search, versus 30–50% with it. Vals also cites a study in which a model answered 44.5% of BrowseComp without search (details). SWE bug-finding bench: In this new benchmark, agents start from an older commit and are scored against real bugs fixed in later commits. Critique: Lucas Beyer argues it mainly tests recall and that the construction is easy to train toward (@giffmana). Authors’ response: The authors say training for bug-finding is fine as long as the test set is excluded (@OfirPress). Eval integrity question: David Rein asks whether Harbor, the framework behind Terminal Bench, lets agents modify their trajectories before evaluation. He notes he may be misreading the code (@idavidrein). Offensive capability of open models: The Batch reports GLM-5.3 nearly matched Claude Mythos at exploiting vulnerabilities, 12% vs 14% (@DeepLearningAI). Disputed claim: One commentator says GLM-5.3 Flash exceeds Mythos Preview on ExploitBench (@teortaxesTex). Uncensored variant: An uncensored GLM-5.3 is circulating on Hugging Face (@kimmonismus). Safety research and safeguards: A new paper proposes using internal signals during training to improve alignment without degrading white-box monitoring (@lenalibon). NeurIPS acceptance: “Models That Know How Evaluations Are Designed Score Safer” was accepted at NeurIPS 2026 (@HaritzPuerto). False positives: Opus 5.5 frequently triggers “reasoning extraction” safeguards during spectrogram syllable labeling (@ChaseBrowe32432). Emergent world knowledge: Asking a model “land or water?” for 16,200 lat/long coordinates and plotting the answers yields a recognizable world map (@karpathy). Inference, Hardware and Systems Ascend 950 via DeepSeek kernels: An analysis of DeepSeek’s open-sourced DeepGEMM, FlashMLA, TileKernels and DeepEP infers the chip’s layout (@ZhihuFrontier). Estimated specs: The chip has 32 AI cores, each pairing one Cube core with two Vector cores. Estimated peaks are about 432/865/1,730 TFLOPS in BF16/FP8/FP4. Capacity: Supply may be limited, despite claims that 950s went on sale in August (@teortaxesTex). Prime Inference: Prime Intellect stores the MLA latent in NVFP4, shrinking rows from 576 to 352 [truncated for AI cost control]
[AINews] not much happened today
Summary
a quiet day.
[AINews] not much happened today
Report an error
The correction channel is not available yet. You can copy the article reference below for later.
Correction instructions