Skip to content
AI News HubLIVE
Public articles 17Collected articles 18Trust 87Refresh 720 min
Health HealthySource type ResearchFull-text rights Full text allowedLast ingested 2026-09-29ID ahead-of-aiStatus Enabled

Public Substack newsletter; free posts allowed.

Latest public articles

GPT-6 Astra, Looped Transformers, and Hidden Reasoning

A Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks

Ahead of AI (Sebastian Raschka)In-site articleGPT-6 Astra, Looped Transformers, and Hidden Reasoning

How Claude Watermarks AI-Generated Text

A 48-minute video walkthrough of token sampling, watermark detection, and removal

Ahead of AI (Sebastian Raschka)In-site articleHow Claude Watermarks AI-Generated Text

Building an AI Text Detector From Scratch

This article walks through building an AI text detector end-to-end: dataset construction, model training, local deployment, and using it as a verifier to train an SLM via RLVR. Inspired by Substack's AI detector, it explains how detectors work, their limitations, and how they can help writers avoid over-polished, AI-sounding text.

Ahead of AI (Sebastian Raschka)In-site articleBuilding an AI Text Detector From Scratch

Controlling Reasoning Effort in LLMs

This article explores how to develop reasoning models with multiple effort modes, covering the evolution from o1 and DeepSeek-R1 to GPT-5.6, and key techniques such as RLVR training, inference scaling, think tokens, and reasoning mode toggles.

Ahead of AI (Sebastian Raschka)In-site articleControlling Reasoning Effort in LLMs

Using Local Coding Agents

A tutorial on setting up a production-ready local coding agent using open-source tools and open-weight LLMs, with a focus on Qwen3.6 and Qwen-Code harness, covering motivation, setup, performance assessment, and alternatives.

Ahead of AI (Sebastian Raschka)In-site articleUsing Local Coding Agents

LLM Research Papers: The 2026 List (January to May)

The author continues an annual tradition by curating and categorizing notable LLM research papers from January to May 2026, covering architecture, training, inference efficiency, reasoning, reinforcement learning, agent systems, and more, with emphasis on hybrid architecture trends and representative works like Nemotron 3.

Ahead of AI (Sebastian Raschka)In-site articleLLM Research Papers: The 2026 List (January to May)

Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

From Gemma 4 to DeepSeek V4, this article explores how new open-weight LLMs are reducing long-context costs through architectures like cross-layer KV sharing, per-layer embeddings, attention budgeting, compressed convolutional attention, and mHC.

Ahead of AI (Sebastian Raschka)In-site articleRecent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

My Workflow for Understanding LLM Architectures

A learning-oriented workflow for understanding new open-weight model releases, starting with official reports but relying more on config files and reference code due to less detailed papers.

Ahead of AI (Sebastian Raschka)In-site articleMy Workflow for Understanding LLM Architectures

Components of A Coding Agent

An overview of the six core components that make coding agents effective, including live repo context, prompt caching, tool use, context management, session memory, and subagent delegation, explaining how these harness features enhance LLM performance in coding tasks.

Ahead of AI (Sebastian Raschka)In-site articleComponents of A Coding Agent

A Dream of Spring for Open-Weight LLMs: 10 Architectures from Jan-Feb 2026

A comprehensive review and comparison of ten open-weight large language model releases from January to February 2026, including Arcee Trinity, Kimi K2.5, Step 3.5 Flash, Qwen3-Coder-Next, GLM-5, MiniMax M2.5, Nanbeige 4.1, Qwen3.5, Ling 2.5, Tiny Aya, and an update on Sarvam. The article focuses on architectural similarities and differences, highlighting trends like hybrid attention, multi-token prediction, and mixture-of-experts.

Ahead of AI (Sebastian Raschka)In-site articleA Dream of Spring for Open-Weight LLMs: 10 Architectures from Jan-Feb 2026

Categories of Inference-Time Scaling for Improved LLM Reasoning

Inference-time scaling is one of the most effective ways to improve answer quality in deployed LLMs. This article categorizes various inference-time scaling techniques and provides an overview of recent papers, including chain-of-thought prompting, self-consistency, best-of-N ranking, rejection sampling with a verifier, self-refinement, and search over solution paths. The author shares personal experiments from drafting a book chapter on the topic.

Ahead of AI (Sebastian Raschka)In-site articleCategories of Inference-Time Scaling for Improved LLM Reasoning

The State Of LLMs 2025: Progress, Problems, and Predictions

A comprehensive review of large language models in 2025, covering key developments like DeepSeek R1's reasoning via RLVR/GRPO, the rise of inference-time scaling and tool use, the problem of benchmark overfitting (benchmaxxing), and predictions for 2026 including diffusion models and broader RLVR applications.

Ahead of AI (Sebastian Raschka)In-site articleThe State Of LLMs 2025: Progress, Problems, and Predictions

LLM Research Papers: The 2025 List (July to December)

The author shares a curated list of interesting research papers from July to December 2025, categorized by topics like reasoning models, reinforcement learning, and architectures, as a thank-you to supporters.

Ahead of AI (Sebastian Raschka)In-site articleLLM Research Papers: The 2025 List (July to December)

From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates

This article provides an in-depth analysis of DeepSeek V3.2's technical evolution, covering architectural changes (including the sparse attention mechanism DSA), reinforcement learning updates (such as GRPO improvements, self-verification and self-refinement), and the development of hybrid reasoning models. V3.2 matches the performance of GPT-5 and Gemini 3.0 Pro and is released as an open-weight model, making it a significant milestone.

Ahead of AI (Sebastian Raschka)In-site articleFrom DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates

Beyond Standard LLMs

This article explores alternatives to standard autoregressive decoder-style transformers for large language models, including linear attention hybrids, text diffusion models, code world models, and small recursive transformers. It analyzes the strengths and limitations of each approach and discusses their potential impact on efficiency, reasoning, and modeling performance.

Ahead of AI (Sebastian Raschka)In-site articleBeyond Standard LLMs

Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)

This article explains four main approaches to evaluating large language models: multiple-choice benchmarks (like MMLU), verifiers for free-form answers, leaderboards based on user preferences (like Chatbot Arena), and LLM-as-a-judge evaluations. It includes from-scratch code implementations and discusses the trade-offs of each method.

Ahead of AI (Sebastian Raschka)In-site articleUnderstanding the 4 Main Approaches to LLM Evaluation (From Scratch)

All sources