Language Models for Text Classification: From Bag-of-Words to Jev
A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency
Source profile
AI News Hub tracks Ahead of AI (Sebastian Raschka) AI updates with visible source status, reuse boundaries, collection method, and published articles.
Public Substack newsletter; free posts allowed.
A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency
A Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks
A 48-minute video walkthrough of token sampling, watermark detection, and removal
This article walks through building an AI text detector end-to-end: dataset construction, model training, local deployment, and using it as a verifier to train an SLM via RLVR. Inspired by Substack's AI detector, it explains how detectors work, their limitations, and how they can help writers avoid over-polished, AI-sounding text.
This article explores how to develop reasoning models with multiple effort modes, covering the evolution from o1 and DeepSeek-R1 to GPT-5.6, and key techniques such as RLVR training, inference scaling, think tokens, and reasoning mode toggles.
A tutorial on setting up a production-ready local coding agent using open-source tools and open-weight LLMs, with a focus on Qwen3.6 and Qwen-Code harness, covering motivation, setup, performance assessment, and alternatives.
The author continues an annual tradition by curating and categorizing notable LLM research papers from January to May 2026, covering architecture, training, inference efficiency, reasoning, reinforcement learning, agent systems, and more, with emphasis on hybrid architecture trends and representative works like Nemotron 3.
From Gemma 4 to DeepSeek V4, this article explores how new open-weight LLMs are reducing long-context costs through architectures like cross-layer KV sharing, per-layer embeddings, attention budgeting, compressed convolutional attention, and mHC.
A learning-oriented workflow for understanding new open-weight model releases, starting with official reports but relying more on config files and reference code due to less detailed papers.
An overview of the six core components that make coding agents effective, including live repo context, prompt caching, tool use, context management, session memory, and subagent delegation, explaining how these harness features enhance LLM performance in coding tasks.
A comprehensive review and comparison of ten open-weight large language model releases from January to February 2026, including Arcee Trinity, Kimi K2.5, Step 3.5 Flash, Qwen3-Coder-Next, GLM-5, MiniMax M2.5, Nanbeige 4.1, Qwen3.5, Ling 2.5, Tiny Aya, and an update on Sarvam. The article focuses on architectural similarities and differences, highlighting trends like hybrid attention, multi-token prediction, and mixture-of-experts.
Inference-time scaling is one of the most effective ways to improve answer quality in deployed LLMs. This article categorizes various inference-time scaling techniques and provides an overview of recent papers, including chain-of-thought prompting, self-consistency, best-of-N ranking, rejection sampling with a verifier, self-refinement, and search over solution paths. The author shares personal experiments from drafting a book chapter on the topic.
A comprehensive review of large language models in 2025, covering key developments like DeepSeek R1's reasoning via RLVR/GRPO, the rise of inference-time scaling and tool use, the problem of benchmark overfitting (benchmaxxing), and predictions for 2026 including diffusion models and broader RLVR applications.
The author shares a curated list of interesting research papers from July to December 2025, categorized by topics like reasoning models, reinforcement learning, and architectures, as a thank-you to supporters.
This article provides an in-depth analysis of DeepSeek V3.2's technical evolution, covering architectural changes (including the sparse attention mechanism DSA), reinforcement learning updates (such as GRPO improvements, self-verification and self-refinement), and the development of hybrid reasoning models. V3.2 matches the performance of GPT-5 and Gemini 3.0 Pro and is released as an open-weight model, making it a significant milestone.
This article explores alternatives to standard autoregressive decoder-style transformers for large language models, including linear attention hybrids, text diffusion models, code world models, and small recursive transformers. It analyzes the strengths and limitations of each approach and discusses their potential impact on efficiency, reasoning, and modeling performance.
This article explains four main approaches to evaluating large language models: multiple-choice benchmarks (like MMLU), verifiers for free-form answers, leaderboards based on user preferences (like Chatbot Arena), and LLM-as-a-judge evaluations. It includes from-scratch code implementations and discusses the trade-offs of each method.