Traj-Evolve is a self-evolving multi-agent system for patient trajectory modeling from longitudinal EHRs. It uses an Experience Pool (ExPool) for non-parametric memory and multi-agent reinforcement learning (MARL) for parametric optimization. On a lung cancer prediction task, it outperforms 9 baselines, with ExPool improving specificity and MARL improving sensitivity.
Live AI News Intelligence
Live monitoring
Live updates
Trusted sources, attribution, rights, and in-site reading distilled into a signal-first AI brief.
Live updates
ChatHealthAI is a multimodal reasoning framework that aligns structured EHR representations from a pretrained EHR foundation model with the semantic space of a frozen LLM via a task-aware resampler. It integrates longitudinal patient data with refined clinical event descriptions to enable interpretable natural-language reasoning while maintaining predictive performance. Evaluated on three tasks from the EHRSHOT benchmark, ChatHealthAI improves reasoning quality and interpretability without sacrificing accuracy.
A new study compares encoder-only Transformer and LSTM for upstream streamflow inference in ungauged basins using NOAA National Water Model simulations. LSTM outperformed Transformer overall, and incorporating downstream information boosted median NNSE by over 60%. The findings highlight the importance of architectural inductive bias.
AURA-Mem proposes a constant-size recurrent memory for robot policies that writes only when an observation would change the next action, drastically reducing memory writes while maintaining accuracy. It uses a learned gate trained on action-error signal, achieving fixed 4,224-byte inference state vs growing KV-cache. Experiments show matching success rates with up to 7x fewer writes.
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.
Dropstone 1.5 is an AI coding agent for the terminal, offering roughly 450 deep coding sessions per week for $15/month—about twice what Claude Code Pro delivers for $20. It runs on DeepSeek and Kimi models hosted in the US, with no data stored. Safety features require permission for file writes, shell commands, and network calls.
A new study examines how AI coding agents differ from human programmers in consuming error messages. Through controlled experiments, it finds that more detailed type error messages significantly improve an agent's ability to fix errors, and that the presence of a type system is more helpful than test suite failure reports alone.
A practical look at four essential signals for AI observability: versioned prompts, detailed traces, user feedback, and model-based evaluation.
Sydney Morning Herald removes piece by Cath Ellis, despite Western Sydney University saying her use of AI was ‘appropriate’
Basic extractors often return broken lines. This pdf to md converter uses layout understanding so headings, tables, captions, and images stay clearer in the output.
South Korea’s Kospi stock market has hit record highs thanks to AI, but experts urge caution over boom-bust cycles and a heavy reliance on two chipmakers.
LiteHarness is a unified SDK that provides a single TypeScript and Python interface for multiple AI agent harnesses, including Claude Agent SDK and OpenAI Agents SDK. It allows easy switching between harnesses and models, and supports streaming messages. The project is in preview.
The article about leaving the prestigious Alan Turing Institute to join a groundbreaking AI lab.
At Code w/ Claude SF 2026, Director of Engineering for Claude Code and Claude Cowork Fiona Fung walked through how the team’s processes and structure changed once agentic coding became the default way of working.
Replicas launches on Product Hunt, enabling developers to run coding agents like Claude Code and Codex in isolated cloud VMs. Integrated with Slack, Linear, and GitHub, it allows background agent execution with real dev environments. Founders Connor and Saai share their vision and early traction.
Learn to fine-tune LFM2 with QLoRA, supervised fine-tuning, DPO, and adapter merging using TRL and PEFT on Colab.
AI search summaries flatten nuance, undermine verification, and harm the web ecosystem. This article argues for results-only search to preserve accuracy and critical thinking.
Instagram resolved a security issue that allowed several users’ accounts to get hacked by tricking its AI-powered support chatbot into granting access.
Google is secretly offering to pay Android app developers for access to their codebases under a confidential pilot program, aiming to improve its AI coding tools. Developers retain intellectual property rights, but Google seeks to catch up with competitors like Anthropic and Microsoft in AI code generation. The move highlights the growing scarcity of public training data.
Brontosaurus is a web-based generative canvas that uses voice commands to create widgets almost instantly. Inspired by Thinking Machines and Ink & Switch, it emphasizes human-AI collaboration, prioritizing speed so users can bring ideas to life as fast as they can speak.
Weaviate announces the general availability of Engram, a managed memory and context service for agentic applications. It addresses long-context degradation, messy raw data, and multi-agent context fragmentation through asynchronous pipelines, templates, and built-in scopes, helping agents compound value over time.
Reachy Mini can now use remote tools hosted in Hugging Face Spaces via MCP, allowing it to check weather or search the web with a single command. The article covers built-in tools, profile-based control, tool installation, naming conventions, and current limitations.
Anthropic's 'On the Biology of a Large Language Model' (2025) is a landmark in mechanistic interpretability. Using circuit tracing, researchers reveal multi-step reasoning inside models, showing they use human-interpretable concepts like 'Texas' for pseudo-symbolic inference. This work helps identify misbehavior, steer models, and design better algorithms.
This documentary examines the potentially disruptive impact of AI on the internet, including disinformation, algorithmic manipulation, and changes to the online ecosystem.
Project Brain is a Claude Code skill that creates a lightweight, navigable memory (.project-brain/ folder) for each project, recording stack, decisions, pitfalls, and history, eliminating the need to re-explain the project every session and reducing token usage and hallucinations.
Titan Network aggregates unused computing power from consumers' connected devices into a decentralized cloud, offering AI firms infrastructure at up to 75% lower cost. Clients include Tencent, Alibaba, and Kling AI. The company pays 80% of revenue from data tasks to individuals who share their devices and bandwidth.
ContextWall is an open-source context firewall that intercepts and scans documents before they enter an AI model's context window, preventing prompt injection, credential leaks, and PII exfiltration. It requires no code changes to agents, runs in your infrastructure, and offers three detection layers with source trust tiers.
Scholar Sidekick is a free tool that generates formatted citations from DOIs, PubMed IDs, arXiv IDs, and more. It also verifies citations, checks open access, and retraction status. It offers a REST API and MCP server for developers and AI agents.
Microsoft announced two new text LLMs: MAI-Thinking-1 (35B parameters, reasoning) and MAI-Code-1-Flash (5B parameters, code). Both are trained on clean, licensed data without distillation, with MAI-Thinking-1 claiming preference over Sonnet 4.6. MAI-Code-1-Flash is rolling out to GitHub Copilot users in VS Code.
Hermes Desktop is an AI agent that grows with you. Currently featured on Product Hunt with discussion.