本文にスキップ
AI News HubLIVE

ソース分布

  • Hacker News AI10
  • arXiv Computational Linguistics7
  • TheSequence7
  • arXiv AI6
  • MarkTechPost4
  • arXiv Machine Learning3
  • Simon Willison's Weblog2
  • The New Stack AI2

トピック分布

  • モデル38
  • 研究37
  • Agent36
  • スタートアップ4
  • チップ2
  • 政策1

タイムライン

  • 2026-09-186
  • 2026-08-183
  • 2026-08-263
  • 2026-08-132
  • 2026-08-172
  • 2026-08-212
  • 2026-09-012
  • 2026-09-102

最新動向

翻訳待ち:DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek has released official macOS and Windows desktop apps for DeepSeek Harness v0.2, its MIT-licensed agent harness. The preview adds a plugin manager, file and code-change review, and scheduled Automation Tasks. It also supports non-DeepSeek models through OpenAI-compatible endpoints. The post DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness

翻訳待ち:LWiAI Podcast #258 - Opus 5.5, Sol and Luna, Muse, DeepSeek-V4.1-Flash, Xi

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:(Belated post :/ ) Anthropic releases Opus 5.5 with lower prices and Fable-level performance, OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes

Last Week in AI原典の内容 · 翻訳・分析待ち翻訳待ち:LWiAI Podcast #258 - Opus 5.5, Sol and Luna, Muse, DeepSeek-V4.1-Flash, Xi

翻訳待ち:Activation-Conditioned Self-Distillation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.38342v1 Announce Type: new Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verif…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:Activation-Conditioned Self-Distillation

翻訳待ち:What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.00018v1 Announce Type: new Abstract: Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface. We introduce a message-intervention diagnostic that fixes the evidence and candidate answer while varying only the rationale passed across the reasoner-to-verifier boundary. On 400 MuSiQue, HotpotQA, and 2WikiMultiHopQA examples with DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy over no rationale, while corrupted rationales strongly alter support judgments. Under a blind verifier prompt, harmless paraphrases shift support by only 0--2.5%, whereas corrupte…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA

翻訳待ち:The Sequence Learning Loop - Issue 942: Learning About Opus 5.5, DeepSeek’s Training Grounds, and Claude’s DNA Discovery

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Cheaper agents, environments at massive scale, and a biological discovery reveal how AI turns capability into useful work.

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Learning Loop - Issue 942: Learning About Opus 5.5, DeepSeek’s Training Grounds, and Claude’s DNA Discovery

翻訳待ち:Will Chinese AI companies slow down? A top House Democrat wants answers

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Image: Tom Williams/CQ-Roll Call, Inc via Getty Images As President Donald Trump prepares to meet tech and AI CEOs in Washington, Rep. Ro Khanna (D-CA) is calling for a treaty between the US and China to keep AI from wreaking havoc on the world. But wrangling leaders in both countries to take action could be a long shot. In letters shared exclusively with The Verge, Khanna - the top Democrat on the House Select Committee on China and a possible presidential contender - asked a US intelligence agency and Chinese AI companies whether they would be prepared for an incident like OpenAI agents' hack of Hugging Face this summer. The companies included DeepSeek, Alibaba, and Moonshot AI, all of which … Read the full story at The Verge.

The Verge AI原典の内容 · 翻訳・分析待ち翻訳待ち:Will Chinese AI companies slow down? A top House Democrat wants answers

翻訳待ち:AI Sovereignty: Bargaining with Big Tech and the Promise of Full Stack Open Source AI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The early rapid expansion of AI capabilities that focused on frontier models was largely ushered into the world by a few powerful, US-based AI labs. Open-weight models released from labs in China, early on from DeepSeek, and later from Moonshot, Z.ai, and others, have in part disrupted that dominance. But growing concerns about the concentration […]

O'Reilly AI & ML Radar原典の内容 · 翻訳・分析待ち翻訳待ち:AI Sovereignty: Bargaining with Big Tech and the Promise of Full Stack Open Source AI

翻訳待ち:Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Jina AI has released jina-ocr-v1, a visual document parser that converts PDFs, scans, tables, charts and invoices into Markdown. The model has 3.4B total parameters, with about 570M active per token, and builds on DeepSeek-OCR. A built-in FastMTP speculative decoding head drafts 3 tokens per step while keeping output lossless. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and parses 2.57 pages per second on 1 A100. Weights are on Hugging Face under CC BY-NC 4.0, with hosted access through Jina Reader. The post Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

翻訳待ち:Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.19156v1 Announce Type: new Abstract: Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning trajectories suffer from a Scaling Collapse: when the problem set is limited, increasing positive examples fails to yield continuous improvement. However, during inference, an LLM can not guarantee that every intermediate step is correct and is therefore prone to errors. Once such errors arise, the LLM often struggles to recover and may be further misled by the accumulation of previous mistakes. To address this, we propose Reflective Recovery, a simple yet effective self-supervised a…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

翻訳待ち:Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.19154v1 Announce Type: new Abstract: While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. To address this issue, we introduce Neo-Classic, an evaluation benchmark that combines a constructionist Out-of-Sample (OOS) dataset with a suite of reverse understanding probes. Unlike traditional benchmarks that rely on verification or generation over historical corpora, Neo-Classic comprises strictly metrical poetry authored by contemporary experts, reducing the possibility of direct retrieval. We evaluate state-of-the-art models, including Qwen3-Max, Gemini-3-Pro…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

翻訳待ち:What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.19151v1 Announce Type: new Abstract: Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applications (ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity), comprising 17,012 English-language reviews from Google Play and the Apple App Store. We combine BERTopic topic modeling with RoBERTa sentiment classification and evaluate cross-application differences using chi-square, Kruskal-Wallis, and multinomial logistic regression with Bonferroni correction. Both components are validated against human coding using a str…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

翻訳待ち:Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.19150v1 Announce Type: new Abstract: Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative: we repeatedly sample completions of a single prompt at elevated temperature, apply Principal Component Analysis (PCA) to the pooled hidden activations, and label the resulting axes automatically from the pole generations. We validate the discovered axes against 245 human-elicited stylistic annotations in a two-phase study. On our strongest model (Qwen-3.5-4B-Instruct), the top two axes match spontaneously requested human dimensions wit…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations

翻訳待ち:Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.19244v1 Announce Type: new Abstract: Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

翻訳待ち:BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.19180v1 Announce Type: new Abstract: Language models face unique challenges in analyzing interdisciplinary scientific research literature. In biophysics research, faithful answers require grounding observed data in source evidence, interpreting it through a quantitative physics model, and linking it to a biological mechanism. To address this challenge, we introduce BioPhys-Bridge, a novel benchmark dataset for evidence-grounded scientific reasoning over biophysical literature. Each case contains evidence blocks, stable evidence IDs, quantitative values, units, equations, assumptions, mechanisms, and next decisions as grounding targets for question answering (QA) and retrieval-augmented generation (RAG). The initial release contains 500 ca…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

翻訳待ち:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.17538v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for structured information extraction from documents, yet their behavior under realistic OCR noise remains poorly understood. We present a systematic benchmark of open-source instruction-tuned LLMs for key-value pair (KVP) extraction under both clean-text and noisy OCR conditions. We evaluate representative decoder-only models (Gemma, Mistral, Qwen2.5, LLaMA 3, and DeepSeek) on the FUNSD, CORD, and SROIE benchmarks using both Gold-text annotations and OCR outputs from PaddleOCR, EasyOCR, and Tesseract. A unified evaluation protocol isolates the effects of input quality, model design, and prompting under consistent conditions. The results show that mode…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

翻訳待ち:The Sequence Learning Loop - Issue 934: Understanding DeepSeek V4.1 Flash, DeepMind’s AlphaGenome Atlas and Muse

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:How architecture and system design turn model capability into useful work

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Learning Loop - Issue 934: Understanding DeepSeek V4.1 Flash, DeepMind’s AlphaGenome Atlas and Muse

翻訳待ち:Self-reported archetypes and behavioral failures in Large Language Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.15998v1 Announce Type: new Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet the structure of LLM character remains poorly understood. We map the self-reported personality archetypes of 22 LLMs spanning closed-source frontier systems (GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6) and open-source models (Llama, DeepSeek, OLMo, and Qwen series). Each model self-rated across 464 bipolar semantic-differential trait pairs, and the resulting profiles were projected into a six-dimensional archetypal…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Self-reported archetypes and behavioral failures in Large Language Models

翻訳待ち:How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:On 25 August, OpenAI fully unveiled Jalapeño, the company’s debut AI accelerator chip. Jalapeño delivers up to 13.4 petaflops of 4-bit compute and accesses 232 gigabytes of the most advanced memory available, linking to it at a blazing 15.4 terabytes per second. Benchmarks cited by OpenAI show that Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6x when compared to Nvidia’s GB300—a chip the company currently relies on—and do so while consuming less power. Whether these figures translate into real-world gains once Jalapeño enters widespread service in OpenAI’s inference fleet remains to be seen, but performance is only half the story. The other half is how the chip was designed—a process which, as you might e…

IEEE Spectrum AI原典の内容 · 翻訳・分析待ち翻訳待ち:How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

翻訳待ち:Why DeepSeek-V4.1-Flash Is Such an Exciting Open Model Release

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek-V4.1-Flash shows how Causal Encoder-Decoder architecture, MoE, KV cache compression, CSA2, cheaper prefill, and efficient decoding can make powerful open-source AI models far more efficient to run.

KDnuggets原典の内容 · 翻訳・分析待ち翻訳待ち:Why DeepSeek-V4.1-Flash Is Such an Exciting Open Model Release

翻訳待ち:The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Faster models, nine billion genetic predictions, personal agents, and a proof that could make mathematical history.

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough

翻訳待ち:[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We agree with Sebastian: this should have been DeepSeek v5

Latent Space原典の内容 · 翻訳・分析待ち翻訳待ち:[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale

翻訳待ち:Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

翻訳待ち:Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09448v1 Announce Type: new Abstract: As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

翻訳待ち:DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. DeepSeek AI built its newest release around that exact bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters, 196B additional Engram parameters, and a 1M-token context window. It […] The post DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

翻訳待ち:Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05658v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. This paper presents a controlled metamorphic evaluation of LLM-based SVA generation under semantics-preserving RTL transformations. Starting from the VERT dataset, we construct a quality-filtered conditional-control pool and a stratified 40-program evaluation set containing 295 assignment behaviors. We evaluate two open code models, Qwen2.5-Coder-7B and DeepSeek-Coder-V2…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

命令チューニングされた小型言語モデルによる高齢者向け段階的金銭詐欺のインクリメンタルリスク評価

高齢者を標的とした金融詐欺は、電子メールやSMS、電話など複数ターンの会話を通じて行われることが増えており、リスクシグナルは各ターンで徐々に現れます。そのため、リソース制約のある環境で継続的にリスクを更新できるモデルが求められています。本研究では、累積ターンベースのリスク評価フレームワークを提案し、投資詐欺・慈善詐欺・テクニカルサポート詐欺のシナリオを含む多ターン対話データセットを構築しました。Phi-4、LLaMA-3.2、DeepSeek-R1、Qwen3の4つの小型言語モデルを微調整した結果、Phi-4とLLaMA-3.2がパラメータ規模に対して優れたターン認識リスク推定性能を示し、プライバシー保護とオンデバイス詐欺防止の可能性を示唆しています。

arXiv AIサイト内本文命令チューニングされた小型言語モデルによる高齢者向け段階的金銭詐欺のインクリメンタルリスク評価

翻訳待ち:The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:From DistilBERT and Gemini Flash to Gemma, Llama, Qwen, DeepSeek, Phi, Ministral, and PrismML’s Bonsai 27B.

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

翻訳待ち:The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28859v1 Announce Type: new Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts th…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

翻訳待ち:DeepSeek-V4-Flash-Vision-Exp

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"pad_token":{"__type":"AddedTok…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek-V4-Flash-Vision-Exp

翻訳待ち:Subs, the cloud native agent harness

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Run a local agent. Create a subs.toml. subs.toml name = "example" [llm.openrouter] type = "openrouter" [agent.teammate] llm = "openrouter" model = "deepseek/deepseek-v4-flash-0731" system = "You are a helpful teammate."…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:Subs, the cloud native agent harness

翻訳待ち:Just a rumour of a bug is enough to find a security exploit these days

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Just a rumour of a bug is enough to find a security exploit these days Anil Madhavapeddy is a professor of computer science at Cambridge and a core maintainer of the OCaml compiler. In this somewhat alarming post he reports that security issues in OCaml projects are seeing evidence of attempted exploits within minutes of patches being shared for discussion: This normally takes a few days and a release within a week or two is reasonable. Within about ten minutes (!) this website was fielding probes for percent-encoded traversal sequences, indicating that automated watchers are keeping an eye on public repositories. Modern coding agents have become so effective at finding flaws that the slightest hint at a new bug can be enough information for them t…

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:Just a rumour of a bug is enough to find a security exploit these days

翻訳待ち:Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled appeared first on The New Stack.

The New Stack AI原典の内容 · 翻訳・分析待ち翻訳待ち:Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

翻訳待ち:The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Distilling three major AI releases to keep you current.

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched

翻訳待ち:AutoRouter – Enterprise AI gateway for every model

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The AI API Platform Built for the Agent EraUnleash unlimitedAI capabilities From Seedance 2.0 and Kling 3.0 to GPT-5.5, Claude Opus 4.7, Gemini 3.1, DeepSeek V4, and Qwen 3.6, connect to the world's leading AI models th…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:AutoRouter – Enterprise AI gateway for every model

翻訳待ち:Kraftapp AI – Describe it. We build it. Customers find it

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bring the intent. Agents do the rest, inclu…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:Kraftapp AI – Describe it. We build it. Customers find it

翻訳待ち:DeepSeek V4 Flash Vision Intelligence, Performance and Price Analysis

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Artificial Analysis DeepSeek • DeepSeek V4 Flash 0731 • Proprietary model • Released August 2026 DeepSeek V4 Flash Vision (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Benchmarks Model…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek V4 Flash Vision Intelligence, Performance and Price Analysis

翻訳待ち:The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:OpenRouter, Ramp, Etched, and DeepSeek reveal the emerging economic stack beneath modern intelligence.

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy

翻訳待ち:DeepSeek debuts multimodal language model competitive with Opus 4.8

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer platform. The company may release a free version later on given that it has open-sourced many of its earlier models. Those models include V4 Flash, […] The post DeepSeek debuts multimodal language model competitive with Opus 4.8 appeared first on SiliconANGLE.

SiliconANGLE AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek debuts multimodal language model competitive with Opus 4.8

翻訳待ち:We burned 11.7B tokens to find the best cyber AI model

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We burned 11.7bn tokens to find the best cyber AI model GLM5.3 and DeepSeek are now frontier-tier models Debarshi Philippe Dourassov Published on: Aug 21, 2026 We burned 11.7 billion tokens to benchmark the cyber capabi…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:We burned 11.7B tokens to find the best cyber AI model

翻訳待ち:DeepSeek-V4-Flash-Vision-Exp Is Now Live on the DeepSeek API Platform

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Post Log inSign up Post DeepSeek on X: "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, re…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek-V4-Flash-Vision-Exp Is Now Live on the DeepSeek API Platform

翻訳待ち:Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered t…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions

翻訳待ち:The Sequence Frontier Learning - Issue 917: Understanding DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:A mini deep dive into some of the most important AI releases of last week.

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Frontier Learning - Issue 917: Understanding DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard

翻訳待ち:DeepSeek Harness

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 16.5k Star 158k BranchesTags Open more actions menu Latest commit History 12,404…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek Harness

翻訳待ち:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

Together AI Blog原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

翻訳待ち:Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.6B parameters, and Luna is size unknown but presumably a whole lot bigger than 27B. Qwen 3.8 27B is a truly astonishing model. Via Hacker News Tags: ai, generative-ai, llms, qwen, ai-in-china, artificial-analysis

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

翻訳待ち:DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. Four runtime modes, append-only session logs, and provider-agnostic model routing. The post DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

翻訳待ち:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

Together AI Blog原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

翻訳待ち:DeepSeek open sources an agent harness where everything is a plugin

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek on Thursday open sourced the DeepSeek Harness, a new agent runtime for developers. The Node.js-based harness is now available The post DeepSeek open sources an agent harness where everything is a plugin appeared first on The New Stack.

The New Stack AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek open sources an agent harness where everything is a plugin

翻訳待ち:DeepSeek v4 Price Increase

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek (@deepseek_ai): "API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek v4 Price Increase

翻訳待ち:DeepSeek-AI/DeepSeek-V4-Pro-0813

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"pad_token":{"__type":"AddedTok…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:DeepSeek-AI/DeepSeek-V4-Pro-0813

企業ナビゲーション