AI News HubLIVE
站內改寫2 分鐘閱讀

待翻譯:AI tool claims to pick the top% of preprints. Should researchers trust

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:QED Science has developed an AI tool that examines manuscripts before publication and assesses the originality and validity of their findings.Credit: deepblue4you/Getty A growing number of private firms are offering res…

來源Hacker News AI作者: sbulaev

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

QED Science has developed an AI tool that examines manuscripts before publication and assesses the originality and validity of their findings.Credit: deepblue4you/Getty A growing number of private firms are offering researchers artificial-intelligence tools for scrutinizing manuscripts before publication. One such system, developed by a start-up called QED Science in Tel Aviv, Israel, aims to judge whether life-sciences research is original and valid. QED Science’s tool is trained to assess whether the claims in a given manuscript are supported by the data presented, and to identify gaps in the work. It is available to researchers for free, and has so far been used by more than 10,000 laboratories across 1,500 institutions in more than 70 countries. In November last year, openRxiv — the non-profit organization that operates the preprint servers bioRxiv and medRxiv — announced it would be piloting the QED Science system on bioRxiv. In an analysis posted in June, QED Science used the tool to rank more than 57,000 preprints that were posted on bioRxiv between May 2025 and April 2026, and selected the ‘top 1%’. But the ranking has sparked debate among researchers, with some arguing that it risks creating another badge of prestige — and reinforcing metric-based culture in academia. Others have expressed doubts about AI's ability to reliably and transparently judge the quality of scientific research. Nature spoke with Niv Mastboim, co-founder and chief executive of QED Science. Niv Masboim co-founded the AI company QED Science, which has developed a metric for measuring the quality of science manuscripts.Credit: Itay Rokban How does QED Science’s AI platform assess the science claims in a given article? There are a lot of tools for reviewing research articles. We go about it a bit differently, in a couple of aspects. We have trained the AI platform on multiple data sources, from open reviews and user feedback to synthetic data. One of the elements is establishing what would have been the negative results of the experiments — results that do not support the original hypothesis being tested, or the conclusion. The published literature disproportionately represents successful and positive findings. To evaluate scientific claims properly, a system also needs to understand what evidence that fails to support a claim looks like. This can be learnt from published null or contradictory findings, failed replication studies and other forms of non-supportive evidence. We focus on creating internal metrics for the system to constantly improve. So, the system validates each of the claims made in a paper and attaches a score to it. We use multiple scoring, from originality to validity. The AI platform is completely autonomous, but we also have a lot of users who give us feedback. We use that feedback to tune the tool constantly. What purpose did the top 1% list aim to serve? The goal was to judge science on the basis of the work itself, not the journal or the prestige of the authors. The 574 preprints in the 1% are the top-scoring bioRxiv preprints among the 57,455 assessed. They were selected solely according to QED’s assessment of originality and validity, independently of author identity, institution or publication venue. We also aimed to identify amazing papers that have been missed by the current publishing system. We did a separate validation analysis involving 2,879 bioRxiv preprints from April 2025 that were subsequently published in peer-reviewed journals, and then we benchmarked the ranking of our tool against journal ranking. In that analysis, QED rated 12.9% of these papers more highly than their eventual journals of publication might suggest. We call these articles ‘hidden gems’. We consulted a panel of experts to judge, in a blinded manner, the strongest cases of disagreement, and found that they preferred the QED-favoured paper in 75% of decisive comparisons. Is the tool the ultimate ‘peer reviewer’? Do you see it being used by publishers in the future? We are not offering our products to journals or publishers. Our goal is to provide free services to authors in a private, secure environment so they can improve the work before it is published.