AI News HubLIVE
站内改写3 分钟阅读

待翻译:Rocky Linux Founder Kurtzer Launches OpenWALDO to Open Up AI Training Data

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Open weights have become something of a civil war in AI circles. True open-source AI? That continues to be a real rarity in AI circles. Gregory Kurtzer, the founder of Rocky Linux and a co-founder of CentOS, wants to ch…

来源Hacker News AI作者: CrankyBear

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Open weights have become something of a civil war in AI circles. True open-source AI? That continues to be a real rarity in AI circles. Gregory Kurtzer, the founder of Rocky Linux and a co-founder of CentOS, wants to change that with the release of OpenWALDO, a community-governed open-source project intended to create a shared, auditable corpus of AI training data and code. OpenWALDO, short for Open Weights, Artifacts, Licenses, Data and Origins, says it will make training data behave more like an open-source dependency: named, reviewable, versioned, attributable and verifiable. Backed by Kurtzer’s enterprise infrastructure company CIQ, it’s designed to address a persistent gap in the “open” AI movement. This brings together the principles behind open-weight and open-source models so that developers can access the code and methods, while the training data is open too, allowing anyone to inspect, verify and leverage that content, so everyone levels up, rather than just proprietary AI code companies. To do this, OpenWALDO will create an AI Bill of Materials (ABOM) that builds on that same transparency. This will start with the piece that has stayed closed the longest: the training data itself. “AI is not open source without the source,” the project says on its website. Its model centers on Git-based review for metadata, content-addressed storage for underlying data objects, and a bill of materials that follows resolved data and model lineage through training and release. OpenWALDO does not plan to build a foundation model. Instead, it wants to provide a common base layer: a licensed, provenance-tracked training corpus that model builders can use, extend with proprietary material, and carry into their own training runs. In the announcement, Kurtzer characterized that as shared infrastructure that would let organizations focus their resources on differentiation rather than repeatedly assembling basic training data. Why? Kurtzer explained in a statement, “I’ve spent my career watching open source turn users into builders, competitors into collaborators, and shared problems into common infrastructure that operates at massive scale. No single organization could build or sustain all of that alone. OpenWALDO brings that proven model to AI. Let’s work together, build its foundation in the open, and collaboratively take AI to the next level.” The project’s website says its public corpus currently indexes 124 billion reference tokens across 75.1 million documents, organized into 20 corpora and 1,051 shards, with 18 asserted license identifiers. OpenWALDO arrives as AI developers, researchers, and policymakers have increasingly focused on the provenance and licensing of training material. The MIT-led Data Provenance Initiative, for example, audited more than 1,800 text datasets and found license information was often omitted or miscategorized. MIT Sloan reported that the initiative found license-omission rates above 70% and license error rates greater than 50% in the data it examined. We can, and must, do better. OpenWALDO addresses this by proposing an AI bill of materials to record the relationship among data sources, licenses, corpus selections, training runs, and model releases. Such a record would allow a company to start from the community corpus, add data that it keeps private, and still retain an auditable path to the source materials used in a resulting model. The project is also framing data provenance as a quality issue. Its launch announcement cites concerns about recursive training on outputs from other AI systems, aka model collapse. OpenWALDO developers argue that a public record of sources can help developers distinguish and select training material. This effort adopts several practices familiar from software development: public review, attributable contributions, Developer Certificate of Origin sign-off and preserved provenance from data ingestion through release. OpenWALDO says contributors will include individuals, researchers, institutions and companies, while CIQ will support the project commercially rather than control it. “Linux didn’t win by being certified safe,” Kurtzer said in the release. “It won by being inspectable, forkable, and community validated. AI is missing that same property, and OpenWALDO is how we build it.” Whether OpenWALDO can become an important part of AI going forward will depend on more than its technical design. The project must attract a sustained supply of high-quality, legally usable data; establish trusted review and governance processes; and persuade AI developers that its provenance records provide enough practical value to become part of their model-development workflows. I hope it does. TECHSTRONG AI PODCAST SHARE THIS STORY