AI News HubLIVE
In-site rewrite3 min read

Thomson Reuters launches proprietary AI model for legal work

Global content powerhouse Thomson Reuters Corp. today launched Thomson, its first proprietary large language model, combining the company’s trove of legal knowledge with LLMs from outside providers to provide legal advice. The company said Thomson will first be deployed in Tabular Analysis, a high-volume document review capability in its CoCounsel Legal AI assistant. CoCounsel will […] The post Thomson Reuters launches proprietary AI model for legal work appeared first on SiliconANGLE.

SourceSiliconANGLE AIAuthor: Paul Gillin

Global content powerhouse Thomson Reuters Corp. today launched Thomson, its first proprietary large language model, combining the company’s trove of legal knowledge with LLMs from outside providers to provide legal advice. The company said Thomson will first be deployed in Tabular Analysis, a high-volume document review capability in its CoCounsel Legal AI assistant. CoCounsel will remain a multimodel product, using Thomson for work where the company’s domain-specific model has an advantage and third-party frontier models for other tasks. Thomson will be the default model for Tabular Analysis, although administrators will be able to select other models. Thomson Reuters spent about $40 million over two years on people and computing for the project, but said economies reduced the cost of the final training run to about $450,000. Instead of building a foundation model from scratch, the company began with an open-weight model and added its proprietary content, training methods and professional expertise. It said the approach reduced both training and inference costs compared with general-purpose frontier models. The company isn’t seeking to compete with the largest AI labs across every field, said Joel Hron, global head of artificial intelligence and TR Labs at Thomson Reuters. “Thomson needs to set the frontier of intelligence for legal,” he said. “That’s a different job than what I think a lot of the frontier labs are doing.” Professional oversight The training process included realigning the base model with Thomson Reuters’ values, pretraining on the company’s content, targeted post-training guided by professionals and reinforcement learning that taught the model to work with company tools such as Westlaw and Practical Law. The firm’s flagship Westlaw platform encompasses over 40,000 individual databases and more than 150 years of legal publishing and editorial curation. Hron said hundreds of subject-matter experts helped define training objectives, create examples of legal questions and judge responses in blind comparisons. Specialization can damage a model’s broader abilities if it is handled poorly, said Jonathan Schwartz, head of foundational research at Thomson Reuters. The research team therefore focused on continual learning, or adding domain skills without erasing existing capabilities. “If you simply take the open-source model without any of these additional steps, it won’t know as much,” Schwartz said. “You won’t necessarily be aligned with your values, and it won’t be as good as using the tools that you’ve built later on.” Thomson Reuters said internal tests showed Thomson is broadly competitive with leading models when all had access only to the web. It moved to roughly equal or slightly better performance when connected to Thomson Reuters content, said Andrew Bean, a senior research scientist at the company. The tests assessed both the completeness of answers and whether citations supported their claims. A technical report to be published in the future is expected to provide additional benchmark results. Those results have not yet received extensive independent validation. Thomson Reuters has begun sharing the model with legal experts and academic institutions for testing and plans to release a smaller open-weight version on Hugging Face under a noncommercial academic license. It is also developing a portal through which outside developers can request application programming interface keys and test the model directly. More to come Only about 10% of the company’s total information base has been used so far, Bean said. The next step is not simply adding more material but turning the most useful content and product activity into better training signals. Ownership is also central to the company’s argument for sovereign AI. Thomson Reuters said customer data is not used to train the model and that controlling the model gives it more authority over deployment, governance and future development. It’s discussing direct model access with large law firms and corporations and is open to the possibility that customers could adapt Thomson to their own knowledge and workflows. Hron acknowledged that maintaining a proprietary model raises questions about whether Thomson Reuters can keep pace with faster-moving AI laboratories. He argued that improvements in open models will give the company stronger foundations for later versions, while its own investment can remain concentrated on professional work. “I don’t see owning an AI model that embodies the knowledge and expertise that TR possesses as something that’s non-core to what we do,” he said. “AI is a new mechanism for expertise delivery.” Photo: Flickr CC A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/ About SiliconANGLE Media