Identity crises?

The rumour that AI companies are bulk buying books and destroying them so that their content can be scanned in to train their AI models has been verified twice: by court documents in California and by journalists at 404 Media who hid an Airtag in a rare book sold to a bulk purchaser. The book was tracked to an Amazon facility near Las Vegas, where books are dismantled, scanned and discarded. Amazon confirmed that it uses older books to develop and improve its products, i.e. for AI training, not for reading. Following 404 Media’s report, book database ISBNdb took down the web page which was advertising a printed books sourcing service for LLM dataset needs. But booksellers worldwide are still seeing suspiciously large orders for second-hand books. As Kathryn James explained in The Guardian, this is to get around the permissions needed to reuse digital content and to provide clean training data from books published before the launch of ChatGPT in 2022 which are less likely to contain AI-generated material.

Source code

Ironically, given that the scale of Anthropic’s ‘Project Panama’ which involved similar destructive scanning of pre-2022 books to train Claude only emerged from unsealed court filings in relation to the northern California district court copyright case Bartz v Anthropic which was finalised in late July, August saw the introduction of Claude Text Watermarking. How does this work? Claude embeds statistical text watermarks during autoregressive token generation so that the watermark is deeply embedded in the text and doesn’t disappear if you copy and paste. However, it is relatively straightforward to get around watermarked content: you can run the text through another process, like translation or summarisation, or if it’s not a massive document, you can simply retype it! But obviously there are risks involved, especially in legal, where there have been high profile instances of lawyers and litigants using unedited AI content.

Legal-specific models

One solution is a legal specific GenAI model. Last week Harvey released a research preview of Harvey Tenet, which is built on Chinese open-weight model Kimi K3 and post-trained on legal data and practice. As Ryan McDonough, head of software engineering at KPMG Law explained on LinkedIn “The model gets matter documents, tools and realistic legal tasks with its work assessed against detailed expert rubrics… to encode how legal work should be performed, not simply what the law says.” Playbooks have been around for a while, and this new model usefully incorporates a playbook for legal work. McDonough also highlighted Harvey’s intelligence per token benchmark.

Thomson Reuters launched Thomson which is based on Chinese OSS model Qwen and trained on its own data, making CoCounsel’s Tabular Analysis more cost-effective. Mike OSS founder Will Chen commented on LinkedIn, “The point is not about having the best model in the market that is competitive with the frontier but a good enough model for some legal workflows. The reality is that it’s more of a cost control measure.” He continued: “I’m sure Thomson is going to be more than enough for summarisation and will help TR save some money on tokens,” highlighting the significance of large legal AI companies shifting to Chinese OSS models.

And turning to frontier models, Google has joined Microsoft, Anthropic, OpenAI etc and launched Gemini Enterprise for Legal, recruiting Tony Ensinger from DeepJudge to head its legal AI division. Gemini Enterprise for Legal includes purpose-built skills, partner agents and model context protocol (MCP) integration with legal AI platforms, including Harvey, Legora, and Thomson Reuters.

Digital twins and avatars

Twin1 AI, a start-up building digital AI twins for lawyers, helps firms preserve lawyers’ tacit knowledge and boost efficiency by automating routine communications in a personalised way. Lawyers can address the risk of oversharing (which is a concern given GenAI’s tendency to be verbose) by managing their digital twin’s privacy/permissions/review settings across multiple channels. While digital twins have been around for a while, Twin1 takes the concept further with deep personalisation and (the option of) setting up a Twin Network where digital twins can communicate and coordinate with each other and lawyers across the firm’s teams and systems.
In education, the AI debate is increasingly polarised. While some universities are strongly against students using AI, with a few banning its use altogether, others are using it to create and deliver personalised e-learning programmes. Harvard has taken this in a new direction, introducing AI avatar tutorials. TechCrunch reported that Harvard’s $699 HBS Foundry startup bootcamp follows up live sessions with instructors with individual feedback on practice pitches and board meetings from AI avatars of its instructors created by HeyGen. This is unusual as educational institutions tend to opt for AI to deliver group/general sessions/e-learning on demand and human tutors to provide personalised feedback and support. 

While digital twins and avatars save time, they are of course replacing human interaction, which doesn’t matter when they are answering repetitive questions or managing routine tasks, but they also risk limiting the personal connection that often underpins business and client relationships.

August announcements included more partnerships and MCP integrations between legal AI vendors. Given last month’s headlines about rogue AI agents, there are surely questions around whether multiple integrations, semi-autonomous agents, including digital twins and avatars, etc make it easier for sophisticated AI models to find pathways into law firm systems and client data.

Planning for an AI future

Bill Gates highlighted the potential for AI cyberattacks when he wrote last week that the choices we make now are critical to ensure that AI is a force for good. He urged governments and organisations to prepare now for the AI era. For example on AI replacing jobs, Gates recommends designating human-reserved jobs and taxing robots and AI tokens because employing people carries tax obligations while technology is treated as a business investment/expense. “The tax system nudges you toward replacing people with machines,” he wrote. In anticipation of a legal AI future, both Harvey and Legora are in talks to raise significant additional funding.

Automating AI slop!

Finally, on 30 July, LinkedIn added a ‘seems like AI slop’ button where you can notify the platform if you suspect posts have been AI generated (which is interesting as it previously offered AI assistance to ‘enhance your post’). Meanwhile, one (not entirely serious) website offers a quick way to add more AI slop to a platform whose brand has been significantly contaminated by GenAI. The viral LinkedIn CringeBot 3000 helps you “become a thought leader in seconds”!

Learn more about upcoming Legal Geek events on our events page.

share
Addleshaw Goddard Workshop

Level up your prompting game: Unlock the power of LLMs

A workshop intended to dive into the mechanics of a good prompt, the key concepts behind ‘prompt engineering’ and some practical tips to help get the most out of LLMs. We will be sharing insights learned across 2 years of hands-on testing and evaluation across a number of tools and LLMs about how a better understanding of the inputs can support in leveraging GenAI for better outputs.

Speakers

Kerry Westland, Partner, Head of Innovation Group, Addleshaw Goddard
Sophie Jackson, 
Senior Manager, Innovation & Legal Technology, Addleshaw Goddard
Mike Kennedy, 
Senior Manager, Innovation & Legal Technology, Addleshaw Goddard
Elliot White, 
Director, Innovation & Legal Technology, Addleshaw Goddard