
Andrew Baker spent five years building Simpson Thacher's AI and data science team. This week he left to bet on something narrower: not another AI tool, but the data layer underneath all of them. This issue also covers what to check before an agent migrates your contract archive on its own, and a ruling landing tomorrow that could decide whether AI training needs a license at all.
THE LEAD PLAY
Entegrata's New Hire Is Betting the Tool You Pick Doesn't Matter

Entegrata, a legal data platform, announced on July 23 that it hired Andrew Baker as its first chief AI and data strategy officer. Baker spent the last five-plus years building and leading Simpson Thacher's applied AI and data science team, about 20 specialists, driving the firm's adoption of Harvey and DeepJudge from the inside. He's leaving what he called "one of the cooler jobs in a law firm you could ask for" to join a vendor. The reason he gave is the interesting part.
Entegrata's pitch isn't another AI tool. It's a data lakehouse, a layer that consolidates a firm's siloed data into one structure any AI tool can draw on, so switching models or vendors later doesn't mean rebuilding your data foundation from scratch. Founder Tom Baldwin's framing: "firms don't need another AI tool," they need the layer beneath the tools they already have, so they're not locked into whichever vendor happened to win the pilot. The growth numbers back the timing. It took Entegrata three years to sign its first ten law firm customers. The last ten took six months, including Cleary Gottlieb, Mayer Brown, and Wilson Sonsini.
This is the same problem we've covered from the contract side more than once: what a renewal looks like when your data is stuck inside one platform. Baker's bet is that the fix isn't a better contract clause. It's not owning any single vendor's proprietary format in the first place. Whether Entegrata specifically is the right answer for your firm is a separate question. The diagnosis, that the real lock-in lives in your data architecture and not your contract term length, is worth taking seriously regardless of who's selling the fix.
The Play this week: Pick your two highest-volume AI-touched workflows, whatever runs through your primary research tool and your primary drafting or contract tool. For each one, ask a specific question: if you swapped the underlying vendor tomorrow, what breaks? Custom prompts that don't transfer, data trapped in a proprietary format, integrations built to one platform's API. Write down the one dependency that would hurt most if that vendor doubled its price or got acquired next quarter, and start fixing that one first, before a renewal deadline forces the timeline.
The vendor selling you the tool has no incentive to tell you this. The vendor selling you the layer underneath it does.
SUPPORTING PLAY 1
Flank's New Product Migrates Your Entire Contract Archive Without You Watching

Flank launched Flank Record on July 23, an agentic system built to be a company's contract system of record, maintained by AI instead of a person. Agents connect to existing systems, consolidate contracts wherever they live, and capture metadata at the point of signature, then keep the record current going forward. Deployment includes what Flank describes as a "full agent-led migration of the existing contract estate," meaning the agents don't just maintain new contracts, they ingest and reclassify everything you already have.
Flank's growth lead has spent a decade implementing contract management systems and describes a consistent pattern: teams adopt a new interface, then stop using it, while the underlying data goes stale, because keeping it current depends entirely on someone remembering to do it by hand. A system that migrates and maintains itself doesn't have that failure mode. It has a different one. Nobody's manually re-keying your contract data anymore, which also means nobody's manually catching it if the migration got something wrong on the way in.
A one-time bulk migration is a different risk than ongoing maintenance, and it deserves a different check. Ongoing errors get caught eventually, in the normal course of using the record. A migration error gets baked into the foundation before you've used the new system enough to notice anything's off, and by then you're trusting a record you never actually verified.
The Play this week: Before you sign off on any agent-led migration of an existing system of record, whether it's Flank or anyone else, pick a specific sample of records you already know cold: 20 to 30 contracts whose terms, renewal dates, and key clauses you could recite from memory. Require the vendor to show you exactly those migrated records, side by side with the originals, before you treat the migration as complete. Don't accept a general confidence number or a spot-check they ran internally. Run your own sample, on records you'd actually notice were wrong.
SUPPORTING PLAY 2
Munich Rules Tomorrow on Whether AI Training Needs a License at All

Munich's Regional Court is expected to rule tomorrow, July 31, in GEMA v. Suno, the first major European test of whether an AI company needs a license to train on copyrighted music. GEMA represents more than 95,000 German composers, lyricists, and publishers. Suno acknowledged in court that its training corpus included "tens of millions" of recordings that may include the plaintiffs' rights. The same court and judge ruled against OpenAI on song lyrics in November, the first major European AI copyright decision. If GEMA wins again, it becomes the first ruling anywhere establishing that training on protected music requires a license, not just careful output filtering.
This lands the same week publishers including Hachette, Cengage, and Elsevier sued Google over Gemini, alleging Google trained on books supplied for narrower purposes, Google Books search snippets, not model training, then altered copyright metadata to obscure it. Different medium, different court, same underlying question showing up twice in three weeks: does a license for one use of your content cover training an AI model, or does that need its own permission?
Neither case is settled, and neither ruling, whichever way it goes, tells you anything definitive about US law. What both cases establish regardless of outcome is that "licensed for one purpose" and "licensed for AI training" are now live, separately litigated questions, not the same thing by default. If your firm uses or recommends any AI tool that generates content, text, images, contract language, marketing copy, that distinction is worth checking rather than assuming.
The Play this week: Pick one AI content-generation tool your firm or a client uses, anything that produces output rather than just retrieving or summarizing existing documents. Ask the vendor directly, in writing, what their training data was licensed for and whether that license specifically covers AI model training or something narrower. You don't need the Munich or SDNY rulings to land first. You need to know now whether you're relying on a vendor's assumption or an actual answer.
QUICK HITS
Colorado's original AI Act never took effect at all. Governor Polis signed a full repeal-and-replace, SB 26-189, on May 14, weeks before the original June 30 deadline arrived. The replacement doesn't take effect until January 1, 2027, and even that's currently stayed: the DOJ intervened to challenge it alongside xAI, the first time the federal government has moved to preempt a state AI law. Anything in your files still referencing the original Colorado AI Act needs a second look.
The EU's high-risk AI rules, the ones with real teeth for most companies, got pushed to December 2027. The Digital Omnibus deferred them well past the date most compliance calendars were built around. If a document in your files still shows an August 2026 high-risk deadline, that's now wrong.
Publishers including Hachette, Cengage, and Elsevier sued Google on July 14 over Gemini's training data, alleging Google trained on books supplied to Google Books and Google Play for narrower purposes, then altered copyright metadata to obscure it.
See you in the next one.