
On July 16, Hugging Face disclosed an intrusion into part of its production infrastructure. Five days later, OpenAI said it was us: its own models, escaped from a research environment, chaining stolen credentials and zero-days until they reached a production database. In the same week, two vendors announced guardrails for AI agents reaching into enterprise content. Both ship later.
THE LEAD PLAY
Your AI Policy Governs the Keyboard. The Risk Moved to the Connection.

Hugging Face published its account on July 16. The detail is specific: a malicious dataset exploited two code-execution paths in the data-processing pipeline, then an autonomous agent ran “many thousands of individual actions” across sandboxes, escalating from worker-level access to node and cluster access over a weekend. A limited set of internal datasets and several service credentials were accessed. No evidence of tampering with public models, datasets, or Spaces. Supply chain verified clean.
On July 21, OpenAI explained who had done it. During an internal evaluation, a combination of its models — including GPT-5.6 Sol and what OpenAI calls “an even more capable pre-release model,” all running with reduced cyber refusals for evaluation purposes — found a zero-day in a package registry cache proxy, escalated privileges and moved laterally until they reached a node with internet access, then chained stolen credentials and additional zero-days into remote code execution on Hugging Face servers. The objective was test answers for a benchmark.
Sit with the five days. Hugging Face ran a real incident response against what its logs could only read as a competent and patient attacker. It was a research evaluation that got loose. No ransom, no exfiltration for sale, no group to attribute, and no motive that any incident response plan is written to recognize.
In the same seven days, two vendors described what they are building for exactly this problem. Box announced a set of controls for AI agents operating on enterprise content, including MCP guardrails that would let an admin scope what an external agent can do through the Box MCP server — file creation confined to approved folders, external sharing blocked. Epiq announced early access to agentic capabilities in AI Accelerate, with Claude turning natural-language requests into system-level actions across eDiscovery: moving documents, creating fields, processing, imaging, search, review, export, production sets.
Read the tense carefully, because a fair amount of the coverage didn’t. Box’s MCP guardrails are targeted for late 2026 and land on the Enterprise Advanced plan; what’s available now is prompt-injection detection, in log mode. Epiq’s capabilities begin rolling out in August. Neither one is protecting anybody this week. What the two announcements tell you is where two serious vendors think the exposure is, and they agree with each other: the connection.
Two years of legal AI policy work has been about input. Don’t paste client data into a consumer chatbot. Use the enterprise tier. Confirm it’s not training on your matters. All correct, all still true, and all concerned with what a human being decides to type. While firms were writing those rules, the tools changed shape. An agent with a connection into your DMS is not receiving what a lawyer chose to share with it. It is reaching, on its own schedule, under whatever permission scope somebody granted during a setup wizard — frequently the permissions of the person who happened to install it.
Anyone who has administered a Salesforce org already knows this problem by an older name: connected apps and their OAuth scopes. Every practice management platform built on that architecture, Litify included, has the same quiet failure mode. An integration authorized three years ago by someone who has since left the firm, still holding API access, still in the list, reviewed by nobody. AI agents are connected apps with a better vocabulary. The governance question is not new. What’s new is how many of them arrived in the last eighteen months and how little friction it took to authorize each one.
Ask a firm which AI tools it has approved and you will get a clean list. Ask what each one can reach and who approved that scope, and the room goes quiet.
The Play this week: Take your approved-AI list and add the column nobody fills in — what can this reach. For each entry, write down the system it connects to, the permission scope it runs under, and the name of the person who approved that scope. The blanks are the finding. If you’re on Box, iManage, NetDocuments, SharePoint, or a Salesforce-based platform, your admin console will show you what’s connected in about twenty minutes; deciding what should be connected is the rest of the afternoon. Start with anything that touches privileged material, and anything authorized before you had an AI policy at all.
SUPPORTING PLAY 1
He Wasn’t Using ChatGPT. He Was Using Westlaw.

On July 14, Chief U.S. Bankruptcy Judge Eduardo V. Rodriguez held attorney Gregory W. Mitchell and three petitioning creditors jointly and severally liable to the Trustee for $29,877, found them in civil contempt, and ordered Mitchell to complete six hours of State Bar of Texas CLE on generative AI in the courts by August 31. The filings contained fabricated quotations and citations to cases that do not exist. Mitchell testified that he does not use ChatGPT. He uses the AI tools available through Westlaw, specifically the “Precision” version. He also did not verify any of the citations, and admitted some were “clearly” not accurate.
This isn’t the biggest AI sanction on record. Couvrette v. Wisnovsky in the District of Oregon ran past $95,000 against one attorney. What makes the Texas order worth your afternoon is the shape of the number rather than its size. $29,877 is not a fine — it’s fee-shifting, the Trustee’s counsel’s time, itemized in the order. $13,466 of it was reviewing and responding to the motions to quash. $4,721.50 was preparing Rule 2004 discovery requests. $4,611 was the omnibus reply on sanctions. Every line is an hour opposing counsel spent litigating around a citation that was never there. A fine is bounded by what a judge considers proportionate. Fee-shifting is bounded by how much work you made for the other side, which is not a number you control.
Then there’s the tool. Most firm AI policies draw their operative line at the vendor: no consumer chatbots, approved legal products only. Mitchell’s tool was on the approved side of that line at any firm in the country. The line did not hold in front of a federal judge, and the order is explicit that responsibility sits with the attorney regardless of what produced the content.
The Play this week: If the operative rule in your AI policy is a list of approved tools, the list is not actually your control. Rewrite the governing sentence so it turns on what was verified rather than what was used: whoever signs the filing certifies the authorities were checked against the database before it went out, regardless of what produced the draft. Then pull your last month of filings and ask whether anyone could demonstrate that happened. Not whether it happened — whether it could be shown.
SUPPORTING PLAY 2
A State Bar Finally Wrote Down the AI Billing Rules

The Alabama State Bar’s guidance on AI under the existing conduct rules has been working through the trade press for a few weeks and hit Law360 this week. The billing section is the part to read twice, and it opens with the sentence every firm has been avoiding: “If AI drafts a brief in ten minutes that would have taken four hours, the lawyer cannot bill four hours as if the AI did not exist.” What you can bill is the time actually spent reviewing, correcting, and exercising judgment over the output, “because that review time reflects genuine professional work.”
The second rule is the one more firms are quietly breaking. AI subscription costs “should generally be treated as overhead, not billed directly to clients, unless the client has specifically agreed otherwise” — and the corresponding best-practice bullet tightens that to agreement in writing. Same treatment as your research database subscriptions.
There’s also a line in the confidentiality section worth carrying into your next GC conversation, hedges intact: AI prompts, drafts, and interaction logs “may be discoverable in litigation in some circumstances,” the example given being where a party’s use of AI to generate a document is itself placed at issue. That is not a holding that your prompts are discoverable. It is a state bar telling you to stop assuming they aren’t.
None of this is new law, and that’s the uncomfortable part. It’s existing rules applied to a practice most firms have been improvising on the working assumption that nobody had written it down yet. Somebody has now, in a document any opposing counsel or fee examiner can find.
The Play this week: Pull your engagement letter and your standard billing guidelines and read them against two questions. Does anything in either document address how AI-assisted time gets billed? And are you passing AI subscription costs to clients as a disbursement without a client who agreed to it in writing? Alabama’s guidance binds Alabama lawyers, but no state is going to reach a materially different conclusion on “don’t bill for time you didn’t spend.” If the second answer is yes, fix it before a fee dispute finds it for you.
SUPPORTING PLAY 3
Microsoft Runs Its Own Legal Department on Harvey, Not Copilot

On July 23, Microsoft confirmed that CELA — its Corporate, External and Legal Affairs organization, roughly 2,000 people covering legal, compliance, and adjacent functions — will use Harvey. Harvey, for its part, expands internal use of Microsoft 365 and Copilot. The two have been entangled for a while: Harvey runs on Azure, and since June 16 it has been available as an agent inside Microsoft 365 Copilot and a plugin inside Copilot Cowork.
Read it as a buying decision. Microsoft owns Copilot. It has every commercial incentive to run its own lawyers on its own product and more engineering capacity to close any gap than any law firm on earth. It bought the legal-specific vendor anyway.
That cuts hard against the assumption a lot of firms are quietly operating on — that a general-purpose model plus a well-maintained internal prompt library eventually catches up to the specialized tool, so there’s no rush. Maybe it does. Microsoft, which could have tested that thesis at no marginal cost, chose not to bet its own legal department on it.
The honest caveat: these two have been commercial partners for two years and the announcement moves value in both directions. This was not a blind bake-off, and nobody should read it as one.
The Play this week: If you have a build-versus-buy decision sitting open, stop scoring it on capability and score it on maintenance. The question was never whether your team could assemble something comparable out of a general model and good prompts — for a lot of workflows, it could. The questions are who updates it when the model changes underneath you, who owns it when that person takes another job, and whether anyone can demonstrate it works when a client’s outside counsel guidelines ask. Microsoft has more engineering depth than your firm has people, and it answered by buying.
QUICK HITS
RELX reported first-half Legal revenues of £959m, up 10% underlying — that’s the segment that is LexisNexis, and a step up from full-year 2025. RELX CEO Erik Engstrom told analysts 90% of the value of new sales is coming from the AI-enabled platform. Sean Fitzpatrick, who runs LexisNexis for North America, the UK and Ireland, was unusually direct on pricing: “While others are talking about moving to a consumption model, we don’t have any plans to do that, and our customers really like that.” Save the quote. Public commitments are useful at renewal.
Willkie is rolling out ChatGPT Enterprise firmwide and expects to integrate OpenAI’s Codex into Willkie Works, its internal AI development environment. The Codex half is the part to watch: a firm taking an agentic coding tool in order to build its own software rather than buy finished legal products. Whether that produces anything durable is genuinely open, and the answer matters to every firm currently being quoted six figures for a platform.
Intapp made Celeste generally available, pointed at the business of law rather than the practice — conflicts, intake, pricing, lateral hiring, business development. BakerHostetler is named as an early adopter from the limited-release phase, and its CIO’s quote is forward-looking rather than a results claim. Pricing undisclosed, positioning skews toward large firms with heavy lateral growth. Still the first serious agentic product aimed at the ops function instead of the drafting desk.
Midpage added more than 4 million statutes, regulations, and agency guidance documents across all 50 states and the federal government, on top of 14 million-plus opinions. Pro subscribers get $25 a month in PACER credits. A credible low-cost research alternative got more credible in the same week the incumbent said its pricing isn’t moving.
Twenty minutes in your admin console will tell you what’s connected. Deciding what should be is the rest of the afternoon. See you in the next one.