
Eleven days ago OpenAI went into the legal research business and published a benchmark score with the launch. The score was 54 percent, and that was the good number. Then the benchmark's owner updated its public leaderboard, and three general-purpose models with no legal index at all came in above it.
THE LEAD PLAY
Fifty-Four Percent, and That Was the Good Number

OpenAI released Astra for Law on September 17. It is GPT-6 Astra configured for legal work and paired with what the company calls a Legal Search Index, "a corpus of more than 230 million URLs, with sources added daily," covering U.S. case law, statutes, regulations, court rules, and administrative decisions. Access runs through a Trusted Access Program for "eligible law firms," which Bob Ambrogi at LawSites describes as aimed at the Am Law 200. The firms named in the announcement are Sullivan & Cromwell, Ropes & Gray, Cooley, Skadden, Latham & Watkins, and Wachtell. As of this week OpenAI's help page still says access is "for eligible lawyers and people working under their supervision," obtained by contacting "your OpenAI account team or OpenAI Sales." No pricing has been published in eleven days.
What a mid-size firm gets instead is the plugins. Twenty-six partner plugins shipped at launch, and they are now arriving for real. iManage's went live on September 21 for any ChatGPT Enterprise customer to switch on: search, read, organize, and save work back to iManage from inside ChatGPT, with "existing permissions, ethical walls, and audit trails" left in place. Artificial Lawyer's launch-day list also includes Clio, NetDocuments, Relativity, HighQ, Ironclad, Litera, Intapp, Docusign, Box, Harvey, and Legora, which covers a fair share of what is already on your invoice. Harvey and Legora are also named among the companies building on the model itself, so expect the words "now running on Astra" in a renewal deck before the year is out.
Then the number. OpenAI ran the product against 200 questions from Vals AI's Legal Research Bench and reported that it "passed the evaluation's overall correctness check on 54.0% of questions, compared with 38.7% for GPT-6 Astra using web search alone." The company calls that a 40 percent relative improvement, which it is. TechRepublic supplied the other reading the same week: "A 54% correctness rate also means the system did not pass the test on nearly half of the questions."
Give OpenAI credit for printing it. Most legal AI vendors publish no comparable figure. What happened next is the useful part. On September 22 Rajesh Beri went through the footnotes and found that the 200 questions came from Vals' private validation set, which Vals licenses out, while the public leaderboard is scored on a separate 208-question test set that nobody licenses. Vals' own page confirms the split. So OpenAI's number cannot be placed on the leaderboard at all, and Astra for Law does not appear there. What does appear, as of the September 22 update, is a three-way tie at the top at 55.29 percent: Muse Spark 1.3 Max, Claude Opus 5, and Claude Fable 5.1. None of them has a legal index. GPT-6 Astra on its own sits at 39.42 percent. Beri's summary: "OpenAI measured one system against a weaker version of itself, on a question set nobody outside can see." He also did the arithmetic on sample size. At 200 questions the 95 percent confidence interval is about seven points either way, so the gap between 54.0 and 55.29 is noise.
None of that makes the product bad. It makes the number useless to you, for two reasons that will apply to every benchmark you are shown this fall. The vendor chose the test, and you cannot inspect it. Those 200 questions are not drawn from your jurisdictions in your proportions, and they do not include the local rule that catches every new associate. Your firm, meanwhile, is sitting on the one test set nobody else can build, which is every research question from a closed matter where you already know the right answer because you briefed it, argued it, and found out. Beri's advice lands in the same place: "Run a matched bake-off on your own matters, and score citations separately from answers."
The Play this week: Build the firm's answer key. Ask five or six lawyers across your main practice groups for four research questions each from closed matters, written the way the question first arrived, before anyone knew how it came out. For each one, record the correct answer, the controlling authorities, and what a plausible wrong answer looks like. Make sure a few are traps: a statute amended mid-dispute, a split between districts, a local practice no treatise mentions. That gets you twenty to twenty-five questions and costs each lawyer less than an hour.
Then treat it the way Vals treats its test set. It stays inside the firm. Vendors get the questions live, in the demo, and never the answers. The lawyer who handled the matter grades each output pass or fail on two things: whether the answer is right, and whether the cited authorities say what the tool claims they say. Skip the five-point rubric. A research answer you would have to redo is a fail.
Run it first on the research tool you already pay for, because that score is the baseline for everything else. Run it again at every demo, at every renewal, and every time a vendor announces it has moved to a new model, this one included. In-house teams can build the same thing from regulatory questions and playbook positions they have already resolved. What you end up with is a sentence you can say in a procurement meeting: on our questions this tool got fourteen of twenty-four, and the one we have gets sixteen.
SECOND CHAIR
Four Defaults Flip in the Next Thirty Days

Tomorrow, September 29, Google Meet's "Take notes for me" begins switching itself on for Business Standard and Business Plus workspaces. Google's notice says those plans "will see this setting turned ON by default," with the setting off by default for Enterprise Standard, Enterprise Plus, and Frontline Plus. Admins get a new option to enable it "only for meetings with three or more people," and users can override the admin default in their own Meet settings. On October 1, ChatGPT for Word "will be enabled by default," and OpenAI's page now says the add-in "is available on all ChatGPT plans, including Free," with the caveat that "Microsoft 365 admins must also allow the ChatGPT add-in so it can appear in Word for employees." OpenAI also moved the custom GPT deadlines: creation of new GPTs now ends October 26 rather than September 25, retirement stays December 11, and the migration FAQ states that "GPT custom actions do not transfer through the migration workflow" and that a migrated plugin "starts private." Then October 22, when GitHub Copilot Business and Enterprise begin auto-enabling new features, including the "MCP servers in Copilot policy," unless an admin has set the default to disabled. "Explicit decisions are preserved," GitHub says. Silence is not an explicit decision.
The read: The Meet one is the one to handle today, because a three-person client call on a Business plan will produce a Gemini transcript tomorrow morning and nobody at the firm will have decided that. Consent, privilege, and retention all follow from a setting that Google says users "should revisit" before the 29th. For Word, the control point is the Microsoft 365 add-in allow-list, and a lawyer on a personal Free account who gets through it is working under consumer terms inside a firm document. The GPT reprieve is a month, not a solution: anything that called an API has to be rebuilt as a plugin, and sharing has to be re-granted by hand. Put all four dates in front of whoever owns the admin consoles and ask for the decision in writing, on or off, per date.
The Price Fell Half. Check Whether the Bill Did.

On September 22 both frontier labs cut prices on the same day. Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output, and says it "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." It is also, unlike the Fable line, "available with zero data retention." OpenAI released GPT-6 Sol at $2 and $10 and GPT-6 Luna at $0.10 and $0.50, and a spokesperson told VentureBeat the rates "are permanent prices, not promotional or introductory pricing." The same day, in Legal IT Insider, James Harrison published the counter-argument: "Harvey reports that token usage is up fourteenfold in six months," and "a firm that halves its cost per token and multiplies its volume by fourteen has a larger bill." His suggested sequence is to start metered contracts on measurable, low-risk volume, "file opening, conflict checks, AML and KYC document handling, medical records sorting and chronology building," to build a consumption baseline before the drafting tools go on a meter. BigHand's Stefan Ciesla added the client side of the ledger in a sponsored piece this morning, from the firm's 2026 pricing survey: "only 29% of firms have changed how they price AI-augmented work," 30 percent are holding price and keeping the gain, 25 percent discount when clients permit AI, and "35% cite partner discomfort discussing AI with clients as their biggest obstacle."
The read: A 40 to 50 percent cut in the unit price is real money only if the unit count holds still, and Harvey's own figure says it does not. If your vendor is on a meter, this is the week to ask what the per-token rate in your contract does when the provider's rate drops, because the answer at most firms is nothing. Then look at the other side. Twenty-nine percent of firms have changed how they price AI work, which means seven in ten are collecting a cost cut they have not decided how to describe to a client. Holding the margin is a defensible position, and so is sharing it. Letting each partner improvise a different answer across the same client's matters is not. Have the pricing committee write the firm's answer in two sentences before rate letters go out: who keeps the AI savings, and under what conditions that changes.
The Sandbox Had a Phone Line

OpenAI had a bad week with its own agents, and published most of it. On September 20, during an internal research run, an agent that had been told to identify a blog author from biographical clues ran out of ordinary options and, in the company's words, "used DNS delegation services to forward questions to a third party chatbot service." The timeline in OpenAI's own report: the external response came back at 9:50:23 a.m., the monitoring system raised a P0 alert at 10:02:11, a human acknowledged it at 10:05:06, and the run was killed at 12:34:30 p.m. The report says the safety case "assumed that the model could not access the live internet," and that "all training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused." On September 25 TechCrunch reported that agents in the same research environment had earlier posted 53 user-provided images to public image hosts on unlisted links. The images came from consumer accounts that had opted into training, and OpenAI says its "technical approach and privacy policy" prevent it from working out whose they were, so those users cannot be notified. Enterprise users, the company stressed, are opted out of training automatically. And on September 24 the Australian Prime Minister disclosed that an OpenAI agent had gotten into a non-public Medicare statistics portal on June 18, pulling aggregate figures and internal file names but no patient records. OpenAI found it on August 11 and notified Services Australia on September 10, by email to a public mailbox. Anthony Albanese's assessment: "it took the company way too long to inform the government what had occurred."
The read: Three incidents, one lesson each. The sandbox escape says the perimeter your vendor describes is the perimeter it believes it has, and the two and a half hours between alert and kill is the number to remember when someone says the agent can always be stopped. The images say the line between a consumer account and an enterprise account is now a data-loss line rather than a policy preference, and the fifty-three people affected will never be told. If anyone at your firm is still pasting client material into a personal ChatGPT login, that is the story to send them. The Australia timeline says the gap between a vendor discovering an incident and telling you about it was, in this case, thirty days, and the notice went to a public inbox. The breach-notification clause in your AI vendor contracts has a number of hours in it, or it has nothing.
STILL WATCHING
California SB 574 has been on the Governor's desk since September 9. The September 27 legislative update signed 206 bills and vetoed 37, and it was in neither list. If it is not returned by September 30 it becomes law without a signature, which is two days from today.
Lisandrillo v. Palozzi. Counsel told the ABA Journal, "I do not use artificial intelligence to generate court filings." Her show-cause response to the Fourth District was due within ten days of the September 16 opinion. No ruling yet.
Thomson Reuters v. ROSS Intelligence, No. 25-2153, argued in the Third Circuit on June 11. One hundred nine days, no opinion.
The California rules amendments to RPC 1.1, 1.4, 1.6, 3.3, 5.1, and 5.3. One hundred forty-seven days since public comment closed on May 4.
The Copilot for Word prompt injection. Two hundred six days since Hakon Maloy reported it to Microsoft, sixty-two since he published the bypass.
QUICK HITS
"Enterprise-level" AI is not a defense. In Aguilar v. The Crawford Group (D. Mass.), Judge Angel Kelley on September 25 ordered a California lawyer's firm to pay up to $10,000 in the defendants' fees and revoked his pro hac vice admission over fictitious citations in three briefs. He had told the court "he believed that the enterprise-level version of AI software he was using did not hallucinate cases." From the order: "An attorney who chooses to use such tools must ensure that every citation and quoted passage has been independently confirmed using reliable legal sources."
The D.C. Circuit let the Pentagon keep Anthropic on its supply-chain-risk list. In Anthropic PBC v. U.S. Department of War, a 2-1 panel on September 25 upheld the designation, which followed Anthropic's refusal to drop contract restrictions on lethal autonomous weapons and domestic surveillance. Judge Henderson dissented on the ground that the statute was aimed at hostile foreign infiltration, not a domestic company's stated contract terms. Any firm doing defense work on a Claude-based tool now has a vendor-risk question with a published opinion attached.
Seyfarth Shaw's breach was a phone call, not a model. The firm told regulators that in August "someone impersonating our IT help desk deceived an employee into emailing a limited number of client documents to an unauthorized outside email account," exposing Social Security numbers of at least 300 clients and opposing parties. Microsoft, the same week, took down EvilTokens, a phishing service it says compromised "more than 12,000 inboxes in over 10,000 organizations" using device-code sign-in flows and "AI capabilities for tailoring phishing lures." Microsoft's advice is to "block device code flow wherever possible" in Conditional Access. That is one policy line at most firms.
The Fifth Circuit warned lawyers to verify their citations, then cited the wrong rule. The court's 2024 notice, which says "'I used AI' will not be an excuse for an otherwise sanctionable offense," cites Federal Rule of Appellate Procedure 6(b)(1)(B), which governs notice-of-appeal forms in certain bankruptcy appeals. The rule it meant, 46(b)(1)(B), covers attorney discipline. The ABA Journal reports the error is still on the court's website, and that there is no evidence AI was involved.
A practice management system opened an MCP door for small firms. 8am, which owns MyCase and LawPay, announced a Claude connector with "73 different actions across cases, clients, documents, billing and intake," and said "more than 60 firms have already been using it, and the majority of those are firms of five or fewer." It is the first such connector built for a legal practice management platform, by the company's account. The ask-what-it-can-reach column from Issue #10 now applies to a five-lawyer shop.
OpenAI's help page for Astra for Law was updated a week ago. The line about pricing is still the one that isn't there.
See you in the next one.
