01 Research & evidence
Claude Opus 4.8 got 42.9% of expert legal questions right
Anyone relying on an AI assistant to research a regulated question needs a human check, because the best one tested failed most of the time.
The benchmark counted an answer correct only when every required point was present and every authority it cited actually checked out.
Katrina Drozdov, Oliver Chen, Langston Nashold and Rayan Krishnan built Legal Research Bench from 413 open-ended United States legal research questions written by experts, each paired with a gold answer, the authorities that support it, and a pass-or-fail rubric. They ran thirteen frontier models through it with web search, case-law search and page-parsing tools available.
The scoring is the part that matters. A response counted as correct only if it satisfied every required criterion and its cited authorities verified. On that bar the strongest model, Claude Opus 4.8, was fully correct on 42.9 percent of the questions.
One finding should change how anyone buys this. More turns, more tool calls and more inference cost did not predict higher accuracy. A longer and more expensive run did not come back more reliable, so paying for more thinking is not a fix.
Difficulty also tracked the kind of question. Pass rates were lower on questions that required reconciling conflicting authorities, which is most of what a professional is actually paid to do. The authors checked their automated grader against practicing attorneys so the scores track attorney judgment.
What to do about it
Take the question in your own field you would least want answered wrongly and put it to whichever assistant you or your staff already use. Then check two things: whether it names the authority you would have named, and whether the sources it cites say what it claims they say.
This is a benchmark of research questions, not of your intake or your drafting. Read it as a ceiling on unsupervised research, and keep the human review step wherever an answer carries professional exposure.
Source arXiv, Drozdov, Chen, Nashold and Krishnan, “Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents,” September 30, 2026 · arxiv.org
02 Operational change
Google’s spam update is still rolling after nine days
If your traffic moved on September 30, Google’s spam update is the likeliest reason, and it has not finished rolling out.
Google said this update would take about two weeks rather than the usual two days, and a second wave landed on September 30.
Google announced the September 2026 spam update on September 24. The first sites felt it on September 25, 26 and 27, which this brief reported on September 28.
A second phase hit on September 30 and carried into October 1. Glenn Gabe, who tracks these rollouts, reported that he “received emails from several site owners explaining they saw a huge drop right on 9/30.”
The reason it keeps moving is in Google’s own note. This update was set to roll out over about two weeks rather than the usual two days, and two weeks from September 24 runs into the week of October 5. A position you check this morning is not necessarily the one you keep.
Spam updates target manipulative practices, not ordinary publishing. A drop during one is not automatically a judgment on your business, because it can equally be sites above you moving up or down. The practical consequence is to wait for the rollout to finish before rewriting anything.
What to do about it
Write down today’s numbers for your top ten pages and leave them alone. If you rewrite during a rollout you will not be able to tell afterwards what the update did and what your own edit did.
Check back in the second full week of October. If a drop holds once the rollout is called complete, that is the point to look at the pages involved.
Source Search Engine Roundtable, Barry Schwartz, “Phase Two Of The Google September 2026 Spam Update Hits On 9/30,” October 1, 2026 · seroundtable.com
03 Research & evidence
ChatGPT, Claude and Gemini found 39.2% of expert-chosen studies
An AI assistant finds the large well-known studies in your field and quietly leaves out the smaller ones, which may include your own work.
Sample size was the only thing that significantly predicted whether a study got retrieved, so size mattered more than relevance did.
Five researchers put 20 clinical questions drawn from 2026 Cochrane reviews to three assistants: Claude Sonnet 5, Gemini 3.1 Pro and ChatGPT GPT-5.5. They asked as a patient, as a clinician and as an evidence-synthesis researcher, four times each, producing 720 responses.
On average a single response retrieved 39.2 percent of the studies the matching Cochrane review had included, and 5.0 percent of the ones it had deliberately excluded. Recall split hard by assistant: ChatGPT 63.1 percent, Claude 37.0 percent, Gemini 17.3 percent.
The predictor was size, not merit. Each doubling of a study’s sample size came with 50 percent higher odds of being retrieved, an odds ratio of 1.50. Asking in the researcher voice helped a little, 42.8 percent against 36.1 percent for the patient voice, far less than the choice of assistant did.
Cochrane reviews are the closest thing medicine has to a settled answer, so this is close to a best case. If your standing rests on work that is specialized or small, an assistant is likelier to miss it than to find it. Veterinary specialists, boutique advisory practices and niche engineering consultancies all sit in exactly that position.
What to do about it
Ask an assistant the question your field would consider settled and see which sources it returns. If the work you are known for is absent, that absence is what a prospective client sees too.
The practical response is not to publish more, it is to make the work you already have easier to find and cite: on your own site, in plain language, with the numbers and the method stated.
Source arXiv, Liu, Jin, Menke, Kahnt and Lu, “Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions,” revised September 29, 2026 · arxiv.org
04 Consumer & market data
202,590 donated ChatGPT conversations show what people ask
The AI usage statistics everyone quotes come from the platforms themselves, and until now nobody outside could check a single one of them.
The data covers India, Nigeria, Brazil and Pakistan, so it is a check on the statistics rather than a picture of your own customers.
Shreyasi Roy Chowdhury and Kiran Garimella collected complete ChatGPT export files from 1,252 people across India, Nigeria, Brazil and Pakistan, covering 202,590 conversations from December 2022 to February 2026, with self-reported age and gender.
Their reason for doing it is the finding. What the public knows about AI use comes from OpenAI’s and Anthropic’s own aggregate reports, which apply fixed categories to hundreds of millions of users and release only summary statistics that outside researchers cannot re-analyze.
In this sample personal use accounted for 55 to 64 percent of conversations. Conversations in which people express themselves rose from a few percent to roughly a fifth or more across the period.
Four countries is not the world and it is not your market, and the honest way to read this is as a check on the numbers rather than a measure of your buyers. Forever Cited quotes AI usage figures in its own work, and nearly all of them trace back to the companies that benefit from them. That is worth saying plainly rather than repeating them as though they had been audited.
What to do about it
Next time a figure about AI adoption is quoted at you, ask one question: who counted, and can anyone else check it? For most of the numbers in circulation the answer is the platform, and no.
Nothing here needs you to change anything this week. It is a reason to discount confident percentages, including the ones in your own inbox.
Source arXiv, Roy Chowdhury and Garimella, “How People Use ChatGPT: Conversation-Level Evidence from India, Nigeria, Brazil, and Pakistan,” September 29, 2026 · arxiv.org
05 Research & evidence
AI models pick AI-written descriptions, but not their own
Which AI wrote your text does not change which AI picks it, so matching the assistant you hope will cite you buys nothing.
These models still prefer machine-written text to the human-written kind, and that older finding is the one that should worry you.
Dmitrij Żatuchin rebuilt the data behind a 2025 paper by Laurito and colleagues in the Proceedings of the National Academy of Sciences, which found that models choosing between two descriptions of the same product, paper or film prefer the machine-written one by a wide margin over what human judges do. The rebuild covers 21,828 valid trials and every cell matches the published figures.
The new question was narrower. Does a model favor text from its own model, beyond what its general preferences already predict? Across three datasets the own-model premium came in at +0.013 on products, -0.010 on paper abstracts and +0.054 on films, pooled at +0.019 with a p-value of 0.14. There is nothing there.
The caveat is in the paper, which is why it is worth trusting. The design had the power to catch a premium of 0.05 pooled at 97 percent, and the smallest effect it could detect reliably is 0.034. So it rules out a sizeable own-model bias and says nothing at all below about 0.04.
The part that holds is the awkward one. GPT-4’s product descriptions were chosen 77 to 95 percent of the time by every model tested. Set that beside Google’s guidance this week that unreviewed AI text shows little to no effort, and the two point in opposite directions. Nobody should pretend that question is settled.
What to do about it
Stop choosing a writing tool on the theory that a particular assistant will recognize its own output. The measurement says it does not, and that theory is sold more often than it is tested.
The finding that should guide you is the older one. Machine-written text reads as more selectable to a machine and, by Google’s own written standard, as less worthy to a human reviewer. Writing that a person with the expertise actually shaped is the only version that satisfies both.
Source arXiv, Żatuchin, “A Shared Taste for Model-Written Text: The Generator-by-Selector Matrices of AI-AI Bias Show No Detectable Own-Model Premium,” September 30, 2026 · arxiv.org