Inside 750,000 Claude Conversations: Anthropic Opens AI Usage Data to Independent Researchers

Table of Contents

Artificial intelligence companies know more than almost anyone else about how people actually use advanced AI systems. They can observe which tasks users bring to their models, where conversations fail, when people rely on AI for consequential decisions and how behavior changes as new models become more capable.

Most of that information remains inside the companies operating the platforms.

Anthropic is now testing a different approach. The company gave three external research groups access to privacy-protected, aggregate findings derived from real Claude conversations, allowing the researchers to develop their own questions and independently analyze the resulting data.

The pilot involved Stanford University’s Social and Language Technologies Lab, the University of Oxford’s Human Information Processing Lab and the nonprofit AI-evaluation organization METR. Each group designed a separate study using a unique sample of approximately 250,000 Claude conversations from April and May 2026.

Stanford and Oxford examined Claude.ai conversations. METR studied Claude Code interactions from consumer users who had opted to allow their information to be used for model improvement. Taken together, the three studies drew on approximately 750,000 conversations, although the groups did not receive the underlying transcripts or work from one combined dataset.

The early results challenge several common assumptions about AI use.

People are not limiting AI to harmless administrative work. More than half of the actionable conversations studied by Stanford involved consequential or high-stakes tasks. Human-led collaboration remained dominant, but users varied substantially in whether they learned from AI, adapted its work or allowed it to substitute for their own effort. Friction appeared in nearly half of conversations, although that friction frequently caused users to clarify their objectives and become more engaged.

Oxford’s preliminary analysis suggests that users’ apparent emotional experiences move alongside the model’s behavior. METR’s early findings indicate that newer coding models may deliver greater time savings than older versions.

The project is also a case study in the unresolved tension between AI transparency and user privacy. External researchers were allowed to ask independent questions of real-world usage data, but Anthropic retained control of the raw information, ran the analyses on its own infrastructure and reviewed every output before releasing it.

The Research Problem: AI Labs Control the Best Evidence About AI

Independent researchers generally have two imperfect ways to study real-world AI use.

They can analyze reports published by AI companies. These reports may draw from authentic platform behavior, but the companies decide which questions to ask, which findings to emphasize and which results to publish.

Alternatively, researchers can use public collections of AI conversations. Public datasets give researchers greater freedom, but the conversations may be voluntarily posted, scraped from limited sources or disproportionately focused on entertainment, experimentation and unusual behavior. They may not represent how paying customers use AI for professional work.

This creates an evidence gap. Policymakers are being asked to make decisions about AI’s effects on employment, education, mental health, productivity, privacy and professional services without access to the data most capable of answering those questions.

Anthropic’s pilot attempts to create a third option: outside researchers formulate the questions, but a privacy-preserving system analyzes the conversations inside Anthropic’s environment and releases only aggregate findings.

Key Facts About Anthropic’s Research Pilot

Category Pilot Details
Participating researchers Stanford’s SALT Lab, Oxford’s Human Information Processing Lab and METR.
Research relationship began February 2026.
Conversation window April-May 2026.
Approximate sample Approximately 250,000 unique conversations for each research group.
Stanford sample 249,834 Claude.ai conversations.
Oxford sample Approximately 250,000 Claude.ai conversations.
METR sample Approximately 250,000 Claude Code conversations, with a separate follow-up sample used for an addendum.
Plans represented Free, Pro and Max consumer plans.
Data excluded No Team, Enterprise or API customer conversations were included.
Raw conversation access External researchers never received raw transcripts, user identifiers or organization identifiers.
Analysis period Researchers received approximately 60 days to analyze the results and prepare initial findings.
Public data release Anthropic released the aggregate outputs received by the researchers, comprising 2,077 rows and approximately 5.63 MB at publication.

How Anthropic Insights Analyzes Conversations Without Releasing Them

The program uses a system now known as Anthropic Insights and previously called Clio.

A participating researcher writes a question about the conversations. Anthropic calls each question a “facet.” A facet might ask what task the user is attempting, how consequential the task appears, whether the user challenged the model or whether the interaction contains evidence of frustration.

Claude then applies that question to every conversation in the selected sample.

For multiple-choice and numerical questions, the responses can be counted directly. For open-ended questions, the system generates short summaries, groups similar answers into clusters and writes a name and description for each cluster.

Researchers receive the aggregate categories, counts, proportions and cluster descriptions. They do not receive the prompts, responses or identities associated with an individual conversation.

Clusters and intersections must contain a minimum number of conversations and users before they can be released. Results below those privacy thresholds are suppressed.

Anthropic also manually reviewed every cluster name and description before providing the data to the research groups. The reviews considered privacy, safety, confidential company information and whether the output accurately represented the methodology.

This arrangement resembles a controlled research environment more than a conventional dataset release. The data remains with the platform, approved analysis runs inside the platform and researchers receive filtered statistical outputs.

Case Study One: Stanford Finds People Entrusting AI With Consequential Work

Stanford’s Social and Language Technologies Lab conducted the most complete of the three studies released with the pilot.

The researchers analyzed 249,834 real-world Claude.ai conversations to investigate three questions:

  • How important or consequential were the tasks people brought to AI?
  • How much control and responsibility did people retain?
  • Where did human-AI collaboration experience friction?

The researchers developed 17 facets to measure issues such as task criticality, human agency, engagement with AI output, teaching behavior and conversational friction.

More Than Half of Actionable Conversations Involved Consequential Work

The Stanford study found that 56% of conversations containing actionable tasks were consequential or high-stakes.

Approximately 20% were categorized as ephemeral tasks with limited and easily reversible effects. At the other end of the scale, 12% were classified as high-stakes, meaning that the work could create substantial professional, legal, financial or personal consequences.

Legal and financial guidance produced the highest Task Criticality Index identified in the study, at 2.14. Nearly one-quarter of conversations in that category were classified as high-stakes.

This finding complicates the belief that people reserve AI for routine work while retaining consequential decisions for themselves. Users were willing to seek help with matters capable of affecting other people or producing outcomes that could be difficult to reverse.

The study does not establish that users acted on the responses, that the model’s advice was correct or that Claude made the final decision. It does show that users introduced high-consequence work into the conversational process.

Users Invested More Effort as the Stakes Increased

Higher-stakes conversations were longer and more involved.

Conversations concerning ephemeral tasks averaged 6.5 turns. High-stakes conversations averaged 12.4 turns—nearly twice as many.

Users also provided more specific instructions and added context incrementally as the importance of the task increased. That suggests people were not simply asking one question and accepting the first response. They often refined the interaction as the model produced information or exposed ambiguity in the original request.

However, users did not necessarily break high-stakes tasks into smaller and more manageable steps more frequently than they did with moderately important work. A person might provide substantial context while still expecting the model to manage a complex assignment as a whole.

That distinction matters because longer conversations do not automatically represent safer use. More interaction may indicate careful oversight, but it can also increase reliance on the model’s framing of the problem.

Humans Led the Collaboration in 72% of Conversations

The Stanford researchers created a five-level Human-AI Agency Scale ranging from AI-driven automation to human-controlled completion with limited AI assistance.

Approximately 72% of conversations fell into the category in which the human led the task and AI provided assistance.

This is an important counterweight to predictions that AI systems are already taking complete control of professional work. In the observed conversations, people generally established the direction and retained primary responsibility.

Yet conversational leadership is not the same as independent judgment. A person can decide what outcome to request while relying heavily on AI to determine the substance of the work.

The study therefore examined what users did with Claude’s output.

Adaptation Was More Common Than Verbatim Use

The researchers classified engagement into five general modes:

  • Using the output directly.
  • Reading the output to understand an issue.
  • Adapting the output to fit the user’s needs.
  • Critiquing or challenging the result.
  • Rejecting the output.

Adaptation was the most common form of engagement and reached approximately 60% for consequential work.

Direct use declined as the stakes increased. Only approximately 13% of high-stakes output was categorized as being used directly without modification.

In high-stakes conversations, users became more likely to read the response for understanding. This suggests that Claude often functioned as an advisor or explanatory resource rather than merely as a generator of finished work.

That pattern is encouraging, but it should not be confused with verification. Understanding what an AI system says does not prove that the user confirmed the underlying facts, sought qualified professional advice or recognized a persuasive but incorrect explanation.

Claude Displayed Teaching Behavior in 67% of Conversations

Stanford found some form of active teaching in 67% of the analyzed conversations.

Teaching behavior appeared across multiple domains rather than being concentrated in formal education. Software development and health or lifestyle topics each accounted for approximately 18% of the observed learning interactions. Business topics accounted for approximately 12%.

Mixed teaching methods appeared in 61% of teaching conversations. These interactions combined techniques such as conceptual explanations, examples and step-by-step instructions.

The findings suggest that AI’s effect on employment cannot be measured solely by counting completed tasks. AI systems may also change how people acquire knowledge, solve unfamiliar problems and enter fields in which they previously lacked expertise.

The study nevertheless found that access to teaching did not necessarily produce equal learning outcomes.

AI Readiness Was Associated With Different Types of Learning

The researchers compared behavior across 19 countries using a national AI-readiness measure incorporating factors such as technological infrastructure and human capital.

The overall presence of AI teaching remained relatively stable across countries. The difference appeared in how users engaged with the instruction.

Higher national AI readiness was positively associated with understanding-oriented engagement, with a reported correlation of r = 0.54 and p = 0.018.

Lower readiness was associated with greater emphasis on directly executing the model’s instructions, reflected in a correlation of r = -0.46 and p = 0.049.

These findings should be interpreted cautiously. Country-level correlations do not establish that national infrastructure caused an individual user’s behavior. The sample was also limited to Claude users and was not designed to represent the populations of the 19 countries.

Even so, the pattern raises an important policy question: Will AI close knowledge gaps by making high-quality instruction broadly available, or will it widen them because better-resourced users are more prepared to interrogate, understand and build upon what AI provides?

Friction Appeared in 49.7% of Conversations

The Stanford study found some form of collaborative friction in 49.7% of conversations.

Among conversations with friction, the researchers attributed:

  • 38.6% to model-initiated problems, including capability limitations, incorrect information or refusals.
  • 19.4% to user-initiated problems, such as ambiguous or insufficiently specified requests.
  • 42% to compounding failures in which user underspecification led the model to misunderstand the objective and the misunderstanding generated additional errors.

The prominence of compounding failures is especially significant. Human and model errors cannot always be separated cleanly. An unclear request can cause the model to adopt an incorrect interpretation, and a confident response can then lead the user further away from the original goal.

AI-governance programs that focus only on model accuracy may miss this interaction-level risk. A technically capable system can still fail when the surrounding workflow does not help users communicate objectives, detect misunderstandings and recover safely.

Users Attempted to Recover From Friction in 78.7% of Cases

Users did not simply abandon most conversations after encountering difficulty. They attempted some form of recovery in 78.7% of friction cases.

Recovery strategies included clarifying the request, correcting a particular mistake, supplying additional information and challenging the model’s reasoning.

Strategies that directly challenged the model’s reasoning were productive in 81.5% of the relevant cases, according to the study’s classification.

This led the researchers to distinguish harmful friction from productive friction. A misunderstanding can waste time or produce an unsafe result, but it can also force a user to clarify the task, inspect assumptions and become more engaged with the work.

The implication for product design is counterintuitive. The best AI interface may not be the one that makes every interaction feel effortless. In consequential settings, some carefully designed friction—asking for missing facts, exposing uncertainty or requiring confirmation—may improve outcomes.

Case Study Two: Oxford Examines Emotion and AI Behavior

The University of Oxford’s Human Information Processing Lab studied how people appeared to feel while using Claude and how those apparent states related to the model’s behavior.

The final paper was not available when Anthropic announced the pilot, so its findings remain preliminary.

The researchers identified several patterns in which human and model behaviors appeared together:

  • Warmer responses from Claude appeared alongside more positive user behavior.
  • Refusals or disagreements from Claude appeared alongside greater user pushback.
  • More eccentric model behavior appeared alongside greater intellectual engagement.
  • Straightforward helpful behavior appeared alongside apparent user satisfaction.

The researchers also compared patterns involving absorption, frustration and enjoyment with results from a separate study of ordinary web browsing. The relationships among those states appeared broadly similar, suggesting that using conversational AI may share important experiential characteristics with other forms of digital activity.

These findings do not demonstrate that Claude’s behavior caused a person’s emotional state. A frustrated user might elicit a different response from the model, just as a refusal might increase the user’s frustration. The direction of the relationship cannot necessarily be determined from the aggregate conversation data.

There is also a deeper measurement problem. The study asked an AI system to infer human feelings and evaluate AI behavior from text. Those inferences are interpretations, not clinical or objective measurements of emotion.

Anthropic removed one Oxford facet entirely because it suspected that a poorly worded question produced misleading cluster descriptions. That removal is important evidence of the method’s limitations: small differences in how a research question is phrased can materially alter what the analysis appears to find.

Case Study Three: METR Tests Whether Newer Coding Models Save More Time

METR used Claude Code conversations to estimate productivity gains from coding agents and examine whether those gains increased across model generations.

The data window was selected to span a Claude model release, allowing METR to compare use before and after the change.

The study’s preliminary findings suggest that newer models produced greater speed improvements than older models.

The approach compares the estimated time a task would have required without AI with the time apparently required when using different Claude models. Because the counterfactual time is not directly observable, the method depends in part on Claude estimating how long the work would have taken without assistance.

METR compared Claude’s time estimates with known completion times from a previous developer study. The model’s estimates correlated with the developers’ actual completion times, providing some support for using the estimates in the larger analysis.

Correlation does not establish that the estimates are sufficiently accurate for every kind of coding task. The conversations may also omit work occurring outside Claude Code, including planning, debugging, review and integration.

METR has not yet published its complete findings, and Anthropic appropriately characterized the productivity results as preliminary.

The organization is also exploring whether the same method can help estimate how much AI accelerates research work. That question could become increasingly important if AI systems begin contributing more substantially to the development and evaluation of future AI models.

The Researchers Were Independent—but the Data Environment Was Not

Anthropic structured the collaboration agreements to give the research groups authority over their questions, study designs, analyses and conclusions.

The company limited its contractual review rights to:

  • User privacy.
  • Information that could help people violate platform rules.
  • Anthropic’s confidential information.
  • Research accuracy.

The researchers were permitted to publish findings that were unfavorable or inconvenient to Anthropic.

That is more independence than researchers receive when merely analyzing a company-authored report. It is not the same as unrestricted access to an independently controlled dataset.

Anthropic selected the initial participants from organizations with which it already had relationships and which it trusted to work through the difficulties of a pilot. Anthropic funded the Insights runs, provided necessary credits, operated the analysis and reviewed outputs before releasing them.

Researchers could not inspect the underlying conversations to investigate unexpected categories, determine whether a classification was wrong or develop new questions after observing individual examples.

Anthropic also reviewed initial write-ups and requested corrections concerning descriptions of the methodology. The company says those corrections did not alter the researchers’ findings or conclusions.

The result is a meaningful but bounded form of independence. Outside researchers controlled the research questions and interpretation, while Anthropic controlled access to the underlying data and the mechanism through which it could be analyzed.

How Anthropic Protected User Privacy

No outside researcher received raw conversation text, user identifiers or organization identifiers. All raw information and computation remained on Anthropic’s systems.

Anthropic’s privacy controls included:

  • Aggregate outputs rather than conversation-level records.
  • Minimum conversation and user thresholds for clusters.
  • Suppression of small clusters and intersections.
  • Removal of identifying information.
  • Manual review of every cluster name and description.
  • Legal and privacy review of research configurations.
  • Testing proposed questions on the public WildChat dataset.
  • An external privacy audit before publication.

Anthropic developed a threat model informed by two established privacy concepts.

The first was the European approach to anonymization, including whether released data allows an attacker to isolate a record, link the information to outside data or infer a meaningful fact about an identifiable person.

The second was the UK regulator’s “motivated intruder” concept: assume a competent person actively attempts reidentification using publicly available information and reasonable resources.

Anthropic’s modeled harms included identifying a specific Claude user, learning a sensitive fact about an identifiable individual, attributing information to a small recognizable group and connecting an organization to confidential activity.

Imperial College London Red-Teamers Could Not Reidentify Users

Researchers from the AI Security and Privacy Lab at Imperial College London received the proposed public data and Anthropic’s threat model. They were given two weeks and $1,000 in Claude credits, with additional credits available if required, to attempt to reidentify users or otherwise violate the threat model.

The red-teamers did not reidentify any users and did not identify a violation of the defined threat model.

They did discover a separate weakness. Distinctive language in one cluster allowed them to connect it with a popular open-source project with a high degree of confidence.

That did not identify a person, small group or organization, so Anthropic did not classify it as a privacy violation. It nevertheless demonstrated that aggregate descriptions can retain recognizable features of tools used across many conversations.

Anthropic said future releases would use higher minimum-user thresholds, reduce the likelihood that cluster descriptions reproduce distinctive wording and repeat the auditors’ tool-linking technique as part of future review.

Anthropic Removed or Redacted a Small but Material Portion of the Results

Some clusters exposed attempts to violate Anthropic’s Acceptable Use Policy or Terms of Service.

Anthropic generally released categories describing what users attempted. It withheld information explaining how users allegedly bypassed safeguards because publishing those techniques could facilitate additional misuse.

The company reported the following interventions:

Study Clusters Redacted or Removed Conversations Represented
Stanford 1.9% of reviewed open-ended clusters 4.28% of conversations
Oxford 3.33% of reviewed open-ended clusters 3.85% of conversations
METR 1.8% of reviewed open-ended clusters 2.96% of conversations

The Oxford percentage excludes the separate facet removed in full because Anthropic believed a misphrased prompt produced misleading results.

These figures show the tension inherent in platform-controlled transparency. A company can have legitimate reasons to withhold information that compromises privacy or enables abuse. Every removal also means the external researcher is not working from a completely unfiltered view of platform behavior.

Transparency therefore requires disclosure not only of what was released, but also how much was removed, who made the decision and why.

The Most Important Methodological Limitation: Claude Is Interpreting Claude

Anthropic Insights does not simply count objective events. Many facets require Claude to interpret a conversation.

The model may be asked whether the user retained agency, whether the model displayed warmth, whether friction was productive or what emotion the user appeared to experience.

Those questions do not always have objectively correct answers. The system is effectively using an AI model to evaluate interactions involving an AI model.

This creates several potential problems.

Forced Categorization Can Create Findings That Are Not There

If a facet requires Claude to assign every conversation to a category, the model will produce a judgment even when the text contains insufficient evidence.

A question asking what problem occurred in every conversation will generate apparent problems even when the interaction was ordinary. Reliable research prompts need an explicit option for “nothing notable,” insufficient evidence or inapplicability.

Cluster Descriptions Can Exaggerate Harm

Anthropic’s clustering method is intentionally sensitive to concerning behavior. This may help researchers discover rare safety issues, but it can cause labels to emphasize the most alarming interpretation of ambiguous conversations.

A cluster’s title may sound more severe than the typical conversation assigned to it. In validation of the original Anthropic Insights system, approximately 3% of conversations were not clearly described by the cluster to which they were assigned.

That validation concerned conversation topics. Facets attempting to judge model behavior or user emotion did not receive the same validation, so the prior accuracy figures should not be treated as proof that those more subjective classifications are correct.

Open-Ended Clusters Are Not Precise Categories

Claude generates summaries that are converted into mathematical representations and grouped using a clustering algorithm. Small differences in wording can divide one topic among multiple clusters or combine distinct behaviors into a single category.

Two runs on the same data may produce different cluster structures. Each conversation is also assigned to only one cluster even when it contains multiple topics.

Anthropic estimated that approximately 10% of Claude.ai conversations covered more than one topic as of May 2026.

Open-ended clusters are therefore more appropriate for discovering themes than for making exact prevalence claims. More structured categorical questions are preferable when researchers need defensible percentages.

Testing on Public Data Did Not Fully Prepare Questions for Claude Traffic

Before applying a research configuration to Claude data, Anthropic ran it against WildChat, a public collection of human-AI conversations.

This allowed researchers to compare the system’s classifications with underlying conversations they were permitted to inspect.

The safeguard had an important limitation: WildChat contained a different mix of activity from Claude’s actual consumer traffic. It skewed more heavily toward casual and creative interactions.

Some questions that appeared effective on WildChat produced misleading categories when applied to Claude conversations. By that point, researchers could not repeatedly inspect and revise the questions against raw Claude data because doing so would undermine the privacy structure and require repeated review.

This is a form of domain shift. A research instrument validated on one population may behave differently when applied to another.

Anthropic is considering ways for researchers to develop and test their questions more effectively without gaining access to private conversations.

The Sample Does Not Represent Every Claude User

The pilot provides a substantial view of consumer activity, but it should not be described as a complete picture of Claude use.

The samples included Free, Pro and Max users. They excluded Team, Enterprise and API customers.

Enterprise use may differ significantly from consumer use. Companies can deploy AI through internal applications, governed workflows, API integrations and agentic systems that do not resemble ordinary chatbot conversations.

Claude Code data was limited to consumer users who opted into allowing their information to support model improvement. People who opt in may behave differently from those who do not.

Each study was also a one-time snapshot drawn from April and May 2026. AI products and user behavior change quickly as models, interfaces, pricing and capabilities evolve.

The results should therefore be understood as evidence about the sampled conversations—not universal measurements of all Claude users, all AI users or future behavior.

What the Pilot Means for AI Governance

The project demonstrates that responsible AI research access does not have to mean releasing raw conversation logs.

A controlled system can allow outside researchers to investigate authentic platform behavior while applying privacy thresholds, review controls and reidentification testing.

For AI providers, a mature external-research program should include:

  • Written criteria for selecting researchers and approving projects.
  • Contracts protecting publication independence.
  • Clear limits on company review rights.
  • Documented anonymization and aggregation controls.
  • Testing against identity, attribute, group and organizational disclosure.
  • Independent privacy red-teaming.
  • Transparent reporting of withheld categories and conversations.
  • Validation of model-generated classifications.
  • Disclosure of sample limitations and opt-in conditions.
  • Versioning so results can be tied to a specific model and period.
  • A process for qualified researchers to challenge company decisions.

Regulators and policymakers should also recognize that transparency can take different forms. Public access to raw data may be inappropriate when conversations contain legal questions, health information, source code, financial details and intimate personal disclosures.

Aggregate access can still provide meaningful oversight if researchers control the questions, methodological limitations are disclosed and the platform cannot suppress inconvenient findings merely because they are reputationally damaging.

What the Findings Mean for Businesses Using AI

The Stanford results are particularly relevant to employers adopting generative AI.

Employees are likely to bring consequential work to AI even if company policy imagines that the tools will be used only for low-risk drafting and administrative assistance. Legal, financial and professional guidance appeared among the most critical use cases in the study.

Organizations should therefore govern AI based on observed use rather than intended use.

A defensible program should address:

  • Which decisions employees may delegate to AI.
  • Which information may be entered into a model.
  • When professional or supervisory review is required.
  • How AI-generated claims should be verified.
  • Whether employees may use consumer accounts for company work.
  • How the organization records material AI involvement.
  • What happens when a model refuses, misunderstands or produces conflicting advice.
  • How employees are trained to challenge AI reasoning.
  • Whether AI use is building employee capability or replacing essential expertise.

The finding that users attempted recovery in most friction cases is promising. Organizations should not assume that every employee possesses the same ability to identify or correct an AI failure.

Training should focus on practical collaboration skills: defining the objective, supplying relevant context, recognizing uncertainty, checking sources, correcting mistaken assumptions and knowing when to stop using AI and seek qualified human judgment.

A Significant Step, but Not Independent Access in the Fullest Sense

Anthropic’s experiment provides outside researchers with access to evidence they could not otherwise obtain. It also preserves a strong privacy boundary around individual conversations.

The project deserves attention because it moves beyond asking AI companies to investigate themselves exclusively.

It should not be mistaken for unrestricted independent auditing.

The company still selects participants, controls the infrastructure, runs the queries, reviews the outputs and decides which privacy or safety concerns require redaction. Researchers cannot inspect the source material when a category appears implausible.

Those limitations may be unavoidable in an early privacy-preserving program. They should remain visible as the model expands.

The long-term test will be whether programs like this can support more researchers, adversarial questions, repeated studies and genuine disagreement without compromising user privacy.

Anthropic’s pilot shows that an AI company can open a meaningful window into real-world use without opening the underlying conversations. What can be seen through that window remains shaped by the platform, the research questions and the AI system performing the analysis.

Even with those constraints, the early picture is important: people are bringing consequential work to AI, humans often remain in charge, learning and delegation are occurring at the same time, and breakdowns are a routine part of collaboration.

Understanding those patterns is essential if businesses, researchers and regulators want to govern AI based on how people actually use it rather than how product descriptions say it should be used.

Written by: 

Online Privacy Compliance Made Easy

Captain Compliance makes it easy to develop, oversee, and expand your privacy program. Book a demo or start a trial now.