Can You Sell Your Company’s Data to Train Frontier AI? The Privacy Risks Behind the New Enterprise Data Market

Table of Contents

The cleanest legal answer sounds simple: if information is truly anonymous, privacy law generally stops treating it as personal data. An IAPP-published analysis by Baker McKenzie attorneys made essentially that point while examining whether companies can use customer data to train AI systems: genuinely anonymized information is not subject to the same privacy-law restrictions as personal information. The problem is hidden inside one word: genuinely. How does a company know that years of Slack messages, CRM records, emails, project histories, documents, customer-support conversations and operational data have actually crossed the line from personal data into anonymous data?That question has suddenly become worth real money. AI data company micro1 is openly offering companies substantial payments to license operational datasets that can be used to train and evaluate frontier AI systems. The company is looking for the material that businesses spent years creating: internal documentation, standard operating procedures, CRM information, project histories, workflows, customer-support processes, engineering records, internal communications and the decision-making patterns that explain how real businesses actually work. The economic proposition is compelling because frontier models have already consumed enormous amounts of public information. What they increasingly need is information that is not publicly available: real examples of how companies make decisions, resolve exceptions, communicate internally and execute complex work.

From a privacy perspective, however, this creates a much harder question than whether your company can make several hundred thousand dollars from an otherwise dormant data asset. Most of this information was never collected because the company planned to sell it into an AI-training pipeline. Customers gave companies information because they wanted a product or service. Employees wrote Slack messages because they were doing their jobs. Vendors exchanged emails because they were negotiating or performing contracts. Prospects entered information into CRMs because they were considering a purchase. Turning those records into training material for somebody else’s AI system is a new use of the information, and potentially a very different one.

Captain Compliance would not recommend treating customer or employee personal data as a newly discovered revenue stream simply because a vendor says it can be de-identified. There may be circumstances where an enterprise can safely license operational knowledge after proper transformation, legal review and technical validation. But there is an enormous difference between selling your company’s know-how and selling a dataset containing the people who happened to create, purchase from, work for or communicate with your company. Businesses considering these programs need to understand that difference before the data leaves their environment.

We transform one real-world entity into one synthetic identity, maintaining the same mapping across all sources

Why Frontier AI Companies Suddenly Want Your Internal Data

The economics behind this market make sense. Publicly available training data has become less differentiating. Frontier AI developers increasingly want datasets showing how knowledgeable people actually perform work. A public website can explain accounting. An internal accounting workflow shows how a controller reconciles a difficult account, handles an exception, escalates an anomaly and documents the final decision. A textbook can describe contract review. A company’s actual contract history shows which provisions lawyers negotiate, what language they replace them with, what gets escalated and which risks the business ultimately accepts.

Micro1 describes enterprise data in almost exactly these terms. Its research says that enterprises possess datasets containing decision-making, communications, tools, project execution and long-running relationships that are difficult to reproduce from public information. The company says it works directly with enterprises to acquire and structure large-scale datasets for training frontier AI systems. The materials can include Slack and email exports, documents, databases, spreadsheets, images, audio, video and other internal systems.

That is valuable because training a capable enterprise agent requires more than teaching it what a document looks like. The model needs examples of relationships across documents and time. Who reports to whom? Which customer issue led to which support escalation? Which engineering bug caused which remediation? What information did a salesperson consider before offering a discount? Which clause caused legal to reject a contract? How did a complicated project move from one department to another?

This also explains why simple redaction is unattractive for AI training. If every person becomes “[NAME],” every company becomes “[COMPANY],” every email becomes “[EMAIL]” and every relationship disappears, much of the informational value disappears with it. A model trained on heavily redacted material may see the words but lose the organizational structure that made the data useful in the first place.

Micro1’s Technical Answer: Replace the Identity, Preserve the Workflow

Micro1’s September 2026 research paper proposes a more sophisticated approach called identity-preserving transformation. Instead of simply deleting identities, its flow-transform system attempts to replace each real person with a coherent synthetic person while maintaining relationships across the entire dataset. If Jane Smith appears in an email, a spreadsheet, a Slack conversation and a CRM record, the objective is not to replace Jane independently four different ways. The system maps the real Jane to one synthetic identity and uses that synthetic identity consistently throughout the transformed corpus.

This is technically important. Consider an internal sales dataset in which Jane Smith reports to Robert Chen, manages the Acme account and corresponds repeatedly with Acme procurement officer Thomas Brown. A conventional anonymizer might produce “PERSON 1,” “PERSON 2” and “PERSON 3” inconsistently across documents. The resulting corpus would lose the relationships that teach an AI model how the business actually operates. Micro1 instead attempts to preserve the graph while changing the nodes: synthetic Jane still reports to synthetic Robert, manages the corresponding account and interacts with synthetic Thomas.

The system also tries to replace related identity attributes coherently. A person’s name, email address, username, employer and other permitted attributes can be transformed as a bundle rather than as unrelated fields. That is a meaningful improvement over naïve regular-expression masking, particularly across a large collection of documents where one person may appear hundreds or thousands of times.

Micro1 argues that anonymization should therefore be evaluated as an entire corpus-transformation problem rather than merely a PII-detection problem. Its framework measures privacy, utility, fidelity, coverage and coherence. That is a useful way to think about the problem because a system can score well at detecting obvious names and email addresses while still fail spectacularly at protecting a real enterprise dataset.

The Benchmark Results Are Strong. They Are Not a Guarantee of Anonymity.

Micro1 reports impressive results for its flow-transform 1.0 system. On its enterprise de-identification benchmark, its Transformation Quality Index rises from 74.9 for NVIDIA’s NeMo Anonymizer baseline to 84.1 without agentic review and 88.9 with agentic review. Its public PrivacyBench evaluation reports a 96.00% F1 score for PII detection, while other evaluations report improvements over competing approaches on Nemotron-PII and AI4Privacy.

We transform one real-world entity into one synthetic identity, maintaining the same mapping across all sources

Those numbers are relevant evidence that the technology is sophisticated. They do not prove that any particular enterprise dataset becomes legally anonymous. Micro1 itself acknowledges an important limitation: its primary enterprise benchmark is synthetic, structured and limited to tabular information. Complex formats, multimodal information, ambiguous identity linkage and artifact reconstruction remain areas where additional work is needed.

This distinction matters enormously. A 96% benchmark F1 score is a model-performance metric. It is not a legal conclusion that a delivered dataset has zero personal information. Precision and recall are statistical measures across a test corpus. A company cannot take an average benchmark score and assume its own dataset is anonymous.

Real corporate data is considerably uglier than a benchmark. A PDF can contain someone’s name in a footer. An Excel workbook can contain hidden worksheets, cell comments, formulas and author metadata. A screenshot can display a customer email address. Source code can contain developer names and internal hostnames. Jira tickets can contain pasted credentials. Slack messages can contain screenshots, phone numbers and references that identify people indirectly. Audio contains voices. Video contains faces. Calendar invites expose relationships. Email headers contain addresses even after names inside the message body have been removed.

Micro1’s own paper recognizes many of these challenges. It discusses hostile formatting, aliases, OCR noise, obfuscated emails, Unicode characters, Base64, missing values and contextual information that only becomes identifying when surrounding facts are considered. It also identifies Photoshop, Figma and other complex formats as difficult transformation problems, and notes that replacing faces and voices creates a separate set of technical and ethical issues.

That is exactly why “we de-identify it” cannot be the end of a privacy review.

De-Identified, Pseudonymized and Anonymous Are Not the Same Thing

This terminology causes persistent problems. A dataset can be de-identified in an ordinary engineering sense without being anonymous in the legal sense. It can also be pseudonymized while remaining fully subject to privacy law.

Under the GDPR, pseudonymization is a safeguard. It does not automatically remove information from the GDPR. If additional information can reconnect the records to an identifiable person, the information generally remains personal data. Truly anonymous information is different: the person must no longer be identifiable using means reasonably likely to be used.

The European Data Protection Board has made the standard particularly relevant to AI. In its Opinion 28/2024 on AI models, the EDPB said anonymity cannot simply be presumed because a model or dataset has undergone a privacy process. It requires a case-by-case assessment. For an AI model to be treated as anonymous, the probability of directly or indirectly identifying people whose information was used, and the probability of extracting their personal information through model queries, must be insignificant when considering means reasonably likely to be used.

That is a much more demanding question than asking whether the names were removed.

The Relationship Graph Itself Can Be Identifying

There is a fascinating tension inside identity-preserving transformation. The feature that makes the transformed data more useful for AI training can also be part of the privacy analysis.

Micro1 intentionally tries to preserve relationships. The same employee remains connected to the same synthetic coworkers, projects, departments and events. That is valuable because the model learns organizational behavior instead of receiving a pile of unrelated redacted documents.

But relationships themselves can be identifying. Imagine that a transformed dataset no longer contains the name of a company’s chief financial officer. It still shows that one synthetic person reports to the CEO, approved a particular acquisition, participated in a publicly known transaction in March, corresponded with three identifiable law firms and managed a six-person finance team. Someone with outside knowledge may not need the original name field to infer who the synthetic person represents.

This is known generally as linkage or re-identification risk. Individual data points that appear harmless can become identifying when combined with other information. The problem becomes especially difficult with high-dimensional enterprise data because each person produces a distinctive trail of communications, projects, reporting relationships, transactions and decisions.

A sophisticated anonymization program therefore needs to test more than direct identifiers. It needs to ask whether people can be singled out through combinations of attributes, relationships and external knowledge.

There Is Another Crucial Moment of Risk: Before the Data Is De-Identified

Even if the final dataset is anonymous, somebody has to create it.

Micro1’s research describes ingesting raw enterprise information, segmenting it, detecting PII, resolving identities, transforming those identities, reconstructing the files and validating the results. During at least part of that process, the system necessarily needs access to the source information in order to identify what should be transformed.

This means companies should not focus exclusively on the downstream sanitized dataset. The raw-data ingestion stage may be the highest-risk stage of the entire transaction. Before transformation occurs, the corpus may contain exactly what the company is trying to protect: customer names, employee communications, financial information, account information, passwords, health information, privileged communications, contractual terms, confidential product information and trade secrets.

One particularly important technical question concerns agentic review. Micro1’s paper explains that an optional LLM-based agent can examine candidates together with surrounding context and the shared identity store. The agent can confirm or reject detections, merge aliases, resolve malformed identifiers and perform more open-ended editing. That may materially improve detection performance, but a privacy and security review should determine exactly where that agent runs and which model receives the context. Is everything processed inside an isolated environment? Does any raw content reach an external model API? Are prompts retained? Can model providers use inputs for training? Are human reviewers exposed to raw information? Which countries can they access it from?

Those questions are not answered by an overall benchmark score. They have to be answered contractually and architecturally for the specific engagement.

The Hardest GDPR Question May Be Purpose, Not Security

Assume for a moment that the technology works perfectly and the final dataset contains no personal information. There is still an earlier legal question: was the company allowed to process people’s personal information for the purpose of creating the dataset that it intends to sell?

Article 5 of the GDPR establishes purpose limitation. Personal data must be collected for specified, explicit and legitimate purposes and generally cannot later be processed in a manner incompatible with those purposes. Further processing is not automatically prohibited, but organizations need to analyze compatibility, legal basis, reasonable expectations and applicable safeguards.

This matters because the original purpose for most enterprise data was not “sell this information to help train third-party frontier AI systems.” A customer gives an ecommerce company an email address so it can fulfill an order and communicate about the relationship. An employee writes internal messages because the employee is performing work. A prospect enters a phone number because the prospect wants to speak with sales. A vendor provides bank information because it needs to be paid.

Turning that information into a monetizable training corpus is a new commercial use. GDPR analysis asks whether the new processing is compatible with the circumstances under which the information was originally collected, what the individual reasonably expected, the nature of the data, the consequences of further use and what safeguards have been implemented.

This does not mean AI training is automatically unlawful. The EDPB recognizes that legitimate interests can in some circumstances support development and deployment of AI systems, provided the purpose is legitimate, the processing is actually necessary and the interests or fundamental rights of data subjects do not override the controller’s interests. But “we can get paid for the dataset” is not itself the end of that analysis.

An IAPP-published article on legitimate interests made another useful point: “training AI” is not a sufficiently meaningful purpose by itself. An organization should identify what the AI is being trained to accomplish. Detecting fraud, improving cybersecurity and developing a general-purpose commercial model present materially different interests and reasonable-expectation analyses.

If the Data Is Truly Anonymous, GDPR Changes Dramatically. Getting There Still Involves GDPR.

This creates a distinction companies frequently miss. If the final corpus has genuinely been rendered anonymous, the resulting information may fall outside the GDPR because it no longer relates to identifiable individuals. But the activity used to transform the raw personal data into anonymous information is itself processing.

In other words, a company cannot simply say, “The output is anonymous, therefore we never had a GDPR issue.” The organization still had personal information at the beginning of the pipeline. It copied it, analyzed it, classified it and transformed it. Those operations need a lawful basis and appropriate governance while the information remains personal.

The distinction becomes particularly important if raw European data is sent to an outside company or across borders before anonymization. If the enterprise first transforms information inside its own controlled European environment and only the demonstrably anonymous output is transferred, the analysis may be considerably different from sending raw employee and customer records to a U.S. vendor and asking the vendor to anonymize them after receipt.

California Creates a Different Problem: Is This a Sale?

California privacy law approaches the issue differently but creates an equally important risk. The CCPA defines a sale broadly around transferring personal information to a third party for monetary or other valuable consideration. California consumers have rights to opt out of covered sales and sharing, while additional restrictions apply to sensitive personal information and minors.

If an enterprise transfers personal information to an AI-data company specifically in exchange for money, the word “sale” deserves serious legal attention. The transaction does not stop being a potential sale simply because the receiving party eventually intends to anonymize the information.

An IAPP-published Baker McKenzie analysis makes this distinction particularly clear in the vendor context. If a processor uses customer personal data only under the controller’s instructions, the processor relationship can potentially be maintained. If the customer instead grants the vendor permission or a license to use personal information for the vendor’s own purposes, including development of its own AI offering, the arrangement can affect processor status and may look much more like the customer selling or disclosing personal data for an independent purpose.

This is one reason the contract architecture matters. “Use our information solely to de-identify it for us and return the output” is fundamentally different from “we license you our corporate data so you can create training datasets for your AI customers.”

California Does Exclude Deidentified Information, But Again: Prove It

The CCPA excludes qualifying deidentified and aggregate information from the definition of personal information. California’s statutory framework also emphasizes technical and organizational controls designed to prevent reidentification in contexts where deidentified information is used.

That leads to the same practical problem found under GDPR. Declaring a dataset deidentified does not necessarily make it so. Companies should have an evidentiary basis for that conclusion and should contractually prohibit downstream recipients from attempting to reidentify people.

This is particularly important because the purpose of frontier model training is to extract patterns. If the corpus preserves rich context, long-term relationships and unusual sequences of events, the company needs to test whether the useful patterns being intentionally preserved also preserve too much information about the underlying individuals.

Not All Enterprise Data Carries the Same Risk

Dataset Privacy Risk Other Major Risk Captain Compliance View
Company-authored SOPs, generic templates and process documentation containing no personal information Lower Trade secrets, IP, competitive know-how Potentially viable after legal and IP review
Project histories and operational records with employee identities removed Medium Reidentification, confidential strategy, customer references Requires validation beyond simple PII masking
CRM records, support tickets and customer communications High Contractual restrictions, sale/sharing obligations, customer confidentiality Do not treat as ordinary monetizable data
Slack, email and employee communications High Employee privacy, privilege, investigations, HR records, trade secrets Requires exceptionally careful scoping and transformation
Health, financial, biometric, children’s or other regulated sensitive data Very high Sector-specific laws and heightened consent requirements Generally avoid absent a very specific lawful architecture

PII Is Only One Category of Information You Can Accidentally Sell

One of the biggest mistakes in this discussion is treating PII removal as synonymous with making an enterprise dataset safe.

Imagine a law firm perfectly removes every client name, lawyer name, email address and telephone number from ten years of internal documents. The resulting corpus can still contain privileged legal analysis, litigation strategy and confidential settlement positions. A company can remove all personal information from engineering documents while still exposing source-code architecture, vulnerabilities and its product roadmap. A retailer can anonymize customers while leaving pricing strategy, margins, supplier negotiations and fraud-detection logic intact.

Micro1 openly markets the value of precisely this operational knowledge. That is not a criticism; it is the product. The data is valuable because it contains information competitors and public datasets do not have.

But executives should therefore treat an enterprise-data partnership as an intellectual-property transaction in addition to a privacy transaction. The privacy team can tell you whether Jane Smith is identifiable. It cannot decide whether the workflow Jane created is a trade secret the company should teach to the next generation of AI systems.

Customer Contracts May Matter Before Privacy Law Does

Many businesses have contractual promises that are stricter than privacy statutes. A SaaS provider may promise customers that customer content will be processed solely to provide the service. A consulting company may have confidentiality obligations covering client work product. A law firm may have professional-secrecy and privilege obligations. A healthcare company may have HIPAA obligations and business-associate agreements. A financial institution may have GLBA requirements. A company may have NDAs with nearly every major customer and vendor.

De-identification does not necessarily solve all of those issues. The agreement may prohibit using the information to train unrelated models, disclose confidential business information or derive generalized products from customer content even if names are removed.

The contract review therefore needs to happen before the corpus is assembled. Otherwise, a company could spend enormous effort anonymizing a dataset that it never had the contractual right to commercialize in the first place.

Employee Data Is Especially Complicated

Internal communications are among the most valuable forms of enterprise training data because they show how work actually happens. They are also some of the most privacy-sensitive.

Slack and email contain far more than business processes. Employees discuss medical leave, compensation, performance, family issues, grievances, workplace investigations, accommodations, political views, interpersonal conflicts and countless incidental details about their lives. One employee may include another person’s information without that person’s knowledge. A private channel may contain an investigation. An email chain may contain legal advice.

Consistent synthetic identities do not automatically remove those risks. If the substantive contents remain, a dataset can continue to contain sensitive personal information even after the sender’s name changes. “Synthetic Employee 418 told HR that she had been diagnosed with breast cancer and would begin chemotherapy on June 4” remains health information about a potentially identifiable person if surrounding facts make the employee reasonably identifiable.

For employee communications, de-identification has to consider the content, not merely the account identifier.

Model Training Creates a Different Kind of Irreversibility

Traditional data sharing has an intuitive deletion model. Company A sends Company B a database. If the contract ends, Company B deletes its copies.

Model training complicates that assumption. Once examples influence model parameters, deleting the source files does not necessarily reverse the training. Machine unlearning remains an active research area rather than a universally reliable operational control.

An IAPP analysis published in 2026 described the practical distinction this way: organizations can delete a person’s source training data and exclude it from future training runs, but they do not generally retrain an entire model every time an individual requests erasure. That makes lawful collection and training-stage governance particularly important. You do not want the first serious discussion about whether information should have been used to happen after it has already entered a frontier-model training pipeline.

De-identification can substantially reduce this risk because properly anonymous training data no longer represents an identifiable individual. But once again, that makes the quality of the de-identification process central. If direct or indirect personal information survives transformation and is then learned by the model, downstream remediation is substantially harder than deleting a row from a database.

Memorization and Extraction Need to Be Part of the Threat Model

AI privacy risk is also different from conventional database privacy because models can memorize portions of training information. Researchers have demonstrated that large language models can under some circumstances reproduce memorized training sequences or leak personally identifiable information.

Micro1 itself cites research on memorization and PII leakage in its technical paper. That is an appropriate acknowledgment of the threat. A defensible process should therefore evaluate not merely whether transformed files look deidentified before training, but whether the downstream model can reproduce sensitive sequences, rare information or information that should not have survived transformation.

This becomes particularly relevant for highly unique enterprise events. Common workflow instructions may generalize harmlessly. A unique legal settlement, unusual security incident or one-of-a-kind customer problem can be much easier to associate with a real company or individual.

What Would a Defensible Enterprise Data Deal Actually Look Like?

A lower-risk arrangement would begin by separating corporate know-how from personal information before transfer. The enterprise would identify which systems and datasets are in scope, map the categories of people represented inside them, identify contractual and sector-specific restrictions, determine what processing purpose justifies the transformation and decide what information must never leave the company.

Ideally, the highest-risk transformation would occur inside a controlled environment before the usable dataset reaches the ultimate AI developer. Direct identifiers, sensitive categories, privileged materials, credentials, customer secrets and unnecessary metadata would be removed or transformed. The company would independently test for reidentification rather than relying exclusively on the vendor that financially benefits from approving the dataset.

The agreement would define whether micro1 or another intermediary acts as a processor during raw-data handling, when that role changes, what independent purposes are permitted, which subprocessors can access information, where processing occurs, whether human reviewers see raw records, what LLMs participate in transformation, whether those model providers retain prompts and what happens after the transformation is complete.

Downstream use should also be specific. A company should know whether one customer receives the resulting dataset or whether it can become an off-the-shelf corpus licensed repeatedly. It should know whether the dataset is used for pretraining, post-training, reinforcement learning, evaluations or synthetic-environment creation. Those are materially different uses.

The contract should prohibit reidentification, require appropriate security, define retention periods and deletion requirements, provide incident-notification obligations and address what happens if previously unidentified PII is discovered after delivery. The seller should also obtain meaningful representations about onward transfers rather than assuming the intermediary remains the final recipient.

Ask Which Frontier Model Gets Your Data

This is an area where companies should demand precision. Micro1 describes itself as a data research company serving frontier AI development and says it works with leading AI organizations. Its public enterprise-data materials emphasize that companies contribute operational data to help train and evaluate frontier systems.

That does not establish that any particular enterprise dataset will be delivered to OpenAI, Anthropic, Google, Meta, xAI or any other named frontier lab. Public marketing materials do not identify the ultimate customer for every partnership. Companies should not fill in that blank themselves.

If you are selling or licensing the dataset, ask. Which entities can receive it? Can it be sublicensed? Can the same corpus be licensed to multiple model developers? Can it leave the United States? Can it be used by foreign model developers? Does the agreement restrict specific jurisdictions or competitors? Are model providers permitted to retain training artifacts indefinitely?

The phrase “frontier AI” describes an industry. It is not a data-processing specification.

The Most Important Question May Be Whether You Should Sell It at All

Micro1’s economics are intentionally attractive. The company publicly markets six- and potentially seven-figure enterprise data partnerships. For a company sitting on years of archived operational information, converting those records into revenue can look like found money.

It is not found money. The company is being paid because the dataset contains something valuable that the buyer cannot easily obtain elsewhere.

That should cause the board, CEO, general counsel, chief privacy officer, CISO and relevant business owners to ask what exactly that value is. If the value comes from your proprietary decision-making process, you are monetizing intellectual property. If it comes from customer interactions, you may be monetizing data generated through customer relationships. If it comes from employee communications, you are monetizing a record of how your workforce thinks and operates.

Those can be legitimate business decisions. They should not happen accidentally because someone saw a $500,000 data-partnership estimate on a website.

Is it okay to Sell Know-How, Not People?

The IAPP-published legal analysis contains the most important sentence in this entire debate: truly anonymized data can fall outside ordinary privacy-law restrictions. The European framework says essentially the same thing. California excludes properly deidentified information from personal information. There is a legitimate path for organizations to transform information into something that can be analyzed and used without exposing the people who created it.

But the phrase “truly anonymized” is carrying nearly the entire legal and technical burden.

A company should not assume that swapping names for synthetic names makes a corpus anonymous. It should not assume that a 96% detection benchmark proves its own dataset is safe. It should not assume that because the ultimate recipient gets deidentified data, there were no privacy obligations during ingestion and transformation. It should not assume that removing PII gives the company the contractual right to commercialize confidential customer material. And it should not assume that information is safe simply because the buyer calls the process privacy-preserving.

The strongest version of this market could be very valuable. Enterprises have accumulated extraordinary amounts of institutional knowledge, and much of that knowledge can potentially be separated from the identities of the people who generated it. If a company wants to license its SOPs, processes, templates, workflows and institutional know-how after those assets have been carefully stripped of personal information, confidential customer data and restricted content, that can be a legitimate commercial decision.

What privacy teams should resist is a simpler proposition: “We already have all this customer and employee data, so why not sell it and anonymize it afterward?” Most of that personal information was not collected for resale into frontier AI training. GDPR purpose limitation, U.S. state privacy laws, consumer expectations, employee rights, contractual restrictions and data-minimization principles all exist precisely because possession of data does not automatically create an unrestricted right to invent new purposes for it.

If you want to sell your company’s systems and know-how, know what you are selling. Separate the institutional intelligence from the identities embedded inside it. Determine whether you have the right to commercialize the underlying material. Transform sensitive information before unnecessary disclosure. Independently validate the output. Understand who ultimately receives it. Restrict reidentification and secondary use. Preserve an audit trail showing how the conclusion that the dataset was anonymous was reached.

Because the central privacy question is not whether micro1, or any other AI-data company, has good anonymization technology. The more important question belongs to the enterprise handing over the information:

Can you prove that the data you are selling is no longer the data your customers, employees and partners trusted you to protect?

Until the answer is yes, “de-identified” should be the beginning of the privacy review, not the end.

Written by: 

Online Privacy Compliance Made Easy

Captain Compliance makes it easy to develop, oversee, and expand your privacy program. Book a demo or start a trial now.