Italy Fines IQVIA €7 Million Over Health Data: Why “Anonymous” Data Can Still Be Personal Data

Table of Contents

Italy’s data protection authority has imposed a €7 million fine on IQVIA Solutions Italy after finding that a large health-data analytics database contained information that was not truly anonymous.

The case matters far beyond one company.

It goes directly to one of the most consequential questions in modern privacy compliance:

When does supposedly anonymized data stop being personal data?

IQVIA maintained a database containing health information relating to roughly 1 million patients from about 800 general practitioners. The database was used for studies commissioned in part by pharmaceutical companies and included information such as year of birth, sex, diagnoses, symptoms, prescriptions, exams, vaccinations and location-related data. Italy’s Garante concluded that the data could still be linked back to individuals using reasonable means.

That finding changed everything.

If the data had been genuinely anonymous, much of the GDPR would no longer apply to it.

If the information remained personal data, however, the processing required a lawful basis, transparency, retention controls, security measures, accountability and, because health data is a special category under the GDPR, an additional legal condition for processing.

According to the Garante, IQVIA did not satisfy those requirements.

The enforcement action is a significant reminder that replacing a name with a code is not the same thing as anonymization.

What IQVIA Was Doing

IQVIA Solutions Italy is part of a multinational group operating in health-data analytics and clinical research.

According to the Garante, the company created what it called the LPD database using information collected from approximately 800 general practitioners.

The database contained information relating to around 1 million patients.

The data was then used for observational studies and other healthcare analyses, including research commissioned by pharmaceutical companies.

Health-data research itself is not inherently unlawful.

The problem identified by the regulator was how the information was collected, structured and protected.

IQVIA maintained that the patient data had been anonymized.

The Garante disagreed.

Why the Garante Said the Data Was Not Anonymous

The central issue was the patient identifier.

According to the authority, each patient was assigned a code that allowed the same person to be followed over time.

That is useful for research.

It is also precisely what makes anonymization difficult.

The code was combined with a detailed set of attributes, including:

  • year of birth;
  • sex;
  • diagnoses;
  • symptoms;
  • prescriptions;
  • laboratory and diagnostic examinations;
  • vaccinations; and
  • location information.

The Garante concluded that this combination made it possible to isolate individual patients and potentially reidentify them using reasonably available means.

That distinction is critical under the GDPR.

Anonymous information falls outside the GDPR when an individual is no longer identifiable.

Pseudonymous information does not.

A dataset may remove names, email addresses and tax identification numbers and still remain personal data if an individual can reasonably be singled out or reidentified.

That appears to have been the central mistake in the IQVIA case.

Pseudonymization Is Not Anonymization

The terms are frequently confused.

Pseudonymization means replacing or separating direct identifiers so information cannot immediately be attributed to someone without additional information.

For example:

Maria Rossi

might become:

Patient 847193

That can materially reduce privacy risk.

But if Patient 847193 can be tracked across years of medical records, locations, prescriptions and diagnostic events, the underlying person may still be identifiable.

The GDPR still treats pseudonymized data as personal data.

True anonymization requires a much higher threshold.

The data has to be transformed so that the individual is no longer identifiable by means reasonably likely to be used.

That assessment cannot focus only on whether a name appears in the database.

It has to examine the dataset as a whole.

The Mosaic Problem

The IQVIA case is a good example of what could be called the mosaic problem.

Individual pieces of information may appear innocuous.

Year of birth alone is not necessarily identifying.

Sex alone is not.

A diagnosis alone may not be.

A geographic area alone may not be.

But combine:

a year of birth;

a particular location;

a rare diagnosis;

a prescription;

a vaccination history;

and a sequence of medical events over time,

and the resulting profile can become highly distinctive.

This matters increasingly because modern analytics works precisely by combining information.

A company may believe it has removed obvious identifiers while preserving a sufficiently detailed behavioral or medical profile that reidentification remains possible.

The privacy analysis therefore cannot end with:

“We removed the names.”

The correct question is:

“Can someone still identify or single out the person represented by this record?”

The Database Went Back to 2001

The Garante identified another major problem: retention.

According to the authority, IQVIA had not established appropriate retention periods, and some of the information in the database dated back as far as 2001.

That means some records had potentially been retained for roughly 25 years.

This is an important part of the case because privacy compliance is not only about whether information may be collected.

Organizations also need to determine how long it should remain available.

Under the GDPR’s storage-limitation principle, personal data should generally not be kept in identifiable form longer than necessary for the purpose for which it is processed.

For research databases, retention questions can be complex.

Longitudinal research may legitimately require long periods.

But “research” does not automatically create an unlimited retention period.

Organizations still need documented reasoning explaining why the information is necessary and how long it should remain.

More Than 3,000 Patients Appeared With Directly Identifying Information

The enforcement action became even more serious because the database did not contain only coded records.

The Garante said information directly identifying more than 3,300 patients also entered the database, including names, tax codes, addresses and contact information.

For more than 3,000 of those people, the identifying data appeared alongside health information.

The authority linked this problem to free-text fields in software used by participating physicians.

That detail is particularly important from a privacy-engineering perspective.

Structured databases may be designed carefully to exclude direct identifiers.

Free-text fields can destroy that design.

A physician might enter:

“Patient Maria Rossi called regarding insulin prescription.”

Or:

“Discussed diagnosis with patient’s husband, Carlo.”

If that free text is ingested into a supposedly anonymized dataset, direct identifiers can move with it.

This is a common problem in data analytics and AI systems.

Engineers may sanitize structured fields while forgetting that uncontrolled text can contain practically anything.

The Investigation Also Followed a Data Breach Notification

The regulator’s investigation began after inspections conducted in April 2025 and was combined with proceedings relating to a personal-data breach that IQVIA itself had notified.

According to the formal decision, investigators found that an add-on installed in physicians’ software extracted supposedly anonymized information but also included free-text fields containing directly identifying patient and doctor data.

The Garante found that appropriate technical and organizational safeguards had not been implemented to prevent that outcome.

This part of the case is worth emphasizing.

Privacy-by-design failures frequently happen at the boundary between systems.

One database may be well designed.

One export function may not be.

One unstructured field can bypass an otherwise carefully constructed anonymization process.

That is why privacy engineering has to examine the entire data flow rather than just the final database.

The Garante Found No Appropriate Legal Basis

Health information receives special protection under Article 9 of the GDPR.

Processing it generally requires both:

a lawful basis under Article 6;

and a valid Article 9 condition permitting processing of special-category data.

The Garante concluded that IQVIA processed health information without an appropriate legal basis.

That conclusion followed directly from the anonymization finding.

If the information were truly anonymous, the GDPR legal-basis analysis would largely fall away.

Once the regulator concluded that the records remained identifiable personal data, IQVIA needed to demonstrate the legal authority for processing them.

This is why incorrectly classifying data as anonymous creates such large downstream compliance consequences.

The classification affects everything else.

Patients Were Not Adequately Informed

The authority also concluded that patients had not received adequate information concerning the processing.

That implicates one of the GDPR’s most basic requirements: transparency.

People generally need to know who is processing their personal information, what information is being used, why it is being processed, how long it is retained and who receives it.

Again, the anonymization theory appears to have shaped the operational model.

If a company concludes that information has been anonymized, it may believe individual privacy notices are unnecessary.

If regulators later decide that the information was never anonymous, the company can suddenly find itself facing not just an anonymization problem but also a transparency problem.

No DPIA Was Conducted

The Garante also said IQVIA had not performed the required data protection impact assessment.

That is another striking fact.

A database involving:

approximately 1 million people;

longitudinal health histories;

medical diagnoses;

prescriptions;

testing;

vaccinations;

location information;

and healthcare research

is precisely the kind of processing that should cause a privacy team to ask whether a DPIA is required.

Under Article 35 of the GDPR, DPIAs are required where processing is likely to result in a high risk to individuals’ rights and freedoms, particularly for large-scale processing of special categories of information.

The formal order cites Article 35 among the provisions the regulator found violated.

A DPIA is not merely paperwork.

Done properly, it forces an organization to ask the questions that appear central to this case:

Is the data actually anonymous?

Can people be reidentified?

What information is necessary?

What risks arise from linking records longitudinally?

What happens if direct identifiers enter free-text fields?

How long will the information be retained?

Who receives it?

What security controls are necessary?

What legal basis supports the processing?

Those are precisely the questions that should be answered before a massive health-data repository is built.

The Garante Treated IQVIA as a Controller

Another important issue was role allocation.

The Garante concluded that IQVIA was acting as a data controller from the point at which information was collected from participating physicians.

That matters because companies involved in analytics sometimes characterize themselves as processors or technical intermediaries.

Controller status carries substantially more responsibility.

A controller determines the purposes and means of processing and therefore bears responsibility for questions such as:

lawfulness;

transparency;

data minimization;

retention;

privacy by design;

data subject rights;

security;

and DPIAs.

Companies cannot determine GDPR roles simply by writing “processor” into a contract.

Regulators look at what the parties actually do.

If an analytics company determines why a large dataset exists, how it is structured and how it is used for research, authorities may scrutinize whether that company is really functioning as a controller.

The Fine Was €7 Million, But the Corrective Order May Matter More

The financial penalty is substantial.

The Garante imposed a €7 million administrative fine after finding violations of GDPR Articles 5, 9, 13, 25, 28, 32 and 35.

But the corrective order may be more important operationally.

IQVIA was given 120 days to bring the processing into compliance if it intends to continue the activity.

The Garante said that, alternatively, anonymization would need to be performed independently by the physicians according to safeguards identified by the authority.

That can be far more disruptive than paying a fine.

A company may have to redesign the architecture underlying a major product or research database.

It may need to change:

data ingestion;

identifiers;

physician software;

free-text handling;

retention;

security;

contracts;

transparency;

and governance.

Privacy enforcement increasingly reaches the product itself.

Pharmaceutical Research Does Not Eliminate Privacy Obligations

There is an understandable temptation to view healthcare research as inherently beneficial.

Much of it is.

Large health datasets can help researchers understand:

drug utilization;

treatment outcomes;

adverse events;

disease prevalence;

healthcare costs;

and epidemiological trends.

But beneficial purpose does not erase privacy law.

The GDPR expressly accommodates scientific research, but it also requires safeguards.

The more sensitive and detailed the information, the more important those safeguards become.

Health analytics therefore has to solve two problems simultaneously:

preserve enough data utility to make the research valuable;

and reduce identifiability enough to protect patients.

That is difficult technical work.

Calling a dataset anonymous does not accomplish it.

This Is Increasingly Relevant to AI

The IQVIA decision also has implications well beyond conventional health analytics.

AI companies routinely want large datasets.

Healthcare organizations increasingly want to train models using clinical information.

The temptation will be to “deidentify” records and treat the resulting dataset as unrestricted.

That can be dangerous.

Suppose an AI training corpus contains:

age;

ZIP code;

diagnoses;

medications;

medical procedures;

rare diseases;

dates of treatment;

and longitudinal patient histories.

Removing the patient’s name may not be sufficient.

Rare combinations of characteristics can identify people.

Location makes that easier.

Longitudinal information makes it easier still.

Free-text clinical notes present an even greater problem because they frequently contain names, relatives, doctors, workplaces, addresses or unusual events.

The same problem that affected IQVIA can therefore appear inside AI training datasets.

Anonymization Needs to Be Tested, Not Assumed

Organizations handling sensitive information should treat anonymization as an engineering process that needs validation.

A serious anonymization review should consider:

singling out;

linkability;

inference;

dataset uniqueness;

rare attributes;

geographic precision;

temporal precision;

external datasets;

longitudinal identifiers;

free-text fields;

and the resources reasonably available to an attacker or recipient.

Depending on the use case, organizations may need techniques such as:

aggregation;

generalization;

suppression;

tokenization;

pseudonymization;

differential privacy;

removal of rare attributes;

geographic coarsening;

or synthetic data.

The correct approach depends on what the organization needs to accomplish.

There is no universal “anonymize” button.

Free Text Deserves Special Attention

One of the most practical lessons from the IQVIA case concerns unstructured information.

Privacy teams spend enormous amounts of time reviewing database fields.

Name.

Email.

Phone.

Date of birth.

IP address.

Customer ID.

Those structured fields are easy to map.

Free text is harder.

A notes field can contain every one of them.

This problem exists in:

medical records;

CRM notes;

customer support tickets;

call transcripts;

chat logs;

AI prompts;

email archives;

legal matter descriptions;

and HR systems.

If those fields flow into analytics systems, data warehouses or AI models, the organization may unintentionally move directly identifiable information into an environment that was designed under the assumption that it contained only deidentified data.

Organizations should therefore explicitly test free-text fields when conducting data mapping and privacy assessments.

Data Retention Cannot Be “Forever Because It Might Be Useful”

The presence of records dating back to 2001 is another warning.

Analytics teams naturally value historical data.

More years create better trend analysis.

More records improve statistical power.

Old information might eventually become useful for research that has not yet been conceived.

Those are business arguments.

They are not automatically privacy justifications.

The GDPR requires organizations to connect retention to purpose.

That means asking:

Why do we need 25 years?

Would 10 years accomplish the same research objective?

Can older records be aggregated?

Can identifiers be removed after a period?

Can detailed records be converted into statistical outputs?

Can the retention schedule differ according to the type of information?

A privacy program should be able to explain why information is still there.

“Storage is cheap” is not a retention policy.

The Bigger Lesson: Data Classification Drives Compliance

The IQVIA case can be reduced to one deceptively simple question:

What kind of data was this?

IQVIA reportedly treated it as anonymous.

The regulator treated it as identifiable health information.

That classification difference determined whether major parts of the GDPR applied.

This is why data classification is not merely an administrative exercise.

If a company incorrectly categorizes personal information as anonymous, it can miss:

legal-basis requirements;

special-category restrictions;

privacy notices;

DSAR obligations;

DPIAs;

security requirements;

retention limits;

processor/controller obligations;

and breach-response requirements.

One classification error can cascade through the entire privacy program.

What Companies Should Take From the IQVIA Fine

Organizations performing analytics, research or AI training on supposedly anonymized datasets should review the assumption rather than simply repeating it.

Ask whether individuals can be singled out.

Ask whether identifiers persist across time.

Ask whether location information narrows identity.

Ask whether rare medical conditions create unique profiles.

Ask whether free-text fields contain names or other identifiers.

Ask whether external datasets could be combined with the information.

Ask who determines the purposes of the processing.

Ask what lawful basis would apply if regulators concluded the information remained personal data.

And document the answers.

The Garante’s €7 million IQVIA decision is ultimately not a warning against health research.

It is a warning against treating anonymization as a label.

The distinction between anonymous information and pseudonymous personal data can determine whether an entire regulatory framework applies.

When a database contains the health histories of 1 million people, getting that distinction wrong can become very expensive.

Online Privacy Compliance Made Easy

Captain Compliance makes it easy to develop, oversee, and expand your privacy program. Book a demo or start a trial now.