Help Center

What Is Personally Identifiable Information (PII)?

Human Written & Fact Checked

Cite this Webpage

Copy

Hadley McIntosh. “What Is Personally Identifiable Information (PII)? (Updated August).” Swif, August 6, 2026, www.swif.ai/learn/data-trust/personally-identifiable-information Accessed 20 August 2026.

Personally identifiable information (PII) is information that can distinguish or trace a specific person, either by itself or when combined with other linked or linkable information, and therefore requires handling decisions based on its context, sensitivity, use and applicable obligations.

PII can appear in structured records, documents, messages, images, logs and application data. It may be stored in a cloud service, copied to a laptop, viewed on a phone or transferred through an employee workflow. Protecting it requires knowing both what the data contains and how people and systems use it.

Names and government identifiers are familiar examples, but identifiability is not limited to obvious fields. A combination of location, job title, device identifier and event time may point to one person even when no name appears. The NIST glossary therefore includes information that distinguishes or traces an identity alone or when combined with linked or linkable information.

PII is not one universal legal category with an identical definition everywhere. Laws and regulatory frameworks may instead use terms such as personal information, personal data or protected health information, each with its own scope. Data loss prevention provides useful prerequisite context for how organizations can govern sensitive data movement without assuming that one control settles every privacy obligation.

Why personally identifiable information matters

PII connects data to a person who may experience consequences if the information is exposed, altered, misused or retained without a valid purpose. The possible effects vary with the data and situation. They can include account takeover, identity fraud, unwanted profiling, discrimination, physical-safety concerns, embarrassment or loss of control over private information.

The same field can create different risk in different contexts. A work email address printed on a public conference page is not equivalent to that address attached to a private performance record. A precise location may be routine for a delivery in progress but highly sensitive when it reveals a medical visit or a protected person's home.

Organizations also need to know which rules apply to a particular data set, business activity and location. The label PII can help teams find and govern identifying information, but the label itself does not determine a legal result. Privacy and legal reviewers must map the actual information and processing activity to applicable requirements.

How information becomes identifiable

Information becomes identifiable when it can reasonably single out, trace or be linked to a person in the relevant environment. That assessment considers the data itself, other available data, who can access both and what linking methods are practical.

Three patterns account for most PII decisions:

  • Direct identification. A name, portrait, government identifier or account record points directly to a person.
  • Indirect identification. An attribute such as age, location or occupation narrows a group and becomes identifying when combined with other attributes.
  • Persistent linkage. A cookie, device ID, employee number or pseudonymous account consistently connects actions or records to the same person and may be linkable to an identity elsewhere.

Identifiability can change. A dataset that appears anonymous inside one isolated system may become identifiable after it is joined with a directory, customer database or public record. Conversely, aggregation, deletion of unnecessary fields and controlled de-identification can reduce the likelihood that records point back to individuals.

Removing names alone is not a reliable test. The remaining values may still be unique, and free-text fields can contain names, contact details or personal circumstances that a field-based scanner misses.

Common types of PII

PII is easier to recognize when teams group information by how it identifies a person rather than rely on a fixed list. The examples below are illustrative; whether a value is PII depends on the data set and governing definition.

Information groupExamplesWhy context matters
IdentityFull name, portrait, signature, employee number, government identifierSome values identify directly; common names may need another attribute
Contact and locationHome address, phone number, email address, precise locationA public business contact and a private home location carry different risks
Financial and accountBank details, payment card data, customer account numberA number may become identifying through an account system or related record
Digital and deviceIP address, cookie ID, mobile advertising ID, device identifierPersistence, ownership and access to linking records affect identifiability
Biometric and biologicalFace template, fingerprint, voiceprint, genetic informationThe method, purpose and governing rule determine whether and how it is protected
Personal attributesDate of birth, nationality, family status, education, employment historyAttributes can identify through combination even when none is unique alone
Activity and behaviorBrowsing events, purchases, travel, access logs, communicationsSequences and timestamps can distinguish a person or reveal sensitive conduct
Health and benefitsDiagnosis, treatment, disability, insurance or benefits informationSector-specific rules may apply only to certain entities, records and activities

The table should not become a universal compliance checklist. For example, a device serial number can identify a company asset without identifying a person. It becomes PII when an assignment record or another reliable source links that device to an employee.

PII, sensitive data and sensitive personal information

PII, sensitive data and sensitive personal information overlap, but they are not interchangeable.

PII focuses on whether information identifies or can be linked to a person. It can include relatively ordinary information, such as a work email address, as well as information whose misuse could cause substantial harm.

Sensitive data is a broader operational category. It can include personal information, credentials, API keys, trade secrets, unreleased financial results and proprietary source code. Some sensitive data does not concern a person at all.

Sensitive personal information usually describes a higher-risk subset of personal information under a particular law, policy or classification scheme. The included categories vary. California, for example, defines personal information broadly and separately enumerates sensitive personal information in California law. That state-specific definition should not be treated as a global list or automatically applied outside its statutory scope.

TermMain questionTypical scope
PIICan the information identify or be linked to a person?Direct and indirect identifiers in context
Sensitive dataWould unauthorized access, use or change create meaningful harm?Personal and nonpersonal organizational data
Sensitive personal informationDoes a governing rule or policy place this personal information in a higher-risk category?A defined subset that varies by jurisdiction or framework

An organization can use an internal classification model that maps these overlapping terms to handling rules. It should still retain the original legal or contractual labels needed for each processing activity.

How organizations govern PII

PII governance is a lifecycle process. It connects a justified business purpose to collection, use, access, storage, sharing, retention and disposal.

  1. Define the purpose and scope. Record why the information is needed, whose information is involved and which systems, vendors and teams process it.
  2. Discover and classify the data. Identify structured fields and unstructured content, then label the data according to identifiability, sensitivity and applicable rules.
  3. Reduce unnecessary collection. Limit fields, precision, copies and retention to what the approved purpose requires.
  4. Assign access. Give users, applications and service providers only the access needed for their roles and tasks, with an exception path for unusual cases.
  5. Protect storage and movement. Apply appropriate encryption, file permissions, sharing restrictions and data loss prevention policies without treating any one control as sufficient.
  6. Record important actions. Maintain proportionate evidence of access, changes, exports, policy decisions and approved exceptions.
  7. Review and dispose. Reassess changed purposes, systems and recipients, then securely delete or de-identify information that no longer needs to remain identifiable.

De-identification requires more than deleting a name column. Techniques can include aggregation, suppression, generalization and carefully controlled tokenization, but the acceptable method depends on the intended use and applicable rule. In the health context, for example, HHS guidance describes two specific HIPAA methods and notes that even properly de-identified health data can retain some identification risk. Those methods apply within HIPAA's scope rather than establishing a universal standard for every dataset.

How PII appears on managed endpoints

Employee endpoints often handle PII temporarily or indirectly. A laptop may synchronize customer records, cache email attachments, download a payroll report, store browser data or write identifying values to logs. A phone may display contact details and notifications even when the system of record remains in the cloud.

Endpoint controls can reduce some exposure when an organization knows which devices are managed and which configurations apply. Relevant measures may include device encryption, screen-lock policy, managed application settings, access removal during offboarding, software updates and selective removal of organizational data where the platform and ownership model support it.

These controls do not determine whether information is legally PII, discover every copy or replace privacy governance. They also cannot govern data on an unmanaged personal device merely because the person has a company account. Organizations managing PII stored or handled on enrolled company endpoints can consider unified endpoint management as one endpoint-management layer. Separate data classification, identity, application, DLP, retention and legal processes remain necessary.

A managed-laptop PII example

Northstar Transit, a fictional company, gives a benefits analyst a managed laptop. The analyst is asked to send an eligibility file to the company's approved benefits provider. The file contains employee numbers, work locations, dates of birth and benefit selections but no employee names.

The starting assumption is that removing names made the file anonymous. During classification, the privacy team finds that the employee number links directly to the human-resources directory and that the combination of work location and date of birth can distinguish several employees. The file remains identifiable and includes personal information that the company treats as sensitive.

Policy permits the approved provider to receive only the fields required for enrollment through a controlled transfer channel. The analyst's managed laptop is encrypted and current, but an attempt to copy the file to a personal synchronization folder is blocked by a data-handling rule. The employee removes an unnecessary location field, transfers the approved version and records the provider and retention schedule.

The result does not depend on declaring every birth date sensitive in every context. Northstar evaluates the complete dataset, linking environment, requested action, recipient and purpose, then applies a proportionate decision.

Benefits of identifying and governing PII

Accurate PII identification improves decisions across privacy, security and operations.

  • Proportionate safeguards. Teams can focus stronger controls on data and uses with greater identification or harm potential.
  • Reduced data sprawl. Purpose and retention reviews remove unnecessary fields, downloads and duplicate records.
  • Clearer access decisions. Owners can connect identities, roles, devices and applications to approved uses.
  • More reliable incident scoping. Responders can determine which people, records, systems and recipients may be affected without assuming every event has the same impact.
  • Better vendor oversight. Contracts and technical reviews can describe the actual data, purpose, transfers, return and deletion expectations.
  • Consistent employee guidance. Classification labels translate abstract privacy rules into practical handling expectations.

These benefits depend on maintaining the inventory and rules. A label applied once can become inaccurate when a dataset gains new fields, moves to another system or is combined with a new source.

PII risks and limitations

PII programs fail when they reduce a contextual decision to keyword matching or a static list.

  • Indirect identifiers are easy to miss. Combinations of ordinary attributes can single out a person.
  • Unstructured content resists simple detection. Documents, images, chat messages and free text may contain personal details without predictable field names.
  • Data joins change risk. A pseudonymous value can become identifying when another party has the linking table or related data.
  • False positives disrupt work. A scanner may mistake test values, public contact data or unrelated number patterns for high-risk PII.
  • False negatives create blind spots. Custom identifiers and contextual clues may pass undetected.
  • Endpoint visibility is incomplete. Cloud-to-cloud transfers, unmanaged devices and personal accounts may sit outside endpoint controls.
  • Encryption has a boundary. It can protect stored or transmitted data but does not prevent an authorized application or user from misusing readable information.
  • Definitions vary. A corporate PII label does not by itself prove that a specific privacy, breach-notification or sector rule applies.

Teams should therefore document uncertainty. When a classification is ambiguous, the record should identify the responsible reviewer, the assumed context and the event that will trigger reassessment.

PII and related concepts

Adjacent data concepts answer different questions and should remain separate.

ConceptPrimary questionBoundary from PII
Personal data or personal informationDoes a particular law or framework cover information relating to a person?Uses the controlling definition and scope, which may differ from an internal PII label
Protected health informationIs identifiable health information held or transmitted in a relationship covered by HIPAA?It is a sector-specific US legal category, not a synonym for all health-related PII
Data classificationWhich label and handling rule should apply to this information?Converts PII and other data characteristics into operational categories
Data loss preventionShould this data action be allowed, warned, blocked or recorded?Enforces selected movement or use rules but does not establish legal status
Access controlWhich principal may perform which action on a resource?Restricts use after identity and policy decisions
Data exfiltrationHas data been transferred from an environment without authorization?Describes an event or outcome, not a data category
Anonymization or de-identificationHas identification risk been reduced enough for a defined purpose and governing standard?Requires a method and context; removing direct identifiers may be insufficient

Data classification turns findings about PII and sensitivity into operational labels. Data exfiltration explains unauthorized movement, while file access control governs which identities can act on files containing PII.

The practical objective is not to place every piece of information into the broadest possible category. It is to understand who can be identified, what harm or obligation may arise, which processing is justified and which safeguards remain effective throughout the data lifecycle.