Data licensing for US companies with 20+ employeeshello@getdataoffer.com

Legal & privacy

How to de-identify business data before licensing it

Practical techniques for de-identifying Slack, email, documents and CRM data while preserving the workflows AI labs value, plus common pitfalls.

A hand blacking out lines on a printed page with a marker

De-identification is what makes it possible to license operational data without exposing the people in it. Done well, it removes who did something while preserving what was done and why, which is exactly what AI labs want. Done poorly, it either leaks identities or strips so much context that the data loses its value.

This guide covers the main techniques, how they apply to different systems, and how to verify the result. It's general information, not legal advice. Regulated data (health, financial, children's data and so on) can have specific legal standards for de-identification, so involve counsel where those apply.

De-identification vs. anonymization vs. pseudonymization

These terms are often used loosely, so it helps to be precise:

  • Pseudonymization replaces identifiers with consistent placeholders ("Person 14", "Customer 22"). The same person gets the same placeholder throughout, so relationships and threads stay intact.
  • De-identification is the broader process of removing or transforming information so individuals can't reasonably be identified. Pseudonymization is usually one part of it.
  • Anonymization typically implies that re-identification isn't reasonably possible even with additional information. It's the strongest standard and the hardest to guarantee for rich text.

For operational data, the practical goal is thorough de-identification with consistent pseudonyms, combined with contractual prohibitions on re-identification.

What counts as an identifier

Direct identifiers are obvious: names, email addresses, phone numbers, street addresses, government ID numbers, account and card numbers, and usernames.

Indirect (quasi-) identifiers are subtler. They're details that, combined, could single someone out: a unique job title ("our only Spanish-speaking paralegal"), specific dates tied to personal events, small-office locations, or rare medical or personal details mentioned in passing.

Business identifiers matter too: your customers' names, vendor names, deal values tied to named accounts, and internal project code names you'd rather not disclose.

Core techniques

Consistent pseudonyms

Map every person, customer and vendor to a stable placeholder. Keep the mapping table inside your company (or destroy it after processing) and never deliver it.

Pattern-based detection

Regular expressions and validators catch structured identifiers reliably: emails, phone numbers, URLs, card and account numbers, IP addresses, postal codes.

Named-entity detection

Machine-learning entity recognition finds names, organizations and locations in free text. It's essential for chat and email but imperfect, so combine it with dictionaries built from your own directory (employee lists, CRM account names).

Generalization

Replace precise values with ranges or categories where precision isn't needed. Exact birth dates become age bands, precise addresses become regions, and exact deal values become brackets.

Suppression

Remove content entirely when it can't be safely transformed: certain attachments, images of documents, signatures, or entire channels and folders.

Date shifting

For some datasets, shifting all dates by a consistent random offset preserves intervals (how long things took) while obscuring when events happened.

System-by-system considerations

Slack and chat. Replace user IDs and display names, scan message text, handle @mentions, and watch for pasted emails, screenshots and file names. See selling Slack data.

Email. Headers, signatures, quoted reply chains and disclaimers all carry identifiers. Thread reconstruction must survive pseudonymization so conversations still read in order.

Documents and Drive. File names, author metadata, comments and revision history can all leak identity. Scanned images and PDFs may need OCR before scanning, or exclusion.

CRM and ERP. Structured fields make redaction easier: you can drop or hash contact fields while keeping stages, timestamps, amounts (possibly bucketed) and activity types. Free-text notes still need scanning. See selling CRM and ERP data.

Code repositories. Scrub secrets (API keys, credentials, tokens), author emails in commit history, and customer data in test fixtures.

Preserving value while removing identity

The trick is to remove identity without destroying structure. A few principles help:

  • Be consistent. "Person 14" must always be the same person, or threads become meaningless.
  • Keep roles where safe. "Person 14 (Account Manager)" is often fine and adds a lot of value, unless the role is unique enough to identify someone.
  • Preserve timing. Relative time and sequence matter more than absolute dates.
  • Don't over-delete. Replacing a name is better than deleting the sentence it's in.

Verifying the result

No automated pipeline is perfect, so verification is part of the process:

  1. Automated re-scan of the output to confirm detectors find nothing left behind.
  2. Manual review of random samples by someone who knows the business and can spot indirect identifiers.
  3. Targeted searches for known sensitive terms: executive names, top customers, project code names.
  4. Documented methodology so you (and the buyer) know exactly what was done.
  5. Your sign-off on a de-identified sample before anything is delivered.

Contractual safeguards

Technical de-identification should be backed by contract terms: prohibition on re-identification attempts, restrictions on combining the data with other sources to identify individuals, security requirements, and deletion of raw copies. Our licensing agreement checklist covers these.

Who does the work?

You can run de-identification internally, use specialist tooling, or have a partner's engineers do it under your supervision. DataOffer's engineers can handle export, cleanup and de-identification, or support your team in running it. Either way, you review and approve the approach and a sample before any data is shared, and nothing is delivered until you approve the buyer, price and terms.

Ready to see what your data is worth?

Share rough estimates (systems, approximate volume, years of history, headcount) and we'll come back with competing offers from AI labs. No upfront cost, no commitment, and nothing is shared until you approve.

This guide is general information, not legal, tax or financial advice. Figures and ranges are illustrative; talk to qualified advisors about your situation.