Data licensing for US companies with 20+ employeeshello@getdataoffer.com

Getting started

How to sell company data to AI companies

A practical, step-by-step guide to selling or licensing your company's operational data to AI labs, from what to inventory to how offers and approvals work.

Two colleagues seen from behind working on laptops at a conference table overlooking a small-town main street

AI labs have largely exhausted the easy sources of training data. What they want now is harder to find: realistic examples of how businesses actually operate. That means the conversations, documents, tickets and records that show work moving from request to decision to outcome. Many established companies already have years of this data sitting in Slack, Google Drive, email and their CRM or ERP.

This guide walks through how selling (or, more precisely, licensing) that data works in practice, what you need to prepare, and how to protect your company along the way.

Why AI labs want company data now

Public web text teaches a model how people write. It doesn't teach the model how a purchasing manager weighs three vendor quotes, how a support team escalates a tricky ticket, or how an operations lead turns a messy request into a finished project. Labs building AI systems that do real work need examples of real work.

Operational data from functioning businesses fills that gap. A few years of internal discussion, documentation and system records shows patterns that are hard to fake: handoffs between teams, approvals, revisions, exceptions and the reasoning behind decisions. For more detail, see what AI labs look for in enterprise data.

Step 1: Take a rough inventory

You don't need exact numbers to start. A useful first inventory answers five questions:

  1. Which systems do you use? Slack or Teams, Google Drive or OneDrive, Notion or Confluence, email, a CRM, an ERP or accounting system, ticketing tools, code repositories and any SOP or wiki tools.
  2. About how much is in each? Rough orders of magnitude are fine: thousands vs. millions of messages, gigabytes vs. terabytes of files.
  3. How many years of history? Longer histories are generally more valuable because they show how work evolves.
  4. What's your headcount? More people usually means more varied workflows and handoffs.
  5. How long have you been in business? Mature companies tend to have more established, documented processes.

Most admin consoles show storage and record counts, so whoever manages your IT can often answer these in an afternoon.

Step 2: Decide what's in and what's out

Before anyone talks price, decide what you're comfortable including. Common exclusions include:

  • HR and payroll channels or folders
  • Legal matters and privileged communications
  • Board materials and M&A discussions
  • Specific customer accounts covered by strict confidentiality terms
  • Anything subject to regulations you can't satisfy (for example, protected health information without the right agreements in place)

You can exclude at the system level (no email at all) or within a system (everything in Slack except three channels). Good brokers and buyers expect this. Exclusions reduce volume, but a clean, well-scoped dataset is usually worth more per unit than a messy one.

Step 3: Package and de-identify

Raw exports aren't what labs buy. Data needs to be exported, cleaned, structured consistently and de-identified: names, email addresses, phone numbers, account numbers and other identifiers are removed or replaced with consistent placeholders so the workflow stays readable but the people don't.

Done well, a reader can still follow "Person A asked Person B to approve a revised quote for Customer 17," without learning who any of them are. Our guide to de-identifying business data covers the techniques in depth.

You can do this work internally, or have a partner's engineers do it under your supervision. Either way, you should see and approve the de-identification approach before anything leaves your control.

Step 4: Get competing offers

There isn't yet a public price list for enterprise data, which is exactly why running a competitive process matters. Different labs value different things. One may prioritize long Slack histories, while another wants structured CRM records or a specific industry. Taking a packaged description of your dataset to several buyers tends to surface those differences.

At this stage, buyers typically see a description, a schema and possibly a small de-identified sample. They don't see the full dataset. Offers then reflect volume, history, uniqueness, quality and the license terms requested.

Step 5: Review terms and approve

The agreement matters as much as the price. Key questions include:

  • Is it a license (you keep your data) or a transfer?
  • Is it exclusive or non-exclusive?
  • What uses are permitted, such as training, evaluation, or both?
  • What are the security, retention and deletion obligations?
  • Who is liable if something goes wrong?

Use our AI training data licensing agreement checklist and have counsel review it. Nothing should be delivered until you've approved the buyer, the price and the terms.

Step 6: Deliver securely and get paid

Delivery usually happens through a secure transfer to the buyer's environment once the agreement is signed. Payment terms vary: some deals pay on delivery, others in tranches tied to milestones. Make sure the agreement spells out exactly what "delivered" means.

Common mistakes to avoid

  • Selling to the first buyer. Without competition, you have no way of knowing whether an offer is fair.
  • Skipping the exclusions conversation. It's much easier to exclude sensitive areas up front than to claw data back later.
  • Underestimating de-identification. Free text hides identifiers in surprising places, like signatures, file names and pasted screenshots.
  • Ignoring existing contracts. Customer agreements, vendor terms and privacy notices may limit what you can share. See is it legal to sell company data?
  • Overcommitting your team. If export and cleanup would pull engineers off core work, have the partner's team do it.

How DataOffer fits in

DataOffer acts as a broker between your company and leading AI labs. We help you scope the dataset, package and de-identify it (or support your team in doing so), take it to multiple labs, and bring you the best offer. Offers for US companies with 20+ employees generally range from $50K to $1M+, depending on the data.

There's no upfront cost, no equity and no commitment to get an offer, and nothing is shared until you approve the buyer, price and terms. To get started, we only need rough estimates: which systems you use, roughly how much data is in each, years of history, headcount and years in business.

Ready to see what your data is worth?

Share rough estimates (systems, approximate volume, years of history, headcount) and we'll come back with competing offers from AI labs. No upfront cost, no commitment, and nothing is shared until you approve.

This guide is general information, not legal, tax or financial advice. Figures and ranges are illustrative; talk to qualified advisors about your situation.