The question every operator should be able to answer

If someone asked you today where the data from your AI tools actually lives, who can see it, and what happens to it after the response comes back, could you answer with any confidence? Most operators I talk to can't. They know the vendor name and the monthly bill. Everything else is fog.

That fog is a problem. Your team is pasting customer emails into ChatGPT, feeding invoices to a document extractor, and letting a voice agent handle inbound calls. Each of those actions moves data somewhere. If you don't know where, you can't answer basic questions from a customer, an auditor, or your own attorney.

This is a plain-terms guide to figuring it out. No compliance theater, no jargon parade. Just the questions to ask, where the answers live, and what to do with what you find.

The five hops your data actually takes

When you type a prompt or upload a file to an AI tool, the data typically moves through a handful of stages. Understanding these stages is the whole game.

  • The client: your browser, desktop app, or the automation platform calling the API. Data leaves your device here.
  • The vendor's application layer: the middleware that logs your request, applies rate limits, and routes it. This is often where prompt logs and analytics get stored.
  • The model host: the actual inference infrastructure that runs the model. This might be the same company (OpenAI runs its own models) or a different one (many tools call Anthropic, OpenAI, or an open-source model on AWS Bedrock or Azure).
  • Training pipelines: a separate path where your data may or may not be routed to improve future models. This is the one people worry about most, and it's usually the easiest to turn off if you ask.
  • Subprocessors: hosting, monitoring, storage, analytics, support tooling. Every vendor uses a handful, and they should be listed publicly.

Your job isn't to memorize infrastructure diagrams. It's to know which hops happen for each tool you use, and to have written answers for each one.

Where the real answers live

Marketing pages will tell you a tool is "enterprise-grade" and "secure by design." Ignore that. The real answers live in four specific documents, and every serious vendor publishes them.

1. The Data Processing Addendum (DPA)

This is the contract that says what the vendor can and can't do with your data. Look for: whether they act as a processor or controller, whether they train on your data by default, retention periods, and what happens if you terminate. If a vendor won't sign a DPA, that's your answer about how seriously they take this.

2. The subprocessor list

Usually a public page titled "Subprocessors" or "Trust Center." It lists every third party that touches your data. You'll typically see AWS or Google Cloud for hosting, Datadog or similar for monitoring, and one or two model providers if the vendor doesn't run its own. If the list is missing or vague, ask.

3. The security overview or trust page

This is where SOC 2 Type II reports, HIPAA readiness, and encryption practices are documented. "Encryption in transit and at rest" is table stakes now. What matters more is who holds the keys, how long logs are kept, and whether you can request deletion.

4. The API and product documentation

This is where you find the actual default behavior. For example, OpenAI's API doesn't train on your data by default, but the ChatGPT consumer product used to. Anthropic's API has similar defaults. Google's Gemini has different rules for the free tier versus paid Workspace. Read the docs for the specific product tier you're paying for, not the company's general privacy page.

The checklist I actually use

Before we deploy an AI tool for a client, or before I let one into my own stack, I run through this list. It takes about thirty minutes per vendor and saves you from the "wait, it does what?" conversation later.

  • Where are the servers geographically? (US, EU, and multi-region matter for different reasons.)
  • Is my data used to train models by default? How do I turn that off, and is the switch at the account level or per-request?
  • How long are prompts and outputs retained? Can I set that to zero?
  • Who at the vendor can see my data, under what circumstances (support tickets, abuse review, legal requests)?
  • What model actually runs my request, and is it the vendor's or a third party's?
  • Do they sign a Business Associate Agreement if I'm in healthcare? Do they sign a DPA if I have EU customers?
  • What's the breach notification timeline in the contract?
  • Can I export and delete everything on demand?

You don't need to be a lawyer to ask these. You need to be the person who wrote them down and got answers in writing.

The special case for healthcare front offices

If you run a medical or dental practice, this gets sharper. Any AI tool that touches patient information (call transcripts, scheduling notes, insurance details, intake forms) is handling PHI. That means you need a signed BAA before you send a single piece of data in.

Plenty of general-purpose AI tools will not sign a BAA. That doesn't make them bad tools. It means you can't use them for patient-facing operations. The consumer version of ChatGPT, for instance, is not appropriate for pasting patient messages into. The API with the right configuration and a BAA is a different conversation. Ask your vendor directly, get it in writing, and confirm the specifics with your own compliance counsel. The rules aren't optional and the answers aren't always intuitive.

What to do with what you find

Once you have this information, write it down somewhere your team can find it. A simple spreadsheet with one row per tool works: vendor, purpose, data types it touches, training opt-out status, DPA signed, BAA signed if needed, and the person on your team who owns it. This is the whole of "AI governance" for most small and mid-sized businesses. It doesn't need to be more complicated than that.

The mistake I see most often is treating this as a one-time exercise. Vendors change their terms. New subprocessors get added. A product tier gets renamed and the defaults shift. Put a recurring calendar reminder to review the list every quarter. Ten minutes per tool is enough.

If you'd rather have someone build automations that come with this documentation already sorted, so your team can move fast without inheriting a governance mess, talk to our team at Qintara Corp. We handle the boring parts on purpose.

Frequently Asked Questions

Does using an AI tool mean my data is used to train future models?

Not automatically. Most business-tier API products from major vendors default to no training on your inputs. Consumer products are often the opposite. The answer depends on the specific product tier and your settings, so check the docs for what you're actually paying for and confirm in writing.

What's the difference between a DPA and a BAA?

A DPA (Data Processing Addendum) is a general contract about how a vendor handles your data, and it's what you need for GDPR and most commercial arrangements. A BAA (Business Associate Agreement) is a specific US healthcare contract required by HIPAA when a vendor handles protected health information. You may need both, depending on your business.

Is it safe for my team to paste customer data into ChatGPT?

It depends on which ChatGPT and which customer data. The free consumer product has different data handling than ChatGPT Team, Enterprise, or the API. If you haven't decided which version your team should use and written down what's allowed, assume the answer is no and fix the policy before the pasting continues.

How do I know if a vendor is actually secure or just claims to be?

Ask for their SOC 2 Type II report under NDA. A real one is dozens of pages and describes actual controls and any exceptions found. If they can't produce one, or if the "security" page is all marketing copy with no documents behind it, that tells you what you need to know.

Do I need a dedicated compliance person for this?

For most operators, no. You need one person who owns the spreadsheet, reads the DPAs, and updates the list quarterly. For regulated industries or larger teams, you'll want counsel or a fractional compliance advisor to review the harder decisions. Start with the list, then decide what you're missing.