Valni

Provider data use

Do AI companies read your messages?

When you send a request to a model, it reaches that model's provider, and each provider sets its own rules for what happens to your data: whether they train on it, how long they keep it, and which country's law governs it. Below is how the major providers compare, in plain language, taken from their own published policies and linked so you can check every claim.

As of July 2026. Providers change their terms; the linked source next to each answer is authoritative. This is a plain-language summary, not legal advice.

What Valni does on your behalf: wherever a provider offers a way to opt out of using your data to train or improve its models, Valni takes it by default. For the US providers and Moonshot's overseas tier that is already the default or available on request; where an opt-out has to be requested, Valni requests it. Zhipu's hosted API offers no opt-out, so routing to it means accepting its terms.

Provider Trains on your API data? Retention Zero-retention option Where it's processed
AnthropicClaude No, opt-in only 30 days by default Yes, on request US (California); Ireland for EEA/UK
OpenAIGPT No, opt-in only Up to 30 days (abuse monitoring) Yes, by approval US (California); Ireland for EEA; 11 regions
GoogleGemini (paid API / Vertex AI) No on the paid tier; the free tier does Paid logs up to 55 days; Vertex ~24h cache Yes, paid tier (per project) Customer-selected region; Google Cloud DPA
xAIGrok No, opt-in only 30 days (abuse auditing) Yes, self-serve or via sales US (Tennessee / Texas)
DeepSeekDeepSeek Yes, with an opt-out No fixed period No China; PRC law
MoonshotKimi No on the overseas Business tier; opt-out on the API Not stated; file deletion available By enterprise arrangement Not stated (Singapore entity; may process in China)
ZhipuGLM Broad "optimization" rights, no opt-out No fixed period No China (mainland); PRC law

In their own words

The exact clauses behind each row, with links to the source document and its effective date.

Anthropic / Claude

"Anthropic may not train models on Customer Content from Services." Retained data "is never used for model training without your express permission."

Content flagged by trust-and-safety systems can be retained longer (up to 2 years).

OpenAI / GPT

"As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)."

Content flagged by classifiers may be retained and human-reviewed; Modified Abuse Monitoring and a BAA reduce this.

Google / Gemini (paid API / Vertex AI)

"When you use Paid Services... Google doesn't use your prompts or responses to improve our products." The free tier is the opposite: content is used "to provide, improve, and develop Google products... and machine learning technologies."

The free, unpaid Gemini API tier is used for model improvement and can be human-reviewed. Grounding with Google Search forces 30-day retention.

xAI / Grok

"xAI never trains on your API inputs or outputs without your explicit permission." The Enterprise terms add: "xAI shall not use any User Content to train any foundation models..."

xAI reserves rights over de-identified and aggregated data derived from your use.

DeepSeek / DeepSeek

Personal data is used "to train and improve our technology, such as our machine learning models and algorithms," with a stated "right to opt-out of using your Personal Data for training our models."

Data is collected, processed, and stored in the People’s Republic of China. The policy permits sharing with law enforcement and public authorities.

Moonshot / Kimi

Overseas Business Service Agreement: "We will not use Customer Content submitted to, generated by, or stored through the Business Services to train, optimize, or improve our artificial intelligence models, unless Customer provides express authorization or such use is required by applicable law." The API help page adds that data submitted through the API "is not used to train or improve Kimi's models." The separate Chinese-language consumer terms grant broad training rights with no opt-out, so which applies depends on the tier you use.

The referenced DPA is not published and the processing location is not stated; both require engaging Moonshot sales to confirm.

Zhipu / GLM

Translated from the Chinese: Zhipu "may utilize data generated during your use of the platform... to locate, maintain, and optimize our products and services," and stores data "within mainland China."

Obligated to cooperate when state authorities lawfully review user-uploaded data. No stated no-training or zero-retention option.

How to decide which models to support

The right models to support depend on the data you handle, not just on price and quality. A useful way to read the table above:

Handling regulated or sensitive data

Require no training on your data by default, a zero-retention option, and a jurisdiction your compliance team accepts. In the table, the US-governed providers and Google's paid tier meet that bar; the China-based providers do not offer zero-retention and process under PRC law.

Internal or low-sensitivity work

The data-handling bar is lower, so you can weigh cost and capability more freely and support a wider set of models.

Building a product on top

You inherit your provider's terms and pass them to your own customers, so the choice is a compliance decision, not only a quality one. If you resell inference, your customers' data lives under the terms of whichever model served their request.

Once you have decided, Valni lets you enforce it. A model policy restricts which models your agents can call, and the usage ledger shows what each one costs, so the decision is both made and kept. Decide the set you trust, then hold your agents to it.

This summarizes each provider's published policy as of July 2026, with links so you can verify. For how Valni itself handles your data, see Your data, our Privacy Policy, and Terms.