Est.

How Google Gemini Uses Conversations for Model Training

Google stores your Gemini chats for training even after you delete them.

Editor at Large · · 14 min read
Cover illustration for “How Google Gemini Uses Conversations for Model Training”
Mainstream AI Data · September 18, 2026 · 14 min read · 3,142 words

Google Gemini uses consumer conversations to train its models by default. That default runs through a specific toggle, a human review pipeline, and a set of retention windows that most users never open the settings menu to see, let alone read.

Gemini does not exist in isolation. It sits inside an ecosystem that already holds a user's email, calendar, location history, and search activity. Gemini's defaults carry more weight than a standalone AI product's defaults would. A chatbot with no memory of anything else about a person is one kind of privacy question. A chatbot embedded in the same account as Gmail and Google Photos is a different one entirely, because the data it can eventually touch is not limited to what gets typed into the chat box.

Most people open Gemini and start typing. They do not check what happens to those words afterward, and there's little reason they would, given how the interface is designed to feel like a conversation rather than a data submission. This piece maps that pipeline: it explains what the "Keep Activity" toggle actually does, how human review works and why it can outlast deletion, lays out the retention timelines that differ by product and by tier, and describes the geographic split that gives some users protections others do not get by default. It also covers the one lawsuit that tested these practices in US federal court, and what that case's outcome says about the limits of legal recourse for now.

Three contexts get covered here, and they should not be confused with one another: the consumer-facing Gemini app, the Gemini API (both free and paid), and Gemini inside Google Workspace or Google Cloud. The policies differ sharply across all three. Conflating them, treating a free API key like an enterprise contract, is probably the single most common source of user confusion on this topic.

What "Gemini Apps Activity" controls and what it does not

One setting governs almost everything discussed in this piece: Gemini Apps Activity, sometimes labeled "Keep Activity" in the account interface. It is on by default for consumer accounts. When it's on, Google's Gemini Apps Privacy Notice states that conversations get used to "provide, develop, and improve its services (including training generative AI models)." That is not a paraphrase; it's the language Google itself uses.

What does "improve its services" mean in practice? The stated uses include routing requests to whichever model is best suited to answer them, training new models on the patterns in those requests, and generating usage metrics to tailor the product experience. So a person's phrasing, question style, and follow-up patterns become inputs into how the system routes and responds to everyone else's future questions too.

Turning Keep Activity off changes some of that, but not all of it. Future conversations stop being saved to the account, stop going to human reviewers, and stop feeding the training pipeline. Personalization tied to conversation history disappears for new sessions. Gemini does not stop functioning, though: the product does not degrade when the toggle is off, it just stops learning from that particular user going forward.

What the toggle does not touch is operational data. Google still collects operational and account-level data regardless of the Keep Activity setting. And other Google products run on their own separate activity controls; this toggle is specific to Gemini and does not reach into those other systems. Even with the toggle off, there's a brief operational retention window, up to 72 hours, that Google keeps to actually deliver the service. That window is invisible in the user's activity panel, but it exists.

Temporary Chat mode is the one surface explicitly carved out of the training pipeline, designed for users who want conversations that are not saved to their account history. For anyone who wants to test this directly, the path is myaccount.google.com, then Data and Privacy, then Gemini Apps Activity, then Keep Activity, then turning it off.

Why conversations can outlive deletion in the human review process

When Keep Activity is on, a subset of conversations gets pulled for review by trained human reviewers, including contracted service providers Google works with. Their job is to judge whether a given response was low-quality, inaccurate, or harmful, and to suggest a better one. Those suggestions feed back into the model. This is the human feedback loop that informs model improvement, and it requires actual humans reading actual conversations somewhere in the process.

Before a reviewer sees anything, Google disconnects the conversation from the user's account and says it tries to anonymize the content. That's the theory. But what if a user typed their own name, their employer, or a medical detail directly into the chat? Anonymizing the account handle does nothing to redact the words themselves. The structural disconnection removes the metadata trail back to the account; it does not touch the content of the message. If someone pastes a paragraph containing a home address or a patient's diagnosis, that text is in the transcript a reviewer reads, regardless of what account it came from.

Reviewed conversations get retained for up to three years, on a track that runs separately from the user's own account and its settings. That separation is the crux of why deleted conversations can persist in review. Deleting a conversation from Gemini Apps Activity removes it from what the user can see. It does not pull it back out of the review pipeline if that conversation has already been selected for review. And once a piece of text has been folded into an actual training run, there is no retroactive extraction, the influence is baked into model weights at that point, not sitting in a file that can be deleted on request.

So the timeline can look something like this: a conversation happens today, gets pulled for review next week, and is in Google's review-retention system until roughly 2029, even if the user deletes every trace of it from their own account tomorrow. The deletion button clears the user's view. It does not necessarily clear Google's.

Gems, Google's term for custom Gemini configurations built with specific instructions and uploaded reference files, carry this same three-year review exposure. If Keep Activity is on, a custom Gemini configuration built around a company's internal style guide or a personal project's research notes is subject to the identical pipeline as an ordinary chat message.

Retention timelines across every scenario, mapped

The numbers here vary enough by tier that a single blanket statement about "Gemini's retention policy" would be misleading. It helps to walk through each scenario on its own terms.

For consumer accounts with Keep Activity on, the default retention window is 18 months, stored inside the user's Google Account. Users can adjust that down to 3 months (the shortest non-zero option available) or up to 36 months. Separately, reviewed conversations sit on that up-to-three-year track, disconnected from the account and untouched by whatever retention window the user picked in their settings.

For consumer accounts with Keep Activity off, conversations are not saved to the account and are not used for training or review. The 72-hour operational window for service delivery still applies. No configuration on the consumer product, including the toggle turned fully off, achieves a true zero-retention state.

The Gemini API tells a different story depending on billing status. On the free tier, meaning Google AI Studio or an unbilled API key, inputs and outputs get used to improve Google's models, the same as the consumer default. The line that actually matters here is not which interface someone is using, it's whether a billing account is attached. Once an API key is linked to paid billing, prompts and responses stop being used for training. On that paid tier, inputs and outputs are retained for up to 55 days, and only for abuse monitoring, not model training.

Google Cloud and Vertex AI carry their own separate windows depending on the feature. Grounding with Google Search retains relevant data for 3 days for debugging. Vertex AI grounding with Google Search retains relevant data for 30 days for debugging purposes. The Gemini Live API's session resume feature, when a user enables it, caches text, video, and audio prompt data along with model output for up to 24 hours.

Inside the Gemini Enterprise Agent Platform, the numbers get more granular still. CodeMender sessions, including any source code snippets involved, are retained for up to 7 days, though source code itself gets deleted within seconds of a session ending, with the remainder auto-deleted at the 7-day mark. Grounding with Google Search inside that same platform logs data for up to 3 days for debugging. Separately, data caching for the Gemini Live API session resume feature stores prompt data and model output for up to 24 hours, a setup that falls under Google's enterprise data handling commitments.

The tier structure that separates consumer, API, and enterprise users

Three tiers, three different rulebooks. The consumer tier covers the free version and paid consumer subscriptions, all running on gemini.google.com. Training opt-in is the default here, the 18-month retention window applies by default, and human reviewers can access sampled conversations for up to three years. Here's the detail that surprises people: Google AI Pro, the paid subscription, treats user data identically to the free tier. Paying for the product does not upgrade its privacy protections. There is also no Data Processing Addendum on this tier. It is not a suitable option for anyone processing EU personal data on behalf of an employer.

The paid Gemini API sits in a genuinely different regime. Prompts and responses are not used for model training. Processing happens under the Google Cloud Data Processing Addendum, and inputs and outputs are retained for up to 55 days for abuse monitoring, not model training.

Gemini for Google Workspace, covering enterprise, business, and education accounts, is different again. Google states it does not use Workspace data to train or improve generative AI models outside of Workspace without explicit permission. Interactions stay inside the organization's own environment, enterprise-grade security controls apply, and there's no human review by default. Workspace admins get to set conversation history rules directly: auto-delete windows of 3 months, 18 months, or 3 years, or indefinite retention. That default warrants attention from any admin who has not actively configured a deletion window.

Google Cloud and Vertex AI extend similar protections at the infrastructure level. Prompts and responses are not used to train models, and eligible enterprise customers can negotiate contractual terms equivalent to zero data retention through amendments to the Data Processing Addendum, though that requires direct engagement with a Google Cloud account team rather than a settings toggle. Data residency, for organizations that need it, gets handled through region pinning at project creation.

The practical trap here is straightforward to describe and easy to fall into anyway: a developer testing prompts on a free Google AI Studio key is operating under the same privacy regime as a casual consumer user, not the enterprise one, even if that developer works for a company with a Workspace enterprise agreement elsewhere. The billing relationship determines the policy, not the technical surface or the company's other contracts with Google. Anyone evaluating Gemini for sensitive or regulated work needs to answer one question before anything else: which tier is this actually running on? Assuming enterprise protections apply to a consumer session or a free API key is the single costliest assumption available here.

Personal Intelligence and the Gmail integration: what the ecosystem access means

Google announced a feature called "Personal Intelligence" on January 14, 2026, letting Gemini draw on a user's data across Gmail, Search, Photos, and YouTube to shape more personalized responses. The rollout moved in stages: limited beta for Google One AI Premium subscribers starting on that January date, then extended to all US users on a standard Google account by March 2026. Some of the more advanced features remain locked behind the Google One AI Premium subscription, available at a monthly subscription price,.

The training boundary here matters and deserves a careful read. Google states that Gemini does not train directly on a user's Gmail inbox or Google Photos library. What it does train on is the specific prompts typed into Gemini and the model's responses to them, used to improve the capability itself. So the underlying email or photo gets referenced to generate an answer, but that underlying content is not itself folded into a training run. That's a real distinction. But it's also a distinction that rests entirely on Google's own characterization of its internal pipeline, since there's no external way for a user to verify where exactly the line falls inside the system.

The rollout also split by geography. In the US, Gmail access under Personal Intelligence was structured so that users could control the connection through their account settings, though the rollout details varied and the practical starting position for many users depended on what underlying toggles were already active. In Europe, the same feature required a genuine, explicit opt-in before Gemini could touch Gmail at all, a direct consequence of GDPR's stricter defaults around data sharing.

Control over Personal Intelligence connections is accessible through the Gemini interface settings as of early 2026. Gmail access specifically can also be switched off through Gmail's own settings menu (See all settings, then Smart features) without touching the rest of the account.

Why does any of this matter beyond the mechanics? Because the risk isn't any single data source on its own, it's the combination. AI conversations start merging with search history, email patterns, location data, and YouTube viewing habits inside one profile. Any one of those streams alone tells a limited story. Together, they build something far more detailed than the sum of the parts, and that compounding effect is the real stake in Personal Intelligence's design.

The geographic exception that gives EEA, UK, and Swiss users stronger default protections

On the free Google AI Studio tier and unpaid Gemini API quota, the general rule is that submitted content improves Google's products and human reviewers may see it. There's an exception, and it's a geographic one. Users located in the European Economic Area, Switzerland, or the United Kingdom get the paid-services data-use terms applied to all Gemini services, including the free ones. That means prompts and outputs from those users are not used to improve Google's products, full stop, under the terms as they stood on June 29, 2026.

GDPR gives EU consumers meaningful legal rights, including access and deletion requests. But those rights run into the same wall described earlier in this piece: they do not compel Google to strip a reviewed conversation out of its systems during that three-year retention window. The right exists on paper. The practical outcome, once a conversation has entered the review pipeline, is a gap between what the law grants and what actually happens to the data.

Separately, the EU's AI Act brought its own transparency requirements into force on August 2, 2026, under Article 50: users must be told they're interacting with an AI system, and AI-generated content has to carry machine-readable markings. Systems already on the market by that date get until December 2, 2026 to meet the machine-readable marking piece specifically.

The US regulatory picture looks thinner by comparison, though it is not empty. California's AB 2013, the Generative AI Training Data Transparency Act, took effect January 1, 2026, and requires developers releasing generative AI models in California to publish a high-level summary of their training datasets, including whether copyrighted material or personal data was part of the mix. Federal privacy legislation covering AI training data has been proposed in various forms but has not been enacted. State-level laws like California's CCPA offer some avenue for recourse, though enforcement and available remedies vary. Both the California Privacy Protection Agency and the European Data Protection Board have acknowledged that deletion rights, under CCPA and under GDPR's Article 17 respectively, may extend to AI models that contain personal data. Neither body, though, has laid out a formal, mandated technical process for actually unlearning that data from a trained model.

Geography changes what protection a user gets from the identical product, without that user doing anything differently. Someone in one city and someone in another city can type the same sentence into the same version of Gemini and land under two different sets of rules.

Thele v. Google LLC and the Limits of Legal Recourse in the US

Thomas Thele, along with co-plaintiff Melo Porter, filed suit against Google in the Northern District of California on November 11, 2025. The claim: Google had "secretly turned on Gemini for all its users' Gmail, Chat, and Meet accounts" around October 10, 2025, giving the model access to scan private communications without proper consent, in alleged violation of Section 632 of the California Invasion of Privacy Act.

The plaintiffs' theory turned on a specific mechanism. They argued Google used an already-existing toggle, the Smart Features setting, as a kind of back door, flipping on Gemini's deeper access to user data without presenting a fresh, distinct consent prompt for that new capability. There's a real distinction buried in that argument: switching on a feature nobody explicitly agreed to is not the same as a user actively re-consenting to a new use of their data, even if the toggle involved happens to be one that technically already existed in some form. Google's position, for its part, was that these features are long-standing, optional, and sit fully under user control.

On July 7, 2026, Judge Noël Wise of the Northern District of California granted Google's motion to dismiss. The ruling turned on standing, not on the merits of what Google actually did. The plaintiffs, in the court's view, had not alleged a concrete enough harm to establish Article III standing in federal court. They were granted leave to file an amended complaint, so the case was not closed for good, but the dismissal as it stands says something specific.

What does it say? That describing what an AI tool is capable of doing is not enough to get a case past the courthouse door. A plaintiff has to plausibly allege what actually happened to them, specifically, not what the system theoretically could have done with their data. That threshold, Article III standing, functions right now as the practical ceiling on consumer AI privacy litigation in US federal court. The ruling is not a finding that Google's practices were lawful or acceptable on the merits. It is a finding that federal court, under the current reading of standing doctrine, is not yet the venue where those merits get tested. Whether that changes on appeal, or in a differently framed amended complaint, remains open. For now, it marks the edge of what US litigation alone can currently do to answer the questions this piece has been mapping from the start.

Sources

  1. Gemini Enterprise Agent Platform 및 제로 데이터 보관 | Google Cloud Documentation
  2. Google Gemini Data Retention Policy 2026 - Does Gemini Train on Your Data? | Meetily
  3. Gemini Apps Privacy Hub - Gemini Apps Help
  4. Gemini adds Temporary Chats and new personalization features
  5. knowledge.workspace.google.com
  6. blog.google
  7. techcrunch.com

More in Mainstream AI Data