Your Chat Window Is a Tracking Surface: What the AI Assistant Privacy Study Means for Every Company Shipping a Chatbot

Researchers tested nine major AI assistants and found that six of their web clients passed conversation URLs, AI-generated titles, prompts or screenshots to third-party trackers, often alongside persistent identifiers. Nobody designed those leaks. They came from ordinary analytics tags meeting a new kind of page. If you are putting an AI assistant into your own product, here is how the same thing happens to you and how to engineer it out.

Security & AIYour Chat Window Is a Tracking Surface: What the AI Assistant Privacy Study Means for Every Company Shipping a Chatbot

Nine Assistants, 44 Third Parties, and a Lot of Leaked Titles

Last week a paper titled Prompt like a Butterfly, Sting like a Tracker reached the front page of Hacker News and collected more than 400 points and 140 comments. The researchers, based at IMDEA Networks in Madrid, ran a systematic privacy analysis of nine consumer AI assistants: ChatGPT, Claude, Grok, DeepSeek, Perplexity, Gemini, Microsoft Copilot, Mistral's Le Chat and Meta AI. They covered the web client of all nine and the Android app of the eight that offer one, and tested each under different cookie choices and subscription tiers.

The headline numbers are blunt. Across the services, the team identified 44 third-party organisations, and every assistant integrated at least one third-party advertising or tracking service. Six of the nine web clients and three of the eight Android clients disclosed conversation URLs, titles, prompts or screenshots to third parties, frequently alongside identifiers that let the recipient link the activity to a person. Three web clients sent AI-generated conversation titles to nine third parties, including Meta, TikTok and DoubleClick. The authors went through a responsible disclosure process with the providers and with European data protection authorities.

Some of the specific findings read like a checklist of everything a privacy engineer worries about. On Grok's shared conversation pages, the Meta Pixel PageView event carried the conversation title and the user's latest prompt, and TikTok received a screenshot of the most recent part of the conversation. Perplexity transmitted users' hashed email addresses to the marketing analytics company Singular. The web clients of ChatGPT and Claude sent the globally unique conversation ID to Datadog as a separate parameter.

It would be easy to read this as a story about a handful of AI labs and move on. That would miss the point. The timing matters: OpenAI began an advertising pilot for ChatGPT's Free and Go tiers in the US, and Criteo announced on 2 March 2026 that it was the first ad-tech partner in that pilot. Advertising infrastructure and conversational interfaces are converging. And the same tags the researchers found in these assistants are already running on the websites and apps of most companies now adding an AI assistant to their own product.

Nobody Decided to Send the Prompt

The most useful thing in the paper for engineering teams is not the list of offenders. It is the mechanism. In almost every case the leak did not come from a line of code that said 'send the prompt to Meta'. It came from conversational products putting user content into places that tracking tags have read automatically for fifteen years.

Look at what a chat interface does that a normal web page does not. It generates a short title summarising the conversation and puts it in the browser tab, which means it lands in document.title. It creates a permanent URL for every conversation. When a user shares a chat, it builds Open Graph and meta description tags so the link previews nicely, and those tags contain a summary of the conversation or the prompt itself. The researchers' appendix shows exactly this on Grok: the shared page's meta description held the user prompt, and the Meta Pixel PageView event picked it up as a standard page metadata field.

Every one of those fields is something a generic analytics or advertising tag collects by default. A PageView event records the page URL and title. Many tags also read meta descriptions and Open Graph data. Session replay tools record what is on the screen. Error monitoring tools attach the current URL, breadcrumbs of recent user actions and sometimes request bodies to every exception. None of that was a problem when the page title was 'Pricing' or 'Order history'. It becomes a problem when the page title is an AI-written summary like 'Managing debt after a divorce' or 'Symptoms of early pregnancy at 41'.

The paper makes the same point about titles: they are AI-generated summaries that compress the purpose of a conversation into a few words, which is precisely what makes them revealing. A product team adding an assistant to a banking app, a health portal, an HR platform or a legal tool will produce the same artifacts. If the marketing team's tag manager container is loaded on that route, the same leak follows, whether or not anyone intended it.

Key Takeaways

  • Chat products put sensitive content into page titles, URLs, meta tags and share previews
  • Standard PageView, session replay and error monitoring tools collect those fields by default
  • The leak is an architecture side effect, so code review of the chat feature alone will not catch it
  • Any company embedding an assistant in an existing site or app inherits the risk from its existing tags

The obvious objection is that users agreed to this, or could have refused. The study tested that directly. Rejecting non-essential cookies did reduce what third parties received: in Claude's case, rejection prevented the Meta Pixel, Datadog telemetry and server-side forwarding to eleven advertising platforms. But across the free tiers, third-party trackers still collected data in four of the nine services after users rejected non-essential cookies. Perplexity, DeepSeek, Gemini, Copilot, ChatGPT and Claude all still connected to Google Ads under the reject-all setting.

Paying did not help either. The researchers observed no clear difference between free and premium tiers in third-party data collection. That is worth sitting with if you sell a B2B product and assume your paying customers are protected by default. The paper explicitly excluded enterprise and government tiers, so it says nothing about contractual enterprise offerings, but it does show that 'paid' and 'private' are different properties that have to be engineered separately.

The study also documents a shift that makes client-side blocking less effective. Claude routed Segment Analytics through a first-party domain and used a server-side configuration that forwarded events to eleven trackers, including Facebook, LinkedIn, TikTok, Reddit and Google. Grok routed events through a server-side Google Tag Manager container that sent the conversation URL and title server-to-server to the Meta Conversions API and the TikTok Events API. Both approaches are invisible to the browser and cannot be stopped by an ad blocker. In both cases the researchers observed this forwarding only after users accepted non-essential cookies, so consent was respected, but it means that once consent is given, the payload is whatever your server decides to forward.

Server-side tagging is now standard marketing advice, usually sold as better data quality and better privacy. Both can be true. But it moves the decision about what leaves your infrastructure from the browser, where researchers and regulators can inspect it, to a configuration file that often belongs to the marketing team and is rarely reviewed by engineering. If that configuration forwards the page title by default, it forwards your users' conversations.

For European companies, the uncomfortable legal point is that you do not get to blame the tag vendor. In the Fashion ID judgment (C-40/17), the Court of Justice of the EU held that a website operator embedding a third-party plugin can be a joint controller for the collection and transmission of visitors' data to that third party. The IMDEA authors lean on the same case: enabling third-party access is itself a processing decision, whether or not the data is ever read.

The consent side is equally broad. The European Data Protection Board's Guidelines 2/2023 on the technical scope of Article 5(3) of the ePrivacy Directive make clear that the consent requirement is not limited to cookies and extends to techniques such as tracking pixels and URL-based tracking. Enforcement is not theoretical. In September 2025 the French regulator CNIL fined SHEIN 150 million euros for placing cookies without valid consent and, on the same day, Google 325 million euros.

In the US, the precedent comes from health care, which is the closest analogue to a chatbot people confide in. In 2022, The Markup's Pixel Hunt investigation found the Meta Pixel on 33 of Newsweek's top 100 US hospital websites, sending details such as appointment information to Meta. Advocate Aurora Health, which had pixels on its website, MyChart portal and app, later agreed to a $12.225 million class action settlement. The FTC ordered online counselling service BetterHelp to pay $7.8 million and banned it from sharing health data for advertising, after it disclosed health questionnaire information to platforms including Facebook and Snapchat.

Now put an AI assistant in front of those same users. A patient portal chatbot, a mortgage assistant or an employee benefits helper will receive exactly the information those cases were about, in free text, summarised into a title by your own model. The paper also notes that providers described what third parties received in vague terms such as 'user content' or 'service interaction info'. If your privacy notice says the chatbot's conversations are private and your tag manager says otherwise, the tag manager is what a regulator will read.

Key Takeaways

  • Under Fashion ID, the site that embeds a third-party tag can be a joint controller for what it sends
  • EDPB guidance extends ePrivacy consent rules beyond cookies to pixels and URL-based tracking
  • US pixel cases in health care already produced settlements and FTC orders
  • A chatbot concentrates the same sensitive data into free text, and your own model summarises it

The second class of risk in the study is not tracking at all. It is access control. All nine services exposed shared conversations to anyone holding the link, without authentication, and nine third parties were present on those sharing pages, which gave them full visibility into the shared content. Some providers made conversation permalinks publicly readable by default. The researchers planted canary URLs in prompts and uploaded files, and observed Grok accessing them repeatedly after the conversation had ended. Perplexity fetched canary URLs even when the prompt explicitly told it not to.

We have seen what happens when share links meet search engines. In August 2025, Forbes reported that xAI had made more than 370,000 Grok conversations searchable on Google, because the share feature produced URLs that search engines could index. Users clicking share believed they were sending a link to a colleague. In practice they were publishing a page.

If your product lets users share an assistant conversation, with a teammate, a support agent or a customer, treat the feature as a publishing system and design it like one. That means explicit audience selection, authentication by default, expiry, a noindex header, no marketing tags on the shared page, a preview card that contains no conversation content, and a revocation path the user can find. None of this is difficult. It simply does not happen unless someone writes it into the specification.

The Engineering Fixes: A Privacy-Clean Chat Surface

The good news is that this is a well-bounded engineering problem. You do not need a new privacy platform. You need a small set of architectural rules applied to every route and screen where conversation content exists, and a test that keeps them true.

First, isolate the chat surface. Serve the assistant from a route group, subdomain or app module where the marketing tag manager, advertising pixels and session replay are simply not loaded, rather than relying on each tag being configured correctly. Conversion tracking can still fire on the pages around the assistant, such as sign-up or checkout, without ever seeing a conversation.

Second, keep content out of the fields that tags read. Use a generic document title such as the product name instead of the AI-generated conversation title. Use opaque, unguessable conversation IDs and never put titles, topics or prompt fragments into URL paths or query strings. Generate Open Graph and meta description tags for shared pages from fixed text, not from conversation content.

Third, scrub your own telemetry. Error monitoring, logging and APM tools are third parties too. Configure them to drop request and response bodies on assistant endpoints, strip conversation IDs from URLs before events are sent, and disable breadcrumbs that capture input fields. The paper's finding that conversation IDs reached Datadog is a reminder that 'it is only our monitoring vendor' still counts as disclosure.

Fourth, if you use server-side tagging, make the forwarding configuration an engineering artifact. Put it in version control, give it an explicit allow-list of fields per destination, and require review from someone outside marketing before it changes. Finally, use a Content Security Policy on the chat surface whose connect-src and script-src list only your own origins and the model provider you actually call. The researchers noted that CSP headers reveal the full set of third parties a page is authorised to contact, including dormant ones, so a strict policy is both an enforcement mechanism and evidence for auditors.

Key Takeaways

  • Load no marketing, advertising or session replay tags on routes that render conversations
  • Use generic page titles, opaque IDs and fixed share-preview text so content never reaches tag-readable fields
  • Configure error monitoring and logging to drop bodies, breadcrumbs and conversation IDs on assistant endpoints
  • Version-control server-side tag forwarding with a field allow-list, and lock the chat surface down with CSP

Make It a Test, Not a Policy

A privacy rule that lives in a document will be broken by the next marketing campaign or SDK upgrade. The IMDEA team found these leaks with standard tooling: browser developer tools capturing network traffic, an instrumented Android runtime, and canary tokens. You can do the same thing to your own product, and you can do it on every build.

On the web, run an end-to-end test that opens the assistant, sends a prompt containing a unique marker string, and records every network request. The test fails if any request goes to a host not on the allow-list, or if the marker string, the conversation ID or the generated title appears in any outbound request outside your own API. Run it under both cookie states, accepted and rejected, because the study showed tracking surfaces differ between the two. Repeat it for the shared-conversation page.

On mobile, keep an inventory of every SDK in the build with the data each one can access, since third-party SDKs run inside your app's process and inherit its permissions. Review that inventory on every dependency update, and include a proxy-based traffic capture of the assistant flow in your release checklist. For the share feature, add a test that an unauthenticated request to a shared conversation returns the access-controlled response you expect and that the page carries a noindex header.

None of this is exotic, but it does require someone to own it. In most companies the assistant is built by a product team, the tags are owned by marketing, the monitoring is owned by platform engineering and the privacy notice is owned by legal. The leak lives in the gaps between them, which is why it needs an engineering owner and an automated check rather than another meeting.

Key Takeaways

  • Seed prompts with a unique marker string and fail the build if it appears in any third-party request
  • Test with cookies accepted and rejected, and test the shared-conversation page separately
  • Inventory mobile SDKs and capture assistant traffic through a proxy before every release

Where a Nearshore Engineering Team Fits

Most companies adding an assistant to their product are doing it under time pressure, often on top of a codebase and a tag setup that has accumulated for years. The privacy work described here is not large, but it touches front-end routing, mobile builds, observability configuration, CI pipelines and the share feature at the same time. That is exactly the kind of cross-cutting work that gets postponed when the in-house team is focused on shipping the feature itself.

At Stepto, we build AI features for European and US clients with this boundary designed in from the start. A dedicated development team from Serbia can own the chat surface end to end: isolating it from marketing tags, scrubbing telemetry, building access-controlled sharing and adding the network-level regression tests that keep it clean as the product evolves. Because our engineers work in European time zones with good overlap with US teams, they can sit in the same reviews as your marketing, platform and legal stakeholders, which is where these issues are actually resolved.

For teams that already have an assistant in production, we also run focused audits: capturing the real traffic of your web and mobile clients under different consent states, mapping every third party that receives conversation-derived data, and delivering the fixes rather than just a report. As a nearshore partner operating under GDPR ourselves, we treat data minimisation as an engineering requirement, the same way we approach developer access to production data. If you are planning an AI assistant and want the build and the privacy architecture handled together, our AI developers can start with a short scoping engagement.

Treat the Conversation as the Most Sensitive Page You Have

The assistant privacy study is not really about nine AI companies. It shows what happens when a new kind of interface, one that people talk to candidly and that summarises their words into titles and links, is dropped into a web stack built around tracking every page view. The leaks came from defaults, not decisions, and the same defaults are running in most products that are about to add a chatbot. The fix is straightforward engineering: a chat surface with no marketing tags, content-free titles and URLs, scrubbed telemetry, access-controlled sharing and an automated test that fails when any of that breaks. Build it in before launch and it costs a few sprints. Discover it afterwards, from a researcher, a regulator or a class action, and it costs far more. If you want an experienced team to build your AI assistant with that privacy boundary in place from day one, talk to Stepto about a dedicated development team.

Building a team in Eastern Europe?

StepTo helps European and US companies build senior-led nearshore engineering teams in Serbia. Let's talk about what your next engagement could look like.

Start a conversation
I

Written by

Igor Gazivoda

Founder & CEO · StepTo

Igor has 15+ years in software engineering and business development. He specializes in scaling engineering teams, nearshore strategy, and AI-driven product development. He holds a Master's in Computer Science from the University of Belgrade.

LinkedIn →
Performance-led engineering

Want senior engineers who move work forward, not just tickets?

Work with accountable, English-fluent professionals who communicate clearly, protect quality, and deliver with a steady operating rhythm. Cost efficiency matters, but performance is why clients stay with us.

Delivery signals · senior engineering team
Senior ownership
Lead-level
Delivery rhythm
Weekly
Timezone overlap
CET
1 teamaccountable for outcomes, communication, and execution