OpenAI announced last week that one of its models solved the Navier-Stokes existence and smoothness problem, a 90-year-old mathematical challenge and one of the seven Millennium Prize Problems. The company said it took 88 hours using 10,000 AI agents, but it has already acknowledged applying a method developed by two Spanish scientists and has not clarified whether it also used work from a researcher at its main competitor, Anthropic.
"We did not use their prompts or their proofs to guide our models or agents," an OpenAI spokesperson said, adding that "although unlikely, we cannot rule out that data derived from the use of our products helped improve our models." That admission cuts to a question researchers and writers are asking with growing urgency: can chatbots extract your ideas and hand them to someone else?
How training data flows into models
Before an AI model can be used, it goes through a training phase. An algorithm processes a massive database so the system can learn patterns through deep learning and neural networks. The process is autonomous, though results can later be refined manually through supervised learning. Three ingredients make this work: enormous computing power, well-formulated algorithms, and databases large enough for the learning to produce meaningful results.
Data quality matters directly. Better data produces better models. And data is starting to run short. Some estimates suggest GPT-4 had already consumed all internet content plus many additional archives. Companies are searching for new sources. "It's not entirely clear what OpenAI or Anthropic do for their models to learn. We only have intuitions and what people who have worked there say. It is fairly clear that they train on people's conversations with them, but in theory not on all of them," said Álvaro Barbero, director of the AI Lab at fraud prevention company Lynx Tech and professor at Afi Global Education.
Julio Gonzalo, professor of Computer Languages and Systems at UNED, put it bluntly: "User interactions are a very valuable source of information for training models." Carlos Gómez Rodríguez, professor of Computing and Artificial Intelligence at the University of La Coruña, explained that this information gets integrated with the immense volume of data models already hold, mostly from the internet, along with material that company employees may generate internally.
Can a prompt end up in another user's results?
Could a researcher write prompts containing valuable ideas, only to see those ideas surface months later when a new model version trained on those requests becomes available? "On paper, it could be that if the prompts used to train the model contain useful information for solving a specific problem, that information gets applied in later requests. It's not something that can be demonstrated 100%, but it's logical to think so," Gómez said.
There is no certainty that every prompt gets used. "Today, there is no way to know if any particular prompt ended up represented in the model. You cannot follow that trace, because the training process modifies the model's weights, which are like an enormous matrix of numbers that we cannot really even interpret," Gómez added.
Gonzalo described the learning process as "a kind of lossy compression." The rarer or more infrequent the input, the less likely the model memorizes it well. But the inverse also applies: the more specific and rare a new user's conversation, the more likely the model recalls conversations on that same rare topic. "Overall, it is possible, though not inevitable, that the original ideas you share with a model could end up being used by others," Gonzalo concluded.
What companies offer - and what they don't
The dividing line is payment. Free-tier users get no guarantee their data won't be used for training. "If you use the free version, they give you no guarantee they won't use your data to train, meaning they probably are doing it," Barbero said. Paid enterprise agreements with OpenAI or Anthropic theoretically include clauses preventing the use of conversations to improve models. Barbero added a caveat: "I say in theory because, for example, Anthropic had to pay compensation to book authors when it was discovered they had downloaded many from pirate sites."
Gómez noted that AI had solved other mathematical problems before Navier-Stokes, and experts on that particular problem say the model also made its own contributions rather than simply following third-party research lines. He added it would not be surprising if companies were paying special attention to prompts from highly capable users. "OpenAI is offering prominent scientists, mathematicians, and engineers free access to their frontier models," Gonzalo said. "Thinking they do this altruistically would be very naive, given everything we are seeing."
Why this matters for researchers and writers
If you share original research directions, unpublished findings, or proprietary writing with a free AI tool, you have no practical way to know whether that material later surfaces in someone else's output. The model's internal weights are not auditable, and companies remain opaque about exactly which conversations feed into training. For paid enterprise accounts, contractual protections exist on paper, but the Anthropic book-piracy case shows that enforcement can lag behind discovery. The safest assumption for anyone working with sensitive intellectual property is that free-tier prompts are training data, not private exchanges.
Your membership also unlocks: