Mathematicians Call for Transparency Regarding OpenAI's Data Usage
Mathematician Andreas Thom stated that OpenAI is not transparent about whether it uses private interactions and unpublished research to train its artificial intelligence models.
Mathematician Andreas Thom has raised fresh concerns about artificial intelligence giant OpenAI, stating that the company is not transparent or honest about the origin of its training data.
Data Debate in the Mathematical World
Questions are mounting regarding the data behind OpenAI's recent successes in the field of mathematics. A new addition has been made to the debates over whether the company's AI models benefit from unpublished work.
Following New York University professor Tristan Buckmaster, mathematician Andreas Thom also drew attention to the lack of transparency in the company's data policies.
Mastodon Post and Area of Expertise
Making posts on Mastodon, Andreas Thom voiced concerns that the interactions they had with ChatGPT may have contributed to the company's success.
The process was triggered when one of the results announced by OpenAI involved non-sofic groups, Thom's area of expertise, and the company acknowledged that this work built upon the previous research of Thom and Gábor Kun.
Lack of Attribution and Revised Texts
Following OpenAI's announcement of the result regarding non-sofic groups, criticisms were directed from mathematical circles on the grounds that Thom and Kun's contributions were omitted.
In the wake of the incoming criticisms, the company was forced to quietly change its online post.
Email Correspondence and Unsatisfactory Responses
Thom sent emails to OpenAI researchers Sébastien Bubeck and Mark Sellke, asking whether their chats were part of the training data.
However, he stated that the responses from the company only addressed whether the chats were directly accessed, and failed to clarify whether they encompassed broader training pools.
Burden of Proof and Expectation of Disclosure
Thom emphasized that researchers cannot reverse-engineer OpenAI's training pipeline and that the company holds the sole authority in this regard.
He argued that if the company denies these allegations, it must provide proof by disclosing the necessary datasets.
Millennium Prize and Similar Defenses
OpenAI's avoidance of drawing a clear boundary regarding the use of user data has manifested similarly in its Millennium Prize achievements.
While the company defended that private user data was not accessed in the Navier-Stokes solution, it stated that it cannot entirely rule out indirect effects.
Ethical Concerns and Scientific Secrecy
Thom stated that using unpublished research provided by users to improve models without consent, proper disclosure, or attribution is ethically indefensible.
It was stated that these events have caused concern among mathematicians and could lead researchers to act more covertly in the future.