OpenAI has an army of contractors who read real ChatGPT users’ prompts and other data to help improve the chatbot’s responses. The idea is that the contractors provide an, obviously, human touch to OpenAI’s models. Well, not all of the contractors are doing that. 404 Media has found multiple contractors hired to improve OpenAI’s models have been fired for using AI to train the AI. That’s not great for the models themselves, but there is also obviously a great irony in AI training companies working for OpenAI firing people for using AI when OpenAI’s whole thing is to make people use AI at work.
Some AI models already exhibit signs of “model collapse,” which is where AI models further trained on AI-generated text can become worse and worse. In this case, some of the people hired to partially stop that happening are themselves using AI-generated responses to train OpenAI’s models.
One contractor said they see people using AI “all the time and people are let go for it all the time, it’s pretty much the one thing that will get you kicked off ASAP.” The person said, “in a group of thousands there are tons that have been caught.”
Do you work as a prompt reviewer for OpenAI, Anthropic, or another AI company? I would love to hear from you. Using a non-work device, you can message me securely on Signal at joseph.404 or send me an email at joseph@404media.co.
Last week, 404 Media revealed Project Lily, in which OpenAI has hundreds of contractors reading real ChatGPT users’ prompts and conversations which can include personal information. Those contractors then rate and critique the responses ChatGPT generated, including making sure that the responses are not too sycophantic or anthropomorphize ChatGPT. That reporting was based on internal documents and conversations with a person who works on the prompts.
404 Media has now obtained other internal documents and spoken to three contractors doing work for OpenAI across various projects. Those projects can include more than ten thousand contractors, according to one of the internal documents.
One of those documents says contractors must not use AI themselves for their work.
“Do not use AI detection tools, or AI yourself,” one document describing the work of contractors who are hired to review the work of other contractors, including catching them for using AI, says. “Do not use GPTZero or any other AI detection tool. They are not reliable. Reviewers may not use AI either, including Grammarly and AI translation, to review, write feedback, or write comments.”
The document continues, “Do not tell evaluators why you suspect AI. It is easier for them to hide if they know what you look for. Judge the overall pattern, not one clue.”
All three of the contractors said reviewers are told not to use AI in their work. Two of the sources said people have been fired or offboarded for using AI. 404 Media granted the contractors anonymity as they weren’t permitted to speak to the press.
The contractors who review other contractors’ work are told to be on the look out for tell-tale signs of AI use. That can include repetitive words, AI-style punctuation — which might include over zealous use of the em dash — and contractors finishing their work very quickly.
In related Slack channels where people ask each other for advice, a lot of people will post an example with the question, ‘Is this AI?,’ one contractor said.
“Usually the answer is yes,” the person said.
One contractor said they used AI while helping to train OpenAI’s models and shared what they presented as their termination letter. It said their employer had identified issues with the “authenticity” of their work.
“I’m not a bad person or worker. I just needed a little boost and turned to AI to help me which eventually led to my downfall,” the contractor told 404 Media. “I felt no joy in the work or that I was contributing to society in any way.”
Two of the contractors 404 Media spoke to worked for Mercor, an AI-training company that hires the contractors who in turn review ChatGPT-related material. A Mercor spokesperson told 404 Media in a statement: “Our experts are hired for their expertise and judgement, which is essential to the ongoing advancement of AI. Our contracts strictly prohibit the use of LLMs to complete projects and we enforce that. We invest heavily in our tools and systems to detect misuse and ensure our experts comply with project rules and contract terms. When we confirm an expert has used AI to complete a task, we immediately remove them from the project.”
404 Media spoke to a fourth contractor who has worked on training models for various AI companies. They said they sometimes purposefully chose the worst responses because they wanted to actively sabotage the models’ training.
“I did feel guilty about doing this kind of work at the start,” they said. “I either pay zero attention to the results and choose randomly or purposely choose the [worst] output. I’m not sure how much of a difference it actually makes since there are hundreds of other people also rating prompt results, but it does feel like I’m getting paid to make AI worse.”
OpenAI declined to comment on its contractors being fired for using AI.
About the author
Joseph is an award-winning investigative journalist focused on generating impact. His work has triggered hundreds of millions of dollars worth of fines, shut down tech companies, and much more.