# When an AI trainer uses AI, the missing product is human judgment

> Reported offboardings over prohibited AI use expose a larger business problem: companies must distinguish AI-assisted production from the independent human judgment they intend to buy.

By BIG CHANGE Editorial

Published: 2026-09-22T17:43:18.226Z
Updated: 2026-09-22T23:32:55.869Z
Canonical: https://bigchange.ai/blog/ai-trainers-human-judgment-independent-evaluation

![Three inspection chambers show a machine lens, a human profile alone, and a machine lens beside a human profile.](https://bigchange.ai/api/media/file/ai-trainers-human-judgment-hero-v1.png)
AI-generated conceptual illustration by BIG CHANGE. AI-assisted production, independent human judgment and disclosed assisted review can all be useful, but they provide different kinds of evidence. This conceptual apparatus shows those distinct roles; it is not a product interface, detector or claim that synthetic feedback is always harmful.

Companies hire people to compare model answers, explain what is wrong and provide examples of better work. In that setting, a polished response is only part of the order. The buyer may be paying for a judgment that did not come from the model being tested.

[404 Media reports](https://www.404media.co/people-training-openais-ai-fired-for-using-ai-to-train-the-ai) that three contractors working on OpenAI-related projects described rules against using AI, and two said contractors had been offboarded for doing so. Mercor, which supplied two of the people interviewed, told the publication that its contracts prohibit using large language models to complete project work and that confirmed violations lead to removal. OpenAI declined to comment. We have not seen the private project documents or independently verified individual cases.

The useful question reaches beyond one contractor platform: when a business asks for an independent human check, how should it separate that work from tasks where AI assistance is welcome?

## The big change

- **What changed:** The reported contractor dispute exposes a conflict over what AI training work is meant to supply: completed answers or an independent human assessment.
- **Why it matters:** Mercor's public pilot makes professional judgment the measurement. Substituting model-generated rankings changes what that evaluation measures, even if the submission is well written.
- **What to watch:** The consequential choice for evaluation buyers is which stages require an unaided judgment. That choice belongs in the task instructions, where workers can see it before accepting the work.

## Independence can be the thing being purchased

An AI assistant is useful when the task is to produce a first draft quickly. It becomes a problem when the task is to supply a reference answer or an independent comparison and that distinction is concealed.

Mercor's public listing for a [business-domain model evaluation pilot](https://work.mercor.com/jobs/list_AAABoBbYguDNWJDFwNRJ6Z0n/business-domain-expert-ai-model-evaluation-pilot-admin-marketing-hr-accounting) makes the requirement unusually clear. Experts solve a professional task, rank five model attempts, score them and write evidence-based rationales. The listing says no rubric or golden answer is provided because professional judgment is the measurement. It prohibits AI assistants during solving, ranking, scoring and rationale writing.

These terms describe one pilot. Its comparison depends on a professional's own assessment; delegating that assessment to a model changes the experiment.

## Synthetic feedback is not automatically bad data

It is tempting to connect any AI-generated training response to “model collapse.” That would outrun the evidence here. [Research published in Nature](https://www.nature.com/articles/s41586-024-07566-y) examined recursive training regimes in which generated data replace data drawn from an original distribution. In those experiments, successive models lost information about the underlying distribution. That setup is not a study of the reported contractor work, and the paper cannot tell us whether any particular rating affected a deployed OpenAI model.

Another [model-collapse study](https://arxiv.org/abs/2404.01413) shows why the distinction matters. Its authors found collapse when each generation's synthetic data replaced the original real data, but avoided it in their experiments when real data remained and successive synthetic data accumulated alongside it. The result does not prove that every mixture is safe. It does show that “AI-generated” is not a sufficient diagnosis by itself; provenance, selection and the training process matter.

## Make the permitted workflow explicit

A brief should name permitted tools, identify stages requiring an unaided judgment and state how workers should disclose assistance. Mercor's pilot specifies independence across solving and evaluation. Buyers commissioning assisted work should describe the professional judgment they still expect the worker to supply.

Review needs the same care. The 404 Media report says one internal document warned reviewers against using AI detection tools and described them as unreliable. That is consistent with a basic management principle: punctuation, speed or polished phrasing is not proof that someone used a model. A reviewer can examine whether the rationale cites the source material, ask the worker to explain a decision, compare repeated submissions and investigate tool records that the project lawfully collects. No single stylistic clue should decide a worker's case.

## Workers need a contestable rule, not a guessing game

Independent contractors can lose access to a project quickly. A credible system should therefore make the rule visible before paid work starts, distinguish a suspected violation from a confirmed one, and provide a channel for a worker to explain relevant evidence. This is a recommendation for fair process, not a claim about rights under any contract or law.

Workers, meanwhile, should treat an unaided judgment requirement as part of the product specification. Someone who believes an approved tool is necessary can ask before using it, or decline the task. Concealed assistance does not become acceptable because the model writes well, just as a prohibited outside collaborator would not become acceptable because the answer was correct.

## Specify the source of judgment

The public Mercor pilot makes the product specification explicit: an unaided professional assessment. Buyers commissioning assisted work need an equally clear description of which tools are permitted and what judgment the worker must supply.

## Sources

- [404 Media: People Training OpenAI's AI Fired for Using AI to Train the AI](https://www.404media.co/people-training-openais-ai-fired-for-using-ai-to-train-the-ai) — Original reporting based on contractor interviews and documents describes rules against AI use and reported offboarding. We have not seen the private documents or independently verified individual cases; OpenAI declined comment.
- [Mercor: Business Domain Expert — AI Model Evaluation Pilot](https://work.mercor.com/jobs/list_AAABoBbYguDNWJDFwNRJ6Z0n/business-domain-expert-ai-model-evaluation-pilot-admin-marketing-hr-accounting) — The public listing states that professional judgment is the measurement and prohibits AI assistants during solving, ranking, scoring and rationale writing. It establishes this pilot's stated design, not every Mercor project's rules or enforcement.
- [Shumailov et al.: AI models collapse when trained on recursively generated data](https://www.nature.com/articles/s41586-024-07566-y) — The paper studies recursive training settings where generated data replace data from an original distribution. It does not evaluate the reported contractor cases or establish harm to a deployed OpenAI model.
- [Gerstgrasser et al.: Is Model Collapse Inevitable?](https://arxiv.org/abs/2404.01413) — Experiments distinguish replacing original real data from accumulating synthetic data alongside it. Results show that synthetic data's effects depend on the data process; they do not certify every real-and-synthetic mixture.
