"Data work" is a vague phrase that hides something quite specific. This is an attempt to describe it without marketing language, for people considering whether it is something they could do.
The short version
Modern data systems do not learn from reality. They learn from a record of what people decided reality was. Someone had to decide: this image contains that object. This answer is better than that one. This record duplicates the previous one. This claim is not supported by the source it cites. This document says something different from its own summary.
Data work is the making of those decisions, at volume, consistently, according to a written specification — and the flagging of the cases the specification did not anticipate.
What the tasks look like
Labelling and annotation. Marking up text, images, audio or video against a guideline. The easy items take seconds. The interesting part is the ten per cent that the guideline does not cleanly cover, where the useful contribution is noticing that fact rather than guessing.
Evaluation. Comparing outputs against criteria: which is more accurate, more complete, more appropriate for the context, and being able to say why in a sentence someone else can act on.
Research and verification. Establishing whether a statement holds up. Finding the source, reading it, and recording what it actually says.
Structuring and cleaning. Turning inconsistent material into a usable form without quietly dropping the awkward records — which is the most common way this work goes wrong.
Quality control. Reviewing other people's work, measuring where reviewers disagree, and pushing back when a guideline is producing nonsense.
Why it is skilled
The individual decisions are simple. Making ten thousand of them consistently is not. The qualities that predict good data work are not credentials:
- reading carefully enough to notice when something does not fit;
- following a written rule exactly, including when you disagree with it;
- being as careful on item nine hundred as on item nine;
- asking instead of guessing;
- working alone, remotely, without someone checking on you.
Done well, it is invisible. Done badly, it fails downstream in ways nobody can trace back. That asymmetry is precisely why consistency is valued above speed.
Who it suits
People who read closely and like closure on small problems. Translators, editors, librarians, researchers, lab technicians, paralegals, teachers, medical coders — and plenty of people with none of those backgrounds who simply pay attention.
It is genuinely remote and genuinely flexible, and those two properties are why it reaches people who are otherwise shut out of technical work.
If this sounds like you
If you are good at careful reading, research, annotation, comparison or verification, you may already have skills that data projects need. You do not need to call yourself an AI professional. You need to be able to do the work accurately.
See the Data Contributor path →
Domain expertise matters too. Some evaluation cannot be done correctly without someone who understands the subject — a clinician reading a finding, a lawyer reading a clause, an engineer checking whether a value is possible at all.
Applying does not guarantee work: opportunities depend on project needs, qualification and fit, and applicable compensation is stated before any work begins.