Language models learn from labeled examples, and somebody has to build those examples correctly, one edge case at a time. That is the core of NLP data curation work, less glamorous than the phrase "AI training data" suggests, and more exacting than most people expect going in.
In this role, you will work through NLP-related datasets, applying detailed guidelines to label, review, or correct text data. Much of the work is repetitive by design, since consistency across thousands of similar items matters more than any single judgment call. Where the guidelines run out, and they always eventually do, you flag the edge case for review rather than guessing at an answer.
A typical batch might mix straightforward items, ones that fit the guidelines cleanly, with a handful that genuinely do not: ambiguous phrasing, sarcasm that reads as sincere out of context, a sentence that technically fits two different categories. Senior curators are relied on to handle those harder items directly rather than escalating everything, and to write up the reasoning clearly enough that it helps refine the guideline for everyone else on the project.
You do not need a college degree for this role. A high school diploma or equivalent is the minimum education required, and a technical degree is treated as a plus rather than a requirement. What matters more is 24 months of relevant experience in NLP data work, annotation, or a closely related task, along with the attention to detail needed to catch the difference between a genuine edge case and a mislabeled item from a previous batch.
People who do well here tend to have a certain patience for the work itself, and they notice when something feels off even when it technically matches the guideline on paper. That instinct is hard to teach, and it is usually what separates a curator whose batches need rework from one whose output is trusted without a second look.
This is a part-time position, and pay is set at $101,000 on an annualized, full-time-equivalent basis, prorated against the hours actually worked. Roles like this one are often structured as contract or hourly work rather than traditional employment, and where that is the case, standard benefits typically do not apply. Where the position is offered on a full-time basis instead, health coverage, paid time off, and retirement plan matching are included.
The work is fully remote and open to candidates anywhere, with no office or specific location tied to it. Because most of the work is guideline-driven rather than meeting-driven, there is no requirement to be online during fixed hours. You are expected to hit agreed throughput and quality targets within each work period, and a small amount of async communication handles the rest, mostly flagging questions and receiving updated guidance.
Remoteroles posts a steady stream of AI and data-work roles like this one, and NLP curation openings in particular tend to fill from candidates who already have some annotation experience, even informal experience, rather than requiring a specific credential. Project length varies, some run for a few weeks, others extend for months, and curators who consistently produce clean, well-reasoned work are often invited back onto new projects directly rather than reapplying each time.
The datasets themselves vary a fair amount too. Some projects focus on sentiment or intent labeling, others on more technical tasks like entity tagging or evaluating whether a generated response actually answers the question it was given. Guidelines differ across each one, so part of the job, especially in the first few days of a new project, is genuinely internalizing a fresh rulebook rather than assuming the last project's conventions still apply.
Quality checks run continuously rather than only at the end of a batch. A sample of your work typically gets reviewed against a gold-standard answer set, and consistent accuracy above the target threshold is what keeps a curator on the higher-visibility, better-paying projects rather than the entry-level queue. That feedback loop is usually fast enough that you know within a day or two whether you are calibrated correctly on a new guideline.
Because the role is part-time and project-based, the workload is not perfectly steady week to week. Some stretches bring a full queue and tight deadlines, others are lighter while a new project spins up. Curators who plan around that variability, rather than expecting a fixed weekly hour count, tend to have an easier time with the schedule than those who need strict predictability.
Feedback loops in this line of work tend to be tighter than in most jobs. Instead of a formal annual review, you find out fairly quickly, through accuracy scores and direct notes from a project lead, whether your calibration on a specific guideline is where it needs to be, which makes it easier to correct course early rather than discovering a systemic issue weeks into a project.
Tooling for this kind of work is usually straightforward once you learn the interface, a browser-based annotation platform with a queue, a guideline reference panel, and a way to flag or comment on a specific item. Most curators are fully comfortable with the tools within the first few days, since the harder skill is judgment, not software.
To apply, submit a resume noting any prior data-labeling, annotation, or NLP-related experience, even if it was informal or project-based. Shortlisted applicants typically complete a short paid sample task as part of the assessment before a final decision is made.