Blog
Essays and field guides on AI training, evaluation, annotation and the economy of human expertise behind modern AI systems.
31 essays
If you have been searching for ways to earn money using your academic background, you have probably seen these three terms scattered across listings and platform descriptions. AI training. Data labeling. Data annotation.
If you have read anything about how modern AI assistants are built, you have probably seen the acronym RLHF. It stands for Reinforcement Learning from Human Feedback. It is the method that turned early language models into useful, safe, and coherent AI systems.
The AI training market is growing fast and it is actively looking for people with STEM backgrounds. But most graduates do not know how to enter it, what to expect, or how to position their degree as an asset rather than just a credential.
Most people applying for data annotation and AI evaluation roles focus on the wrong things. They optimise for speed, or they list their credentials and expect that to be enough. Neither reflects what AI companies actually need from the people who evaluate and annotate their training data.
If you are considering AI data labeling or evaluation work, one of the most practical things you can understand upfront is how your work gets scored. Knowing what quality metrics you are being measured against, and how those measurements affect your access to work and pay, changes how you approach every task.
Pay ranges for AI training and evaluation work vary enormously depending on where you look. Some listings show $10 per hour. Others show $100 or more. Both figures are accurate, and the gap between them is almost entirely explained by one variable: the depth and verifiability of your domain expertise.
Red teaming is one of the more unusual categories of AI evaluation work. Where most AI training tasks involve assessing whether a model's outputs are accurate or helpful, red teaming specifically involves trying to make the model fail. You are looking for weaknesses, not strengths.
Image annotation is one of the largest categories of AI training work by volume. Every computer vision application, from autonomous vehicle perception systems to diagnostic radiology AI to satellite land classification tools, relies on human annotators who have labeled the images the model learned from.
Named Entity Recognition, commonly abbreviated NER, is one of the foundational tasks in natural language processing. It is also one of the most common text annotation tasks in AI training work, with applications spanning clinical documentation, legal analysis, financial research, scientific literature, and media monitoring.
The AI training market has a talent problem. Not a shortage of workers. A shortage of workers whose knowledge is deep enough to be genuinely useful for the tasks that matter most.
Remote work in AI is not just for software engineers. Science and engineering graduates with no coding background are accessing well-paid remote roles through the AI training and evaluation market, and the demand for their specific knowledge is increasing faster than the supply of people who have it.
Most people entering the AI evaluation market assume it works like most online freelance markets: more hours means more money, speed is rewarded, and experience accumulates in a straightforward way over time. The AI evaluation market does not work like that. The variable that matters most is not how fast you work or how many hours you log. It is how rare and verifiable your knowledge is.
AI hallucinations are one of the most discussed problems in modern AI development, and one of the least precisely understood outside of technical circles. The term gets used loosely to describe anything an AI gets wrong, but that conflates several distinct failure modes that have different causes, different consequences, and different solutions.
Postgraduate students are in an unusual position in the AI training market. They have exactly the kind of specialist knowledge that commands a pay premium, a schedule with more flexibility than most employment, and a financial situation where meaningful supplemental income matters considerably.
Most people who use AI tools have no clear picture of how they were built. That gap matters more than it used to, because understanding the process helps you understand what AI can reliably do, where it tends to fail, and why human expertise remains central to producing AI that works.
Credentials have always been a proxy for capability. A degree from a recognised institution signals that someone spent time studying a subject seriously, passed assessments, and had their knowledge verified by an academic body. For most of the twentieth century, that proxy worked well enough that employers, clients, and institutions used it as the primary filter for hiring and access.
Most people who apply for AI evaluation roles think the main requirement is knowing their subject. That is necessary but not sufficient. The evaluators who access the best projects and earn the most over time are the ones who combine domain knowledge with a specific set of skills that are not taught in any degree programme but can be developed deliberately.
AI safety evaluation is one of the fastest-growing categories of work in the AI training market, and one of the least well understood by people outside the industry. It is often conflated with red teaming, or dismissed as a niche concern of AI researchers. In practice, it is a broad and growing category of work that draws on expertise from medicine, law, social science, ethics, and technical domains. Understanding what it involves and who is qualified to do it matters for anyone thinking seriously about AI evaluation as a professional path.
One of the most persistent predictions in AI commentary is that human oversight of AI systems will diminish as those systems become more capable. The logic seems intuitive: if an AI can perform a task better than a human, why keep the human in the loop?
Completing a dissertation or thesis is, among other things, a training programme in a specific set of skills: reading critically, evaluating evidence, identifying weaknesses in an argument, writing precisely under constraints, and working independently on problems that do not have a clean answer at the start.
Prompt engineering has gone from a niche technical concept to one of the most searched terms in AI careers. The interest is genuine: companies building AI systems need people who can craft inputs that reliably produce useful outputs, test where models fail, and create the training examples that make models better over time.
AI agents are the most discussed development in applied AI right now. Every major AI company has announced an agentic product. Every major enterprise software vendor is building agent capabilities into their platforms. The market has moved on from the question of whether agents will matter to how to make them reliable.
The pitch for AI training as a side income sounds simple: use your expertise, work from anywhere, earn well. That pitch is accurate as far as it goes. What it leaves out is how the work actually operates, where the gaps are, and what separates people who build a genuinely useful income stream from those who sign up, complete a few tasks, and drift away.
AI deployment in professional services is moving faster than most people outside those industries realise. The products being built and sold today are not demonstrations or prototypes. They are clinical documentation tools used in hospitals, contract review systems deployed at law firms, and financial analysis platforms integrated into investment workflows.
Most of the AI training content written for a general audience focuses on language: text in, text out. That is where the market started and where most evaluation work currently sits. But the AI systems being built and deployed today increasingly handle multiple types of input simultaneously, and the evaluation of those systems requires a correspondingly broader set of human expertise.
No AI experience is not the same as no relevant experience. The AI training market is not looking for people who have worked in AI before. It is looking for people with genuine knowledge in the domains where AI is being built: medicine, law, engineering, science, finance, linguistics, software development.
Synthetic data has moved from a niche technique to a central pillar of AI development strategy at major AI labs. Understanding what it is, why it is used, and where it still depends on human expertise is useful context for anyone working in or thinking about the AI training market.
Most coverage of the AI job market focuses on the engineering side: the machine learning researchers, the software engineers, the data scientists building the systems. The parallel market for the human expertise that trains and evaluates those systems receives less attention. That imbalance in coverage has created a gap between how large that market actually is and how widely it is understood.
People with specialist expertise who are thinking about flexible, remote income typically have a choice between traditional freelancing, consulting, tutoring, or contracting in their field, and AI evaluation work. These are not mutually exclusive, and many people do both. But understanding where AI evaluation work fits in comparison to the alternatives helps you decide how much to invest in it and how it fits alongside your other work.
Every discussion of the AI talent market focuses on engineers. Who is building the models, what the competition for ML researchers looks like, how much companies are paying for AI product managers. The framing is almost always supply and demand for the people building AI systems.
Every AI model has two lives. The first is training, where it learns. The second is inference, where it works. The simplest way to hold the distinction in your head is a student and an exam. Training is the years of study: slow, expensive, and done once. Inference is exam day: the student walks in with everything already learned and simply applies it, question after question. The two phases use the same underlying mathematics but differ in almost every way that matters: cost, hardware, timescale, and who pays for them. Understanding the split explains most of what you read about AI economics, from why frontier models cost hundreds of millions to build to why using one costs a fraction of a penny.