
Reinforcement learning fits one specific shape of problem: nobody can label the correct action in advance, and the only feedback available is the outcome of acting, usually after a delay. Routing a delivery fleet, ranking a chat reply, controlling a robot joint — none of them come with a right answer for every situation.
This guide covers the mechanics, the algorithm families and how to tell which one fits your problem, and reinforcement learning from human feedback (RLHF), the version of RL most ML engineers meet first. If you already know the agent-reward loop and came for the human-feedback part, skip to reinforcement learning from human feedback in LLMs; the sections before it build the vocabulary that section assumes.
What is reinforcement learning?
Reinforcement learning (RL) is the machine learning paradigm where an agent learns a behavior by acting in an environment and being rewarded or penalized for each action, rather than from labeled examples.
The reason the field is built this way is the feedback it has to work with. Most real decision problems deliver delayed, sometimes noisy signals about whether a whole sequence of choices worked out, never a correct answer for each step. RL is the branch built specifically for that setup [1].
It needs no labels, only a reward signal and the ability to try actions repeatedly, and it optimizes cumulative reward over time rather than any single action. That is what puts RL behind everything from game-playing systems to the human-feedback step that aligns large language models.
Core components of a reinforcement learning system
A reward with no policy to shape is just a number, and a policy with no reward signal has nothing to optimize toward. That is why the pieces of an RL system (agent, environment, state, action, reward) are worth taking in pairs rather than one definition at a time.
Agent, environment, state, and action
The agent is whatever makes the decision. The environment is everything it doesn't directly control, including the rules for what happens after it acts. The environment exposes a state, a snapshot of the current situation, and the agent responds by choosing from an action space: the set of moves available at that state.
A warehouse robot makes this concrete. Its state might include its position, what it's holding, and what's nearby; its action space includes moves like "move forward," "rotate," or "release." After each action, the environment returns a new state and a reward: a small penalty for a collision, a larger reward for placing an item correctly.
This is where RL diverges from supervised learning in practice. There's no dataset of "correct" robot trajectories to imitate, only a stream of attempts and outcomes the agent learns from directly.
Reward function, policy, and value function
A reward function and a value function answer two different questions. The reward function says how good this action was, right now. The value function says how good this situation is, accounting for everything that might happen afterward, including rewards that haven't arrived yet.
A chess sacrifice illustrates the gap: giving up a piece produces an immediate negative reward, but a value function that correctly estimates the resulting position rates it highly if the sacrifice leads to checkmate a few moves later. Sutton and Barto formalize this as expected return, the value of a state being the expected sum of future rewards, discounted so sooner rewards count more than distant ones [1].

The policy ties value to action: it's the agent's rule for mapping a state to a chosen action, and a well-trained policy consistently picks actions a good value function would endorse. Getting there means resolving a tension present in every RL system: the exploration-exploitation tradeoff.
How reinforcement learning works: the agent-environment loop
Whatever the algorithm sitting on top, the loop underneath is the same:
- The agent observes the current state of the environment.
- The agent selects an action, based on its policy.
- The environment returns a reward and a new state.
- The agent updates its policy using that reward.
- The cycle repeats, often thousands or millions of times, until the policy converges.

No single step teaches the agent anything on its own. The policy improves only across many repetitions, and reinforcement learning as a field is largely the study of making that convergence fast and stable across very different environments and reward structures.
The Markov decision process and exploration-exploitation tradeoff
Formally, this loop is a Markov Decision Process (MDP): a framework where the next state depends only on the current state and action, not the full history that led there. That property, the Markov property, is what makes the problem mathematically tractable, because the present state already contains what matters [1].
At the action-selection step, the agent has to choose between trying something new and repeating what has worked, and epsilon-greedy is the common answer. The agent picks the best-known action most of the time but chooses a random action with some small probability (epsilon), enough to keep discovering better strategies without abandoning what works. Set epsilon too low and the agent gets stuck on a mediocre policy; too high and training never converges. Most modern algorithms replace fixed epsilon schedules with adaptive exploration, but the tradeoff never goes away. It just moves into a different parameter.
Types of reinforcement learning
Reinforcement learning splits along two independent axes. The first is model-based versus model-free: a model-based agent builds an internal simulation of the environment and plans ahead using it, while a model-free agent skips the simulation and learns directly from trial and error. The second is online versus offline: online learning happens through live interaction with the environment, offline learning trains entirely on a fixed, previously collected dataset.
| Axis | Type | Learns from | Adapts to change | Typical use case |
|---|---|---|---|---|
| Model use | Model-based | An internal model of the environment | Quickly, if the model is accurate | Robot navigation with a known map |
| Model use | Model-free | Direct trial-and-error interaction | Slowly, but holds up when the model is wrong | Self-driving in unpredictable traffic |
| Online vs offline | Online | Live, ongoing interaction | Continuously | Real-time recommendation systems |
| Online vs offline | Offline | A fixed historical dataset | Not until retrained | Settings where live experimentation is unsafe or costly |
Because the axes are independent, the pairs combine, and most production systems do combine them: a model-free policy trained mostly online, pretrained offline on logged data to cut live-interaction cost. (Positive and negative reinforcement, from behavioral psychology, describe how a reward is shaped rather than a third axis.)

Key reinforcement learning algorithms
Which family to reach for depends on what you have to work with. Value-based methods, like Q-learning, estimate how good each action is and pick the best one. Policy-based methods skip that step and optimize the decision rule directly. Actor-critic methods combine both, using one network to act and a second to judge those actions.
Deep reinforcement learning is what happens when any of these families use a neural network instead of a lookup table, which becomes necessary once the number of states or actions is too large to store explicitly. That is true past almost any toy example.
Q-learning and other value-based methods
Q-learning estimates the action-value function, written $$Q(s, a)$$, for every state-action pair: the expected return of taking action a in state s, then acting optimally afterward [2]. It updates that estimate after every step using the reward just received plus the best estimated value of the next state:
$$$Q(s, a) ← Q(s, a) + α [ r + γ · max(Q(s', a')) − Q(s, a) ]$$$
Here α is the learning rate and γ the discount factor, weighing future rewards against immediate ones. Take a small grid-world with $$γ$$ = 0.9: if the agent gets reward $$r$$ = 1 for moving toward a goal and the best next-state value is 0.8, the update target is 1 + 0.9 · 0.8 = 1.72, and $$Q(s, a)$$ moves toward it by a step of size $$α$$.
SARSA (state–action–reward–state–action) is the close on-policy variant: it updates from the action the policy actually takes next rather than the best one available, producing safer but slower-converging policies in noisy or risky environments.
Policy gradient, actor-critic, and deep reinforcement learning
Policy gradient methods optimize the policy's parameters directly, using the gradient of expected reward, with no Q-table and no value-estimation step first. This scales to continuous action spaces, like a robot arm's joint angles, where a Q-table would need infinite entries.
Actor-critic algorithms pair the two approaches: an "actor" network selects actions, while a "critic" network estimates value and scores those choices, giving the actor a steadier, less noisy signal (lower variance) than pure policy gradient alone. Proximal Policy Optimization (PPO) is the actor-critic variant that constrains how much the policy can change per update; its original paper reports it outperforming other online policy-gradient methods on locomotion and Atari benchmarks [4]. It went on to become the algorithm used for the RL step in the first production RLHF pipelines [6].
Deep Q-Networks (DQN) extended Q-learning to use a neural network in place of a table, letting RL agents play Atari games directly from raw pixel input [3]. That shift, from tables to networks, is the difference between classroom RL and the deep RL used in production today.
In practice that comes down to four cases:
- A discrete action space small enough to enumerate: value-based, in practice DQN.
- Continuous or high-dimensional actions, like joint angles: policy-based, or actor-critic if the gradient is too noisy on its own.
- No environment you can safely experiment in: offline RL on logged data.
- No reward function anyone can write down: human feedback, which the next section covers.

Reinforcement learning vs. supervised and unsupervised learning
What separates the three paradigms is the kind of feedback the algorithm receives. Supervised learning requires a known correct output for every input. Unsupervised learning needs no labels and pursues no specific goal, finding structure in data for its own sake. Reinforcement learning sits in between: it has a goal, maximize reward, but no labeled correct action, only delayed reward signals.
| Paradigm | Feedback type | Data requirement | Example task |
|---|---|---|---|
| Supervised learning | Correct label per example | Labeled training data | Image classification |
| Unsupervised learning | None | Unlabeled data only | Customer segmentation |
| Reinforcement learning | Delayed reward signal | Interaction with an environment | Game playing, robot control |
This is why reinforcement learning is the right tool for sequential decision problems, where the right answer to any single step isn't known in advance, only the outcome several steps later. Supervised learning doesn't apply: there's no labeled training data describing the correct move at each intermediate step, only whether the whole sequence ended well or badly.
Reinforcement learning from human feedback in LLMs
For most ML engineers today, RLHF is the version of reinforcement learning they'll actually touch, more so than robotics or game-playing. It adapts the standard RL loop to a setting where no hand-coded reward function exists for "good" language: human raters compare model outputs, and a reward model learns to predict those preferences [5]. The policy, the language model itself, is then trained to maximize the score that learned reward model gives it.
One guardrail is built into that objective: a KL penalty against the pre-RLHF model, so the policy cannot drift far from its starting behavior while chasing reward-model scores [6]. Without it, the policy finds outputs the reward model loves and humans don't.
PPO is the classic choice for that RL step, though direct preference-optimization methods now replace it in many pipelines. DPO trains on the same human comparison data without a separate reward model or an RL loop at all [13]. Both approaches need the same input, which is where the real cost sits.
RLHF is the technique behind the alignment step in ChatGPT: a base model fine-tuned on instructions, then optimized on human preference data so outputs match what raters judged helpful [6]. Related systems use variants of it. Claude is trained with Constitutional AI, where AI-generated feedback against a set of written principles stands in for part of the human comparison signal [14].
Multilingual coverage is one place RLHF breaks down in practice. Public preference datasets are overwhelmingly English, and a reward model cannot score what it has no comparisons for, so the signal for any other language is collected rather than downloaded.
"Covering a language" is not one decision either. On an Arabic annotation project for LLM evaluation the dialect split was the constraint: Gulf and North African speakers can struggle to understand each other, so one undifferentiated Arabic pool yields judgements that are not comparable with one another. The recruitment failure to plan for is the false positive — someone living in an Arabic-speaking country who is not a native speaker. A profile does not catch that and a test task does, which is why proficiency was validated by test task rather than by profile.
How the RLHF training pipeline works
The pipeline runs in four stages. The base model generates multiple candidate responses to the same prompt. Human raters rank or compare those candidates rather than label them individually: the signal is which response is better, not whether any one of them is good in isolation. A reward model trains on those rankings to predict preference scores for unseen responses [5]. The policy is then optimized against that reward model to produce outputs it scores highly.

"Which response is better" is not one judgement. On a project annotating suggested replies in marketplace chats, better decomposed into four separate calls: relevance to the message being answered, absence of provocative content, contextual accuracy within the conversation, and stylistic correctness including the informal phrasing chat actually uses. The hard half was never spotting the bad reply. It was recognising one that stays neutral and context-aware without sounding synthetic, and that is a distinction a rubric can locate where a single opaque vote cannot.
The volumes involved are smaller than most people assume. InstructGPT's reward model trained on 33k prompts, with four to nine responses ranked per prompt, which yields six to thirty-six pairwise comparisons each [6]. That makes preference data a procurement question rather than a scraping one.
The bottleneck is rarely the optimization step. It's the consistency of the human preference rankings feeding the reward model: the same pair ranked differently by different raters, or by the same rater on different days, trains a noisy reward model that optimizes the policy toward the wrong target. A second rater checking ranking consistency before the rankings reach the reward model catches that disagreement early.

Real-world applications and use cases
Reinforcement learning is already in production or active research use across four domains, each of them a problem where the right answer depends on a whole sequence of decisions.
Robotics uses RL for manipulation and locomotion: a robotic arm learning to grasp irregular objects, or a legged robot learning to walk over uneven terrain, where no fixed dataset of correct movements covers every object and surface it will meet [7].
Autonomous vehicles use RL mainly in planning and control, deciding when to merge, how to handle an unprotected left turn, how to balance smoothness against safety. These are sequential decisions under uncertainty that a fixed rulebook struggles to cover [8].
Recommendation systems use RL to optimize for long-term engagement rather than the next click alone, treating a session as a sequence of actions and rewards [9].
Game-playing systems offer the clearest public proof RL works at scale: AlphaGo combined deep neural networks with tree search and self-play, training by playing against copies of itself, to beat top human players at Go [10].
None of these needed RLHF's human-preference framing. They use task-defined reward functions instead, which is the cleaner reward-design case. Defining that function is where production RL usually breaks, and that is the next section.
Common challenges and limitations of reinforcement learning
Textbook RL and production RL diverge sharply, and most of the gap comes down to four obstacles.
Sample inefficiency is the first. Many RL algorithms need millions of interactions to converge, which is fine in a simulator but often infeasible on physical hardware or with real users.
Reward design is the second and the trickiest. An agent optimizes exactly what the reward function measures, not what the designer intended, and a slightly misspecified reward can produce a strategy that scores well while missing the actual goal. That failure has a name: reward hacking [11].
The third is the simulation-to-reality gap. A policy trained in simulation often degrades sharply on real hardware or with real users, because the simulator never captured every variable: lighting, friction, the long tail of unusual situations a synthetic generator didn't anticipate.
One thing that narrows the third gap is making the simulated environment and the real recording the same place rather than two independent builds.
Scanning the rooms where the egocentric footage was captured is where we ran into what that costs. Lighting is the hard part, and its three parameters pull against each other with no universal setting: a high ISO reduces the light you need but adds grain, and the reconstruction then stops matching points between frames; a wide aperture admits light but narrows depth of field, so part of the object defocuses and reconstruction degrades there; a long exposure needs the camera perfectly still. Shooting parameters end up calibrated per object class, which is the line item to plan for rather than the number of scans. The scanned environments are also static, drawers and cabinet doors do not open, and for surface manipulation that turns out not to matter.
The fourth is interpretability. Explaining why a trained policy chose a specific action is often hard even for the team that built it, which complicates debugging and safety review in high-stakes deployments.
Each has a standard mitigation, none of them free:
- Pretrain offline on logged data before any live fine-tuning, which buys back part of the sample cost.
- Test the reward function against adversarial and edge-case scenarios before committing to a full training run.
- Mix real-world data into training instead of relying on simulation alone.
- Restrict high-stakes deployments to well-tested action spaces. Policy visualization helps with interpretability, but there is no clean answer here yet.
Conclusion and next steps
Reinforcement learning is a paradigm for sequential decision-making under reward feedback, not a single algorithm. Its trajectory runs from tabular methods like Q-learning, through deep RL that scales to images and continuous control, to RLHF and its preference-optimization successors.
The two limits that have not moved are reward design and the sim-to-real gap. Both are the same problem at different scales: the distance between what you measured and what you meant.
For hands-on practice, the open-source Gymnasium library, maintained by the Farama Foundation, provides standard environments for testing RL algorithms without building a simulator from scratch [12].
If you need human-preference or evaluation data collected in a language your reward model has no comparisons for — start here.
- AI Training
- Robotics
- Data Annotation
- Python
- AWS
Frequently Asked Questions (FAQ)
Reinforcement learning is a way for software to learn by trial and error. An agent takes actions in an environment, gets a reward or penalty for each one, and adjusts its behavior to earn more reward over time. There’s no list of correct answers to memorize. The agent discovers what works by acting, observing the result, and adjusting.
Supervised learning trains on labeled examples with a known correct answer for each input. Reinforcement learning has no labels, only a goal and a reward signal that arrives after the agent acts, often with a delay. Supervised learning answers “what’s correct for this input”; reinforcement learning answers “what sequence of actions maximizes reward over time.”
Every RL system has five components: the agent (decision-maker), the environment (everything outside the agent), the state (a snapshot of the current situation), the action (what the agent can do), and the reward (feedback on the action just taken). Policy and value function describe how the agent decides and judges long-term outcomes.
A Markov Decision Process (MDP) is the mathematical framework most reinforcement learning problems are built on. Its defining assumption, the Markov property, is that the current state carries everything the agent needs, so the path that led there can be discarded. That is what makes the problem solvable.
A model-based agent keeps its own simulation of the environment and plans against it. A model-free agent has no such simulation and improves only from what actually happened. Model-free is slower to adapt, but it holds up better when the environment is too complex to model accurately, such as unpredictable real-world traffic.
Q-learning is a value-based algorithm that estimates the expected long-term reward of taking a specific action in a specific state, the action-value function. It updates that estimate after every action using the reward received plus the best estimated value of the resulting state, converging on values the agent can act on. Once the state space is large, the deep variant (DQN) holds those estimates in a neural network instead of a table.
Deep reinforcement learning combines reinforcement learning with neural networks, replacing a lookup table of values with a network that handles far larger or continuous state and action spaces. It made it possible to train agents directly from raw inputs like pixels, and underlies most modern RL systems used in production today.
The exploration-exploitation tradeoff is the agent’s ongoing choice between trying a new, untested action (exploration) and repeating an action already known to work (exploitation). Too much exploration wastes time on poor choices; too much exploitation risks settling for a mediocre strategy. Most algorithms manage this with a tunable parameter, such as epsilon in epsilon-greedy methods.
RLHF fine-tunes a model using human preference rankings instead of a hand-coded reward function. Human raters compare model outputs, a reward model learns to predict those preferences, and the policy is optimized against that reward model. Newer preference-optimization methods such as DPO skip the separate reward model and train on the comparisons directly. It is the technique behind the alignment step in ChatGPT; other production systems use variants that substitute AI-generated feedback for part of the human signal.
Reinforcement learning is used in robotics for manipulation and locomotion, in autonomous vehicles for planning and control, in recommendation systems to optimize for long-term engagement, and in game-playing systems like AlphaGo. Its most widespread current application is RLHF, which aligns large language models with human preferences and reaches far more users than any robotics or game-playing deployment.
The main challenges are sample inefficiency (needing huge amounts of interaction data), reward design (a poorly specified reward can produce reward hacking, unintended behavior that still scores well), the simulation-to-reality gap (policies trained in simulation often degrade on real hardware), and interpretability (explaining why a policy chose a given action).
No. Deep learning is a technique using neural networks; reinforcement learning is a paradigm defined by how feedback works, which is reward-based rather than labeled. The two combine in deep reinforcement learning, where a network estimates values or policies inside an RL system, but deep learning also exists outside RL, in supervised tasks like image classification.
Further Reading & References:
- [1] Sutton, R.S., Barto, A.G. "Reinforcement Learning: An Introduction" (2nd ed.) — MIT Press — 2018
- [2] Watkins, C.J.C.H., Dayan, P. "Q-learning" — Machine Learning, Vol. 8 — 1992
- [3] Mnih, V., Kavukcuoglu, K., Silver, D., et al. "Human-level control through deep reinforcement learning" — Nature — 2015
- [4] Schulman, J., Wolski, F., Dhariwal, P., et al. "Proximal Policy Optimization Algorithms" — arXiv — 2017
- [5] Christiano, P., Leike, J., Brown, T., et al. "Deep Reinforcement Learning from Human Preferences" — NeurIPS — 2017
- [6] Ouyang, L., Wu, J., Jiang, X., et al. "Training Language Models to Follow Instructions with Human Feedback" — NeurIPS — 2022
- [7] Kober, J., Bagnell, J.A., Peters, J. "Reinforcement Learning in Robotics: A Survey" — The International Journal of Robotics Research — 2013
- [8] Kiran, B.R., Sobh, I., Talpaert, V., et al. "Deep Reinforcement Learning for Autonomous Driving: A Survey" — IEEE Transactions on Intelligent Transportation Systems — 2022
- [9] Afsar, M.M., Crump, T., Far, B. "Reinforcement Learning Based Recommender Systems: A Survey" — ACM Computing Surveys — 2022
- [10] Silver, D., Huang, A., Maddison, C.J., et al. "Mastering the Game of Go with Deep Neural Networks and Tree Search" — Nature — 2016
- [11] Amodei, D., Olah, C., Steinhardt, J., et al. "Concrete Problems in AI Safety" — arXiv — 2016
- [12] Gymnasium Documentation — Farama Foundation
- [13] Rafailov, R., Sharma, A., Mitchell, E., et al. "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" — NeurIPS — 2023
- [14] Bai, Y., Kadavath, S., Kundu, S., et al. "Constitutional AI: Harmlessness from AI Feedback" — arXiv — 2022