Offline RL Conservative Policies and Out-of-Distribution Action Penalization

Conservative policies prevent overconfident decisions by penalizing actions outside training data.

Cover illustration for “Offline RL Conservative Policies and Out-of-Distribution Action Penalization”
Written by
Saoirse Ní FhaoláinContributing Writer
Published
October 10, 2026
Reading time
11 min read
Sources cited
8 sources ↓

Offline reinforcement learning asks an agent to learn a good policy from a fixed batch of logged data, with no chance to try things out and see what happens. The entire project runs into one hard fact: Q-learning, the workhorse algorithm behind most of modern RL, was never built for a setting where exploration is off the table, and applying it without modification produces policies that look good on paper and fail in deployment.

Why offline RL cannot simply borrow standard Q-learning

Standard Q-learning stays safe because it operates inside a loop. An agent tries an action it hasn't tried before, the environment responds, and that response, whether reward or penalty, feeds straight back into the value estimate. Bad guesses get corrected almost as soon as they're made. Offline RL removes that loop. The dataset is fixed before training starts, so when the Q-function assigns a high value to some action it has barely seen, no environment steps in to say otherwise. The error just sits there, and it gets reinforced.

The specific mechanism runs through the Bellman backup. To compute a target value, the update rule samples actions from the policy currently being learned, not from the policy that generated the data. But the Q-function has only ever been trained on state-action pairs that the original behavior policy actually produced. When the learned policy drifts toward actions outside that footprint, the Q-function has to evaluate something it has no real basis for judging. It drifts systematically toward exactly the actions the Q-function is least equipped to judge, and each gradient step tightens that bias.

Bootstrapping makes this dangerous rather than merely imprecise: one inflated estimate at a given state doesn't stay contained. It feeds into the target value used to update a different state in the next step, which feeds into the step after that, so a single bad guess propagates through the whole value function. Left running, this produces a policy that scores well against its own, badly miscalibrated internal measure of performance, right up until it's handed real decisions to make and selects actions that are disastrous in practice. The gap between training performance and deployment performance is the entire problem offline RL has to solve.

Domains that motivated conservative offline RL

Offline RL matters most in exactly the settings where trial and error is too costly to permit. The domains that benefit most from offline learning are also the ones least able to absorb an overestimation error. Autonomous driving is a case in point: policies there are built from historical driving logs because sending a car out to generate new exploratory data means real risk on real roads. Robotic surgery and sepsis treatment carry the point further. An exploratory action in those settings can mean a damaged instrument or a harmed patient rather than a few points of reward, so conservative behavior during learning is a requirement.

Much of the experimental proof that conservative offline methods work outside simulation has come from robotics. CODAC (Ma et al., NeurIPS 2021) was tested on robot navigation tasks. ETH Zurich's RWM-U research went further, deploying uncertainty-penalized offline policies on physical hardware, a quadruped (ANYmal D) and a humanoid (Unitree G1), and training those policies entirely from existing datasets before putting them on the actual machines. The resulting policies outperformed online model-free baselines, and the researchers describe it as, to their knowledge, the first demonstration of uncertainty-penalized offline model-based RL running full-scale tasks on physical robot hardware.

The stakes made OOD action penalization a central research question. If a field builds policies for driving, surgery, and physical robots, it needs a principled way to keep those policies from acting on inflated confidence in actions the data never actually supports.

How CQL establishes the conservative Q-learning baseline

Kumar et al. introduced Conservative Q-Learning in 2020, and with it the field got its first systematic answer. CQL adds a regularizer to the standard Bellman objective, and that regularizer does something specific: it provably pushes the learned Q-function below the true Q-function for every state-action pair actually present in the dataset. So overestimation, the central failure mode described above, gets swapped out for a guaranteed underestimate instead.

The mechanism penalizes Q-values for actions the offline data doesn't cover well, but it leaves the update rule for well-covered, in-distribution pairs untouched. In expectation, the learned Q-function sits below the true value for in-dataset pairs, a pointwise lower bound that leaves the policy nothing to exploit. It can't chase a phantom high value, because the regularizer has already suppressed it. The strength of that suppression is set by a single scalar, α, which can be fixed in advance or tuned automatically through dual gradient descent, giving practitioners one lever to balance caution against performance.

Part of why CQL spread as fast as it did is practical: Kumar et al. show it drops into existing deep Q-learning and actor-critic implementations without requiring a different training pipeline. On both discrete and continuous control benchmarks, it beat the offline RL methods that came before it by a wide margin, particularly on datasets with complex, multi-modal action distributions. CQL became the reference point that nearly every later method in this space defines itself against.

Extending CQL to risk-sensitive objectives: what CODAC adds

CQL's guarantee covers expected return, a single number that summarizes the average outcome a policy can expect. For the domains described above, that average can hide the exact failure the field is trying to prevent: a policy with a strong mean return can still carry a long tail of rare, severe failures, and expected-value pessimism alone won't catch that. CODAC (Ma et al., NeurIPS 2021) extends CQL's logic to the entire distribution of possible returns, not just its mean.

The mechanism adapts distributional reinforcement learning to the offline setting by penalizing the predicted quantiles of the return for out-of-distribution actions. The theoretical payoff is stronger than CQL's: CODAC converges to a uniform lower bound across all quantiles of the return distribution in finite MDPs, a guarantee that covers CVaR and other risk-sensitive objectives CQL's expected-value bound can't reach. The penalization stays data-driven, pushing OOD actions down more sharply than in-distribution ones, so the separation between the two widens.

The results bear out the theory. On two demanding robot navigation tasks, CODAC learns risk-averse policies from offline data that was collected entirely by risk-neutral agents, a setting where every baseline tested alongside it failed. On the D4RL MuJoCo benchmark, CODAC performs strongly on both expected-return and risk-sensitive metrics at once, which matters: it shows that building in risk-awareness doesn't cost average performance to get it.

The hidden cost of uniform penalization: how blanket pessimism degrades policy quality

CQL and CODAC both share a design choice that looked, at first, like the obvious right answer: penalize OOD actions uniformly, regardless of how far outside the data's support they actually sit. That choice carries a cost that later research has worked to isolate. If you treat all OOD actions as equally suspect, the penalty doesn't stay contained to the actions that deserve it. It bleeds into nearby in-distribution actions too, so it drags down Q-value estimates for actions the policy should actually be taking.

One concrete consequence is the loss of trajectory stitching. Offline datasets often mix trajectories of different quality, so to get the best policy you often need to combine the good parts of several sub-optimal runs. So doing that stitching usually means taking an action that looks OOD in isolation but leads somewhere genuinely valuable. Blanket penalization treats that kind of action the same as a truly dangerous one and suppresses both.

Mildly Conservative Q-Learning names this pattern directly: CQL-style penalties operate implicitly on OOD actions, and the downstream effect is a value function that stays overly conservative and overly tied to the behavior policy, even in cases where some OOD action would have been perfectly safe to take. The cost compounds as the action space grows or the task becomes harder, since the approximator's uncertainty about far-away regions increases and the collateral damage to good, in-distribution predictions grows along with it. CQL and CODAC remain the correct first answer to overestimation. What they also did was expose a structural tension between safety and flexibility that the rest of the field has since taken up directly.

Diagnosing where overestimation comes from: the PARS reframing

Most of the methods discussed so far treat overestimation as a consequence of missing data: the Q-function is uncertain about actions it hasn't seen, and pessimism is the correct response to that uncertainty. PARS proposes a different diagnosis. It argues the real driver of OOD overestimation is a specific architectural property of ReLU networks, not an unavoidable gap in the data.

Standard ReLU networks extrapolate linearly once inputs move past the range they were trained on. So if you apply this to a Q-function, Q-values for actions far outside the data's support don't decay toward some cautious baseline. They can grow without any natural bound, and the network effectively treats those far-OOD actions as if they resembled familiar, in-distribution ones. If that's the real cause, pessimism regularizers are compensating for a symptom rather than fixing the underlying defect in the function approximator.

PARS responds with two mechanisms. Reward scaling combined with layer normalization (RS-LN) weakens the similarity the network perceives between in-distribution and OOD inputs. A separate penalty on infeasible actions (PA) pins Q-values for far-OOD actions down to a fixed, low target directly. The bet underlying both pieces is that critic regularization alone, without generative models or task-specific machinery, is enough to produce robust behavior in both pure offline training and offline-to-online transfer. The empirical case for the diagnosis is the AntMaze Ultra benchmark, where PARS is the only method among those evaluated that enables successful offline-to-online training, a task where the competing approaches stall.

This is a live disagreement, not a settled one. If PARS is right that the mechanism is architectural, it reframes the collateral damage described in the previous section: the bleed from OOD penalties into in-distribution Q-values may be less about algorithm design and more about a network architecture mismatched to the problem it's being asked to solve. Whether that diagnosis holds across a wider range of tasks than AntMaze Ultra is still an open question, but it has shifted some of the field's attention from data coverage toward the function approximator itself.

Changing what gets constrained: state-based and outcome-based flexibility

Every method covered so far, CQL, CODAC, PARS, works by adjusting how harshly the value function penalizes specific actions. A separate line of research changes what gets constrained in the first place, moving the restriction from actions to states. Instead of asking whether a particular action resembles something in the dataset, these methods ask whether the state that action leads to is one the data can vouch for.

StaCQ, from Charles Alexander Hepburn, Yue Jin, and Giovanni Montana, allows the policy to take actions outside the data's support as long as those actions lead back to states that are in-distribution. That gives the agent room to find shortcuts and more efficient paths through a task without ever leaving territory the data has validated as safe. ODAF, from Ke Jiang, Wen Jiang, Yao Li, and Xiaoyang Tan, makes the same move from a different angle, arguing that agents should be constrained by where they end up rather than by the specific route they took to get there, which permits stitching high-value paths together even out of low-quality, non-expert demonstrations. Strategically Conservative Q-Learning (SCQ), from Yutaka Shimizu, Joey Hong, Sergey Levine, and Masayoshi Tomizuka, takes a related but distinct position: rather than reconstructing the constraint around states, it trusts the neural network's own capacity to interpolate between known data points, suppressing values only for actions genuinely far from anything in the dataset and letting the network generalize freely in between.

The Lacuna research direction survey (frozen 2026-10-04) treats this shift from action-level constraints to state-level or outcome-level targets as the idea that organizes the current generation of offline RL methods. The intuition behind all three approaches is the same: how densely an action is represented in the data isn't the same thing as whether that action is safe. A state with no exact match at the action level can still be entirely safe to pass through if it sits between two states the data documents well and rates highly.

Selective regularization: how DOSER distinguishes OOD actions worth suppressing from those worth encouraging

The state-constrained methods above change what's being constrained. DOSER goes a step further and changes how the penalty gets decided in the first place, by trying to tell which specific OOD actions deserve suppression and which deserve encouragement, rather than treating the whole OOD category as one undifferentiated block.

Doing that requires two things: a detector that can reliably flag when an action is OOD, and a way to judge whether a flagged action actually leads somewhere good. DOSER, published 2026-05-06, builds both using diffusion models. It trains two separate diffusion models, one over the behavior policy's action distribution and one over the state distribution, and uses single-step denoising reconstruction error as its OOD signal. That choice matters because older approaches to this problem leaned on another generative model for the same purpose, and its reconstruction error is a noisier signal, prone to misclassifying actions as in-distribution or out of it. Diffusion reconstruction error gives a sharper, more dependable boundary.

During policy optimization, DOSER uses its state model to evaluate what the predicted next state looks like for a flagged OOD action, then decides on that basis whether to suppress it or let the policy pursue it. Risky OOD actions get suppressed. OOD actions that lead to high-value states get left alone, and that is precisely the exploration that blanket penalization methods shut down. The approach carries formal backing as well: DOSER is proven to be a γ-contraction with a unique fixed point and bounded value estimates, and it carries an asymptotic performance guarantee relative to the optimal policy, conditional on bounded errors in both the learned model and the OOD detector. The gains over earlier penalization methods are largest on suboptimal datasets, exactly where uniform penalization does the most harm, since a poor behavior policy combined with blanket suppression leaves almost no room to improve on the data. By sparing the beneficial OOD actions that lead to strong outcomes, DOSER lets an offline policy get more out of limited data than methods built purely around caution.

Lightweight alternatives: PANI's noise-injection approach to OOD penalization

DOSER's accuracy comes at a cost: training and running two diffusion models alongside the main policy adds real computational weight to the pipeline. PANI shows that selective OOD penalization can work without generative modeling. Its noise-injection formulation reaches comparable protection against harmful OOD actions through a much simpler mechanism, so you don't need to train auxiliary diffusion models just to represent the behavior policy and the state distribution separately. Set alongside DOSER, PANI marks the other end of a spectrum the field is now actively exploring: methods differ not just in how conservative they are, but in how much machinery they need to decide where that conservatism should fall.

Methodology & sources

  1. Conservative offline distributional reinforcement learning

    Provided the NeurIPS 2021 CODAC paper (Ma et al.) on conservative offline distributional reinforcement learning, covering the quantile-based penalization and robot navigation results discussed in the article.

  2. Conservative Offline Distributional Reinforcement Learning

    Source for CODAC's mechanism of penalizing predicted return quantiles for OOD actions and its convergence guarantee as a uniform lower bound across quantiles.

  3. Mildly Conservative Q-Learning for Offline Reinforcement Learning

    Source for Mildly Conservative Q-Learning, which identifies the collateral damage of CQL-style uniform penalization on in-distribution Q-values and trajectory stitching.

  4. Overcoming Excessive Conservatism in Offline Reinforcement Learning — Lacuna

    Provided the research direction survey framing the shift from action-level to state-level or outcome-level constraints as the organizing idea of current offline RL methods.

  5. Offline Reinforcement Learning with Penalized Action Noise Injection

    Source for the PANI noise-injection method as a lightweight alternative to generative-model-based OOD penalization.

  6. Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning

    Source for the DOSER method using diffusion-based OOD detection and selective regularization to distinguish harmful from beneficial OOD actions.

  7. Conservative Q-Learning for Offline Reinforcement Learning

    Original CQL paper by Kumar et al. establishing the conservative Q-learning baseline with a provable lower-bound guarantee on Q-values for in-dataset state-action pairs.

  8. Strategically Conservative Q-Learning

    Source for Strategically Conservative Q-Learning (SCQ), which suppresses values only for actions genuinely far from the dataset while allowing free interpolation between known data points.

Saoirse Ní Fhaoláin

Contributing Writer

Saoirse reports on real-world deployments of reinforcement learning from human and AI feedback, with a particular focus on policy, safety, and the gap between benchmark performance and production behavior. She spent several years covering AI governance for a Dublin-based technology policy outlet before joining The Calibration Review.