Skip to content
Everything *[NYC] 2026: see what we announced. →

RLCD (Reinforcement Learning for Calibrated Decisions) definition

Reinforcement Learning for Calibrated Decisions (RLCD) is a reinforcement learning post-training method in which the reward is tied to whether a model's stated probability matches how often its answer turns out to be correct, rather than to human preference ratings or automatic verification. The term was coined by TypeSafe AI for its Jev model in September 2026; the objective is public, the mechanism is not.


A diagram explaining RLCD (Reinforcement Learning for Calibrated Decisions) in terms of related concepts.

What is Reinforcement Learning for Calibrated Decisions (RLCD)?

Why does calibration make sense as a training objective?

How is RLCD different from RLHF and RLVR?

Is RLCD the same as Reinforcement Learning from Contrastive Distillation?

What does RLCD not do?

What does RLCD mean for the content you feed a decision model?

Unlock New Possibilities with Sanity

With RLCD (Reinforcement Learning for Calibrated Decisions) under your belt, it's time to see what Sanity can do for you. Explore our features and tools to take your content to the next level.

Last updated: