Temperature Scaling vs Platt Scaling for Neural Classifiers
Temperature scaling needs one parameter; Platt scaling needs two.

A neural network can classify images or text with impressive accuracy and still produce confidence scores that are systematically wrong. These are separate properties. Calibration asks something accuracy never does: when the model expresses a given level of confidence, is it actually correct that often? A classifier can answer correctly on nearly every test case and still assign confidence values that bear no relationship to the true probability of being right, and no amount of inspecting top-line accuracy will reveal that gap. Closing it requires a distinct step, applied after training, that maps the model's raw outputs, its logits, onto probabilities that match empirical frequency. This step is required because of how large networks are optimized. Modern classifiers are trained by minimizing cross-entropy loss on a finite training set, and that objective rewards confident, low-entropy predictions far more than it rewards probability estimates that track reality. An over-parameterized network can drive training loss to near zero long after it has stopped improving on held-out accuracy, and in doing so it keeps pushing its output probabilities toward the extremes of 0 and 1. The model's predictions look assertive while its confidence scores are, in a precise statistical sense, untrustworthy. In domains where a downstream system has to decide whether to trust a model's output or defer to a sensor, a second opinion, or a human reviewer, that confidence score carries operational weight, and getting it wrong costs more than a few misclassified examples. Post-hoc calibration methods exist to correct this without retraining the base network: they take the trained model as fixed and fit a transformation between its raw logits and calibrated probabilities using a held-out validation set.
How Platt Scaling Transforms Logits
Platt scaling was built for a narrower problem than the one neural networks present, and its mechanics reflect that origin. The method takes a classifier's logit, a single real-valued score z, and passes it through a two-parameter logistic function: the calibrated probability is the sigmoid of (a times z, plus b). Both a and b are fit by minimizing negative log-likelihood against a validation set, so the correction is learned rather than fixed in advance. The parameter a rescales the logit the way a single temperature would, but the addition of b lets the method shift the whole curve, correcting for a global bias in the classifier's outputs that a pure rescaling operation has no way to touch. That extra degree of freedom is also where the method's assumptions start to show. Platt scaling was first proposed for support vector machines, and its formulation is inherently binary: one logit maps to one calibrated probability, full stop on the geometry. Nothing in the two-parameter sigmoid was built with a multiclass softmax output in mind, so applying it to a classifier with more than two classes means layering one-vs-rest or one-vs-one decomposition on top of the core method rather than extending the method itself. The fit also assumes that whatever miscalibration exists in the classifier can be captured by a linear transformation of the logit before the sigmoid is applied, an assumption that holds only as well as the real miscalibration pattern happens to match a straight line. And because the method is estimating two parameters instead of one, it needs a validation set large enough to pin down both the slope and the intercept independently; a thin calibration set leaves the shift parameter poorly constrained, which turns the extra flexibility Platt scaling offers into a source of noise rather than a source of accuracy.
How Temperature Scaling Transforms Logits
Temperature scaling asks for far less from the data, and that restraint is the entire point of the method. Rather than operating on a single scalar logit, it takes the full C-dimensional logit vector produced by a multiclass network, divides every entry by one scalar parameter T, and passes the result through the standard softmax function to produce the calibrated probability distribution. T is fit the same way Platt scaling's parameters are, by minimizing negative log-likelihood on a validation set, but there is only one number to estimate instead of two. Guo and colleagues introduced the method explicitly as a single-parameter variant of Platt scaling, a naming choice that makes the family relationship and the parameter reduction visible at once. Because the correction is a single scalar, the optimization problem is one-dimensional and convex: using a conjugate gradient solver, the optimal T can be found in around ten iterations, a calculation that runs in a fraction of a second on standard hardware. Setting T above 1 softens the distribution and reduces overconfidence; setting it below 1 sharpens the distribution and increases confidence. This single-scalar design assumes that a neural network's miscalibration is low-dimensional and uniform across classes rather than concentrated in a few of them. That assumption is not just a convenient simplification. Guo and colleagues tested vector scaling, a per-class generalization of Platt scaling with one scaling parameter per class, and found that the learned per-class values came out nearly constant across classes once fitting was complete. The data itself was close to scalar in its miscalibration pattern. A single T was capturing almost everything the richer, per-class version could capture.
The one structural difference that determines everything else: where in the pipeline each method operates
Comparing parameter counts misses the real distinction between these two methods. What separates them is where in the computation each one intervenes: temperature scaling rescales the logit vector before softmax is applied, while Platt scaling, in its standard form, operates on outputs after whatever softmax-like transformation the classifier already uses. That single placement decision is the structural fact from which every other difference follows. Because temperature scaling divides the entire logit vector by one positive scalar before the softmax normalizes it, the operation is coherent across the whole probability simplex: every class probability is softened or sharpened simultaneously, and the normalization that guarantees the outputs sum to one happens automatically as part of the softmax step. Platt scaling, working after that normalization in its binary form, has no equivalent built-in consistency when extended to many classes. Applying it to a multiclass problem means decomposing the task into one-vs-rest or one-vs-one binary comparisons, and the resulting set of per-class probabilities is not guaranteed to sum to one, so a renormalization step has to be grafted on, introducing a seam where the method's math stops being self-contained.
The second consequence concerns accuracy itself. Dividing every logit by the same positive scalar T cannot change which logit was largest to begin with, so argmax(z/T) equals argmax(z) for any T greater than zero: temperature scaling can never alter which class the model predicts, only how confident it sounds about that prediction. Platt scaling's shift parameter b has no such guarantee, because it operates in probability space, where its effect on the argmax depends on the local shape of the sigmoid at whatever point the classifier happens to be operating. That makes Platt scaling, in principle, capable of flipping the model's top prediction, a property that makes it a fundamentally different kind of calibrator from temperature scaling rather than simply a more flexible version of the same idea.
Why the Simpler Method Wins
Guo and colleagues' 2017 benchmark produced a result that looks counterintuitive on first reading: temperature scaling, the least expressive of the post-hoc methods they tested, outperformed both vector scaling and matrix scaling, the two richer generalizations built to subsume it. Vector scaling assigns a separate scaling parameter to each class; matrix scaling applies a full linear transformation across the entire logit vector. Both contain the scalar transformation of temperature scaling as a special case, so in principle both should do at least as well. The fact that they did worse is not a quirk of one benchmark. The per-class scaling parameters vector scaling learned came out nearly constant across classes: the richer model, given the freedom to find class-specific structure, found almost none and used its extra capacity to fit noise in the calibration set instead of signal. Kull and colleagues reproduced the same pattern in their 2019 Dirichlet calibration work at NeurIPS, where calibration methods carrying a large number of parameters overfit small calibration sets, which explains why the more expressive variants underperformed consistently rather than occasionally. The same logic extends directly to Platt scaling's shift parameter b: in a multiclass setting decomposed into many binary problems, that extra degree of freedom per class is rarely justified by the amount of calibration data available, so it tends to add variance rather than capture real structure. None of this establishes that one scalar is always enough. It establishes that for the typical neural network, trained with standard cross-entropy loss, the miscalibration pattern is close to uniform across classes rather than concentrated in a subset of them, and under that condition a single scalar correction captures nearly all the available signal while a richer one mostly captures noise.
Where temperature scaling's single-scalar assumption breaks down
The uniform-miscalibration assumption that makes temperature scaling so effective on standard benchmarks is also precisely what limits it, and three distinct failure modes show where the single scalar runs out of room. The strongest of the three is theoretical. Chidambaram and Ge, at ICLR 2024, prove that temperature scaling fails in a specific and unavoidable way when class supports overlap: models trained by empirical risk minimization become extremely confident in small neighborhoods around their training points, and in regions where classes genuinely overlap, the model simply picks a side and commits to it with high confidence. A single T can only squash the entire output distribution toward uniformity or sharpen it uniformly; it has no mechanism for correcting overconfidence in the overlap regions without also destroying legitimate, well-earned confidence everywhere else. Chidambaram and Ge show this degradation worsens with the degree of class overlap and that performance asymptotically becomes no better than random as the number of classes grows large, a limit that no amount of retuning T can patch, because the problem is structural rather than a matter of finding the right value.
The most practically actionable failure mode concerns how individual examples respond to a global correction rather than how the whole distribution behaves in aggregate. Joy and colleagues, at AAAI 2023, document that vanilla temperature scaling applies the same fixed rescaling to every input regardless of whether that particular input was classified correctly or incorrectly. Average calibration across a dataset improves, but the correction is blind to the individual case: easy examples that were already well-calibrated get pushed off target by the same correction that hard examples needed, while the hardest examples remain under-corrected. The SMART framework, from Guo and colleagues at ICML 2026, formalizes this as a logit-margin problem: examples with a wide margin between the top two logits drift toward under-confidence after a global temperature adjustment, while low-margin examples stay overconfident, and no single scalar T can resolve both directions of error at once.
The narrowest of the three failure modes concerns evaluation against soft labels rather than hard ones. A 2025 preprint, "Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions," finds that the gap between calibrated confidence and true soft-label correctness grows with model scale in vision settings, so a larger model fitted with temperature scaling on hard labels is not automatically better calibrated when evaluated against the fuller distribution of human label judgments. The scalar fit to hard-label cross-entropy does not transfer cleanly to that softer evaluation criterion, a limitation that matters most in settings where ground truth itself is a distribution rather than a single class.
Platt Scaling's Failure Modes on Neural Classifiers
Platt scaling's extra parameter and its binary-native design do not make it a fallback for everywhere temperature scaling struggles. These weaknesses appear specifically on neural classifiers with many classes and finite calibration data. The most direct of these is calibration set size. Fitting two parameters, a and b, requires enough validation data to constrain both the slope and the intercept of the sigmoid independently, and when the calibration set is small, the fit becomes unreliable in exactly the way the structural argument in the previous sections predicts: the model has more freedom than the data can support. That weakness sits on a spectrum rather than standing alone. Platt scaling is considerably more data-efficient than fully non-parametric alternatives like isotonic regression, so its reliability zone is narrow on both sides: it demands more calibration data than temperature scaling to fit safely, and less than isotonic regression would require, which leaves it as the right tool only across a specific, bounded middle range of calibration set sizes.
The second weakness is structural rather than a matter of data volume. In multiclass settings, Platt scaling's binary-decomposition approach, whether one-vs-rest or one-vs-one, produces per-class probability estimates that are not guaranteed to sum to one, and the renormalization needed to fix that introduces distortions that compound as the number of classes grows. Vector scaling, the direct multiclass generalization of Platt scaling, underperforms temperature scaling precisely because its additional per-class parameters overfit the calibration set rather than capturing structure that was genuinely there to find. The shift parameter that gives Platt scaling its flexibility in the binary case becomes a liability once it has to be estimated separately for each class without enough data to support the estimate.
The third weakness returns to the accuracy-preservation property established earlier. Because Platt scaling's shift parameter operates in probability space rather than on the pre-softmax logit, its effect on the model's argmax depends on the local shape of the sigmoid, and the method can in principle change which class a classifier predicts as its top choice. In a production system where downstream logic depends on the ranking of predicted classes, that property makes Platt scaling capable of a kind of risk that temperature scaling, which by construction cannot touch the ranking, does not carry. None of this rules Platt scaling out. It has a genuine place in binary classification problems and in settings with calibration sets large enough to support its extra parameter, where its ability to correct a global bias that pure scaling cannot reach is a real advantage rather than a liability.
Sources
- Sample Margin-Aware Recalibration of Temperature Scaling
- On Calibration of Modern Neural Networks
- On Calibration of Modern Neural Networks
- Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities
- Calibration of Neural Networks
- Calibration in Deep Learning: A Survey of the State-of-the-Art
- Classifier Calibration: A survey on how to assess and improve predicted class probabilities
- Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions


