• Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
Sunday, October 11, 2026
newsaiworld
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us
No Result
View All Result
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us
No Result
View All Result
Morning News
No Result
View All Result
Home Machine Learning

Why Temperature 0 Is not Deterministic

Admin by Admin
October 10, 2026
in Machine Learning
0
1791170399775 6xky00.jpg
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter

READ ALSO

Can TypeSafe’s Jev Make AI Brokers Safer With out One other LLM?

AI Brokers Beat PyTorch: Writing Sooner CUDA Kernels


In September 2025, Horace He and colleagues at Pondering Machines Lab ran a easy experiment. They despatched the immediate “Inform me about Richard Feynman” to Qwen3-235B 1,000 instances at temperature 0 and requested for 1,000 tokens every time. Temperature 0 means the mannequin at all times picks its most possible token, so you’d count on 1,000 an identical solutions. They acquired 80 distinctive completions.

What caught my eye is the place the solutions break up. All 1,000 completions have been an identical for the primary 102 tokens. At token 103, 992 of them continued with “Queens, New York” and eight with “New York Metropolis.” Eight runs in a thousand flipped at a single token, after 102 tokens the place none did.

Their submit explains why outputs differ in any respect. This text asks a special query: how usually ought to a token flip, and why does that danger arrive in sudden spikes? I will derive a one-line system, examine it with a simulation, and take a look at it on an actual mannequin. All the pieces runs on a laptop computer, and the code is within the article.

The place the noise comes from

At temperature 0 the mannequin picks the token with the best logit (its uncooked rating). That’s deterministic arithmetic, however your laptop computes logits in floating level, the place addition isn’t associative:

>>> 0.1 + 0.2 + 0.30.6000000000000001>>> 0.1 + (0.2 + 0.3)0.6

At scale it will get worse. I summed 100,000 random float32 numbers 3 ways. Ahead order gave −90.82497, reverse order gave −90.82674, and NumPy’s pairwise sum gave −90.82513 (a float64 reference provides −90.82508). Similar numbers, three solutions.

A neural community is billions of such sums, so the order the {hardware} makes use of issues. A preferred clarification blames GPU threads that end in random order. He et al. present that this is not the primary trigger: a typical LLM ahead go accommodates primarily no atomic provides, and the identical kernel on the identical enter returns the identical bits each time. The true perpetrator is that many kernels aren’t batch-invariant. Their discount order adjustments with the batch measurement, and on a shared server the batch measurement is dependent upon what number of different persons are sending requests at that second. Your reply is dependent upon strangers.

The authors be aware that this is not particular to GPUs, and that with batch-invariant kernels all 1,000 Feynman completions got here out an identical. The repair has a worth: of their serving take a look at (Qwen-3-8B on one GPU), the unoptimized deterministic model took 55 seconds in opposition to 26 for default vLLM, and 42 after an improved consideration kernel.

Solely near-ties matter

Right here is the half that pursuits me as a mathematician. Let z₁ and z₂ be the 2 largest logits, and name their distinction the hole, M=z1−z2≥0M = z₁ − z₂ ≥ 0M=z1​−z2​≥0. Numerical noise adjustments that hole by a small error Δ, and the argmax flips provided that M+Δ<0M + Δ < 0M+Δ<0.

A token with a niche of 8 logits won’t ever flip, nonetheless noisy the kernel. Solely near-ties are in danger. Averaging over all positions, with f the density of gaps, the flip chance is

p = ∫ P(Δ < −m) f(m) dm

If the noise is small, f is roughly fixed, equal to f(0), over the slim vary the place P(Δ<−m)P(Δ < −m) P(Δ<−m) is not negligible. The integral of P(Δ<−m)P(Δ < −m)P(Δ<−m) over m≥0m ≥ 0m≥0 is the anticipated detrimental a part of the error, which is E∣Δ∣/2E|Δ|/2E∣Δ∣/2 for symmetric noise. So

p ≈ f(0) · E|Δ| / 2

The flip chance per token is how crowded the neighbourhood of a tie is, instances how giant the noise is. The primary issue belongs to the mannequin and the textual content. The second belongs to your {hardware}, precision and kernels. That is an elementary small-noise argument, so I make no declare of novelty. A associated formalization, the “background temperature” of Messina and Scotta (TMLR, 2026), treats such perturbations as an efficient temperature.

Schematic. The orange area holds the positions whose hole is smaller than a given error. As a result of the error has a random signal and measurement, averaging provides f(0)·E|Δ|/2.

Does the system maintain?

Earlier than touching an actual mannequin, I checked the system on simulated gaps with a identified density and Gaussian noise of identified measurement σ σσ. The core is 9 traces:

def flip_rate(sampler, sigma, n, rng, chunk=5_000_000):    flips, executed = 0, 0    whereas executed < n:        m = min(chunk, n - executed)        hole = sampler(rng, m)            # top-1 minus top-2 logit hole        err = rng.regular(0.0, sigma, m)  # numerical error on that hole        flips += np.count_nonzero(hole + err < 0)        executed += m    return flips / n

For an exponential and a half-normal hole, the simulation matches the system inside 4% for σ as much as 0.03, and the half-normal stays inside 2% even at σ = 0.3 (Determine 2). It breaks when the noise is as giant as the standard hole: at σ = 1 the exponential case flips 24% of the time in opposition to a predicted 40%, as a result of the density is not flat over the area that issues.

A 3rd case exhibits the system’s key assumption, f(0)>0f(0) > 0f(0)>0. If the hole density vanishes at zero, the flip fee grows like σ2σ²σ2 as an alternative (it matches σ2/4σ²/4σ2/4 inside 8% for σ as much as 0.06), and noise issues far much less. So which regime does an actual mannequin reside in?

Strains: the system. Dots: simulation, 10 million samples per level.

An actual mannequin

I ran GPT-2 on 256 home windows of 128 tokens from WikiText-2 and in contrast a bf16bf16bf16 ahead go with an fp32fp32fp32 one. It is a stand-in for lifelike serving noise: my laptop computer cannot reproduce the batch-size impact of a loaded GPU server, however it will probably measure the logit error a lower-precision kernel produces. I run the transformer physique in bf16bf16bf16and the output projection in fp32fp32fp32, so the ultimate logits aren’t rounded to bf16bf16bf16‘s coarse grid.

import copy, numpy as np, torchfrom datasets import load_datasetfrom transformers import AutoModelForCausalLM, AutoTokenizerm32 = AutoModelForCausalLM.from_pretrained("gpt2").eval().float()m16 = copy.deepcopy(m32).to(torch.bfloat16)W = m32.get_output_embeddings().weight   # fp32 output matrix, each runstok = AutoTokenizer.from_pretrained("gpt2")information = load_dataset("wikitext", "wikitext-2-raw-v1", break up="take a look at")textual content = "nn".be a part of(t for t in information["text"] if t.strip())ids = torch.tensor(tok(textual content)["input_ids"])n = len(ids) // 128ids = ids[: n * 128].reshape(n, 128)decide = torch.randperm(n, generator=torch.Generator().manual_seed(0))[:256]ids = ids[pick]

For each place I preserve the fp32fp32fp32 top-20 logits, the hole, the error on that hole within the bf16bf16bf16 run, and whether or not the argmax flipped. (Plotting code omitted.)

gaps, deltas, flips, top20 = [], [], [], []with torch.no_grad():    for x in ids.break up(8):                 # x: [8, 128] token ids        h32 = m32.base_model(x).last_hidden_state[:, 4:]        h16 = m16.base_model(x).last_hidden_state[:, 4:].float()        ref = (h32 @ W.T).reshape(-1, W.form[0])        alt = (h16 @ W.T).reshape(-1, W.form[0])        v, i = ref.topk(20, dim=1)         # fp32 top-20 logits        a = alt.collect(1, i)               # identical tokens, bf16 run        hole = v[:, 0] - v[:, 1]        gaps.append(hole.numpy())        deltas.append(((a[:, 0] - a[:, 1]) - hole).numpy())        flips.append((alt.argmax(1) != i[:, 0]).numpy())        top20.append(v.numpy())hole, delta, flip, top20 = map(np.concatenate,                              (gaps, deltas, flips, top20))
# density of the hole at 0 (per logit), and the error close to tiesf0 = np.imply(hole < 0.25) / 0.25err = np.abs(delta[gap < 0.5]).imply()print(f"f(0) = {f0:.3f}, error close to ties = {err:.4f}")print(f"flip fee: measured {flip.imply():.4f}, "      f"predicted {f0 * err / 2:.4f}")# inject noise of identified measurement into the actual logitsfor s in np.logspace(-3, 0.5, 8):    noise = np.random.randn(*top20.form) * s    print(f"{s:.4f}", np.imply((top20 + noise).argmax(1) != 0))

The density of the hole at zero is f(0)≈0.83f(0) ≈ 0.83f(0)≈0.83 per logit, and 9%9%9% of positions have a niche under 0.1 logits. The bf16bf16bf16 error on the hole close to ties averages 0.020.020.02 logits, so the system predicts a flip fee of 0.010.01 0.01. I measured 0.010.010.01. Determine 3 (proper) repeats the comparability throughout noise ranges by injecting Gaussian noise of identified measurement into the actual logits.

Left: the distribution of the top-two logit hole in GPT-2 close to zero. Proper: flip fee in opposition to noise degree. Dots are noise injected into the actual logits, the star is bf16 in opposition to fp32.

From tokens to completions

For one completion, survival is a product of per-token survival possibilities, S(t)≈exp(−Σps)S(t) ≈ exp(−Σ pₛ)S(t)≈exp(−Σps​). If each token carried the identical danger ppp, that will be exp(−p⋅t)exp(−p·t)exp(−p⋅t), and half of all completions would have break up by token ln 2/p2 / p2/p. Actual textual content is not like that.

Take the Feynman numbers (that is my arithmetic on their printed counts). Over the primary 102 tokens, 1,000 runs produced zero disagreements. That’s about 102,000 token selections with no flip, so the common per-token flip fee over that stretch is at most about 3×10⁻⁵ (the “rule of three” 95% certain). At token 103 the speed was 8 in 1,000, or 0.8%. That’s greater than 250 instances larger, at a single token. Danger is not unfold evenly. It sits on just a few knife-edge tokens, with lengthy stretches of near-zero hazard in between.

Prompts additionally differ in what number of knife-edge tokens they comprise. To see what that does to survival, I simulated a imply per-token flip fee of 0.005 with fragility various throughout prompts (lognormal, log-sd 1):

p_bar, sigma_ln = 0.005, 1.0t = np.arange(0, 1001)# per-prompt flip chance: lognormal with imply p_barp_i = p_bar * np.exp(rng.regular(0, sigma_ln, 400_000) - 0.5 * sigma_ln**2)uniform = (1 - p_bar) ** t                  # identical p for each immediatecombination = np.array([np.mean((1 - p_i) ** k) for k in t])  # p varies
Simulation, not measured information. Left: fraction of completions nonetheless an identical after t tokens. Proper: the identical curves on a log scale. A mix of exponentials has a heavier tail than any single exponential.

If each immediate had p = 0.005, half of the completions would break up by token 139 and eight% would nonetheless match at token 500. With the identical imply however various fragility, the median strikes to token 210 and 28% nonetheless match at token 500. The common flip fee alone would not inform you how lengthy a completion survives. The unfold of fragility issues too.

The place this breaks

  • My noise is not your server’s noise. bf16 in opposition to fp32 on a laptop computer reproduces the scale of a precision error, not the batch-size mechanism of a loaded GPU cluster. These numbers describe the mannequin, not any hosted API.

  • Two tokens solely. The system ignores the third-best token and assumes the error is symmetric and unbiased of the hole.

  • A flip is not an error. Many flips swap near-equivalent phrasings. The system says how usually outputs differ, not how usually they worsen.

What to do with it

Predict your flake fee. In case your stack has a per-token flip fee p, a completion of L tokens differs between runs with chance about 1−exp(−p⋅L)1 − exp(−p·L)1−exp(−p⋅L). With p = 10⁻³ and 500-token solutions, that’s about 39%. Measure f(0) in your mannequin and the noise in your stack, and you may estimate it earlier than transport an eval.

Do not assert actual equality on temperature-0 outputs. Evaluate with a tolerance, or grade semantically. Should you management inference, batch-invariant kernels take away the impact for a throughput worth. Should you do not, assume two an identical calls can differ.

Takeaways

  1. Temperature 0 is deterministic within the arithmetic, not within the arithmetic.

  2. Solely near-ties matter: the per-token flip chance is about f(0)⋅E∣Δ∣/2f(0) · E|Δ| / 2f(0)⋅E∣Δ∣/2 the crowding of ties instances the noise degree.

  3. Danger is spiky. A number of knife-edge tokens carry most of it, which is why completions keep an identical for a stretch after which break up.

···

Sources and credit

  1. He, Horace and Pondering Machines Lab, “Defeating Nondeterminism in LLM Inference”, Pondering Machines Lab: Connectionism, September 10, 2025. https://thinkingmachines.ai/weblog/defeating-nondeterminism-in-llm-inference/ (supply of the Feynman experiment and its counts, the batch-invariance clarification, and the serving timings).

  2. Messina, Alberto and Scotta, Stefano, “Introducing Background Temperature to Characterise Hidden Randomness in Massive Language Fashions”, Transactions on Machine Studying Analysis, 2026. https://openreview.web/discussion board?id=bz0he4bARF

  3. Radford et al., “Language Fashions are Unsupervised Multitask Learners” (GPT-2), 2019, and Merity et al., “Pointer Sentinel Combination Fashions” (WikiText-2), 2016, each used by Hugging Face Transformers and Datasets.

  4. The flip-probability system, simulations, real-model experiment, figures and the arithmetic on the Feynman counts are the writer’s personal.

Tags: DeterministicisntTemperature

Related Posts

1791231261636 64jqin.webp.webp
Machine Learning

Can TypeSafe’s Jev Make AI Brokers Safer With out One other LLM?

October 9, 2026
Vishnu mohanan pfR18JNEMv8 unsplash scaled.jpg
Machine Learning

AI Brokers Beat PyTorch: Writing Sooner CUDA Kernels

October 8, 2026
1790995354743 w6cg9a.png
Machine Learning

When Do PINNs Beat Classical Numerical Strategies? A 1D vs 5D Experiment

October 7, 2026
MLM Shittu RAG vs Fine Tuning for Domain Adaptation 1024x586.png
Machine Learning

RAG vs. Nice-Tuning for Area Adaptation: When to Use Which

October 6, 2026
Feature image 2 scaled.png
Machine Learning

Pc Imaginative and prescient: SIFT algorithm (Scale Invariant Function Rework)

October 5, 2026
MLM Shittu Local Agentic AI Workflows with Hermes Ollama scaled 1.png
Machine Learning

Native Agentic AI Workflows with Hermes + Ollama

October 5, 2026

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

POPULAR NEWS

Gemini 2.0 Fash Vs Gpt 4o.webp.webp

Gemini 2.0 Flash vs GPT 4o: Which is Higher?

January 19, 2025
Chainlink Link And Cardano Ada Dominate The Crypto Coin Development Chart.jpg

Chainlink’s Run to $20 Beneficial properties Steam Amid LINK Taking the Helm because the High Creating DeFi Challenge ⋆ ZyCrypto

May 17, 2025
Image 100 1024x683.png

Easy methods to Use LLMs for Highly effective Computerized Evaluations

August 13, 2025
Blog.png

XMN is accessible for buying and selling!

October 10, 2025
0 3.png

College endowments be a part of crypto rush, boosting meme cash like Meme Index

February 10, 2025

EDITOR'S PICK

Chatgpt image may 23 2026 05 34 02 pm.jpg

Most AI Brokers Fail in Manufacturing As a result of They’re Constructed Backwards

May 28, 2026
Trump Slams The Judicial System Is The Court Hindering The Crypto Bill.webp.webp

Is Courtroom Hindering Crypto Rise?

May 7, 2025
Be351 Crispr Cas 9 Gene Editing Technology.jpg

The Way forward for Predictive Analytics: Tendencies and Improvements to Watch

October 5, 2024
Bitcoin Profit Taking.jpg

Lengthy-term holders are locking in revenue after Bitcoin’s rally to new ATHs

November 16, 2024

About Us

Welcome to News AI World, your go-to source for the latest in artificial intelligence news and developments. Our mission is to deliver comprehensive and insightful coverage of the rapidly evolving AI landscape, keeping you informed about breakthroughs, trends, and the transformative impact of AI technologies across industries.

Categories

  • Artificial Intelligence
  • ChatGPT
  • Crypto Coins
  • Data Science
  • Machine Learning

Recent Posts

  • Why Temperature 0 Is not Deterministic
  • Construct Your First MCP Server in Python (Stateless Spec Version)
  • How Can AI Brokers Learn Untrusted Sources Safely?
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy

© 2024 Newsaiworld.com. All rights reserved.

No Result
View All Result
  • Home
  • Artificial Intelligence
  • ChatGPT
  • Data Science
  • Machine Learning
  • Crypto Coins
  • Contact Us

© 2024 Newsaiworld.com. All rights reserved.

Are you sure want to unlock this post?
Unlock left : 0
Are you sure want to cancel subscription?