QlarifyLabs
← Evals
ExploratoryEmerging

Chain-of-thought faithfulness probing

Test whether a model's stated reasoning actually determines its answer, or is a post-hoc rationalization.

Published August 22, 2026

How it works

Models often produce plausible, fluent step-by-step reasoning that does not actually reflect the computation behind the final answer — the explanation can be a post-hoc rationalization rather than a faithful account. Probing this takes a few forms: injecting a biasing hint into the prompt (subtly suggesting the 'expected' answer) and checking whether the stated reasoning ever acknowledges the hint even when the answer clearly shifted toward it; truncating or corrupting steps in the reasoning trace and observing whether the final answer changes the way it should if the trace were actually load-bearing; and reordering steps to see whether an answer that's supposedly derived sequentially is in fact insensitive to the derivation.

When to use it

High-stakes or regulated uses that rely on the model's stated reasoning as an explanation — audit trails, decision justifications, anything a human downstream is expected to trust the 'why' of, not just the 'what'; interpretability research into how models actually arrive at answers.

Limitations

Faithfulness itself is hard to define and measure precisely — there's no ground truth for 'the real computation' to compare the stated trace against, only indirect signals like sensitivity to hints and perturbations. Methods in this area are still maturing and results can be sensitive to exactly how the perturbation is designed.

Cite this

Qlarify Labs. (2026). Chain-of-thought faithfulness probing. Retrieved from https://labs.qlarify.fi/evals/cot-faithfulness-probing