QlarifyLabs
← Evals
ExploratoryEmerging

Chain-of-thought & explanation faithfulness probing

Test whether the reasons a system gives — its step-by-step reasoning, or the explanation it shows a user for a decision — actually determine its answer, or are a plausible story told afterwards.

Published August 22, 2026 · Updated October 3, 2026

How it works

Models often produce plausible, fluent step-by-step reasoning that does not actually reflect the computation behind the final answer — the explanation can be a post-hoc rationalization rather than a faithful account. Probing the reasoning takes a few forms: injecting a biasing hint into the prompt (subtly suggesting the 'expected' answer) and checking whether the stated reasoning ever acknowledges the hint even when the answer clearly shifted toward it; truncating or corrupting steps in the reasoning trace and observing whether the final answer changes the way it should if the trace were actually load-bearing; and reordering steps to see whether an answer that's supposedly derived sequentially is in fact insensitive to the derivation. The same question applies to the explanations an application shows its users: why a claim was flagged, why a candidate was ranked low, why a request was declined, why this product was recommended. Those explanations name factors, and each named factor is a testable claim. Change the input so the cited factor is absent or reversed and check the decision moves the way the explanation says it should; change a factor the explanation never mentions and check the decision stays put. An explanation that cites one reason while the decision tracks another — a protected attribute, the order the options were listed in, the length of the input — is unfaithful in the way that matters to the person reading it. A second check asks whether the explanation is useful for prediction at all: given only the explanation, could a reader correctly guess what the system would do on a nearby case?

When to use it

High-stakes or regulated uses that rely on the model's stated reasoning as an explanation — audit trails, decision justifications, anything a human downstream is expected to trust the 'why' of, not just the 'what'; any feature that shows end users a reason for a decision about them, such as screening, ranking, moderation, credit or claims handling, where a wrong reason misleads the person and can hide a bias the decision itself contains; review workflows where a human approves the system's output on the strength of its explanation; and interpretability research into how models actually arrive at answers.

Limitations

Faithfulness itself is hard to define and measure precisely — there's no ground truth for 'the real computation' to compare the stated trace against, only indirect signals like sensitivity to hints and perturbations. Counterfactual checks on user-facing explanations depend on changing one factor without changing anything else, which is easy for a structured field and hard for anything expressed in free text; a clumsy counterfactual tests the edit, not the explanation. An explanation can pass every check and still leave out a factor that mattered, since the method tests the factors named and a factor nobody thought to vary stays invisible. When several cited factors are each sufficient, removing one need not move the decision, and every explanation omits some minor factor — so results are a matter of degree, comparing the size of each factor's effect with the weight the explanation gives it, not a pass/fail on whether the decision moved. Methods in this area are still maturing and results can be sensitive to exactly how the perturbation is designed.

Cite this

Qlarify Labs. (2026). Chain-of-thought & explanation faithfulness probing. Retrieved from https://labs.qlarify.fi/evals/cot-faithfulness-probing