Internal Verdicts Track Evidence: A Template-Controlled Neutrosophic Reading of Epistemic States in Large Language Models via the Jacobian Lens
Keywords:
neutrosophic logic, single-valued neutrosophic sets, large language models, mechanistic interpretability, Jacobian lens, epistemic auditing, template controlAbstract
For twenty-five years, neutrosophic logic [1, 2, 3] has assigned independent degrees of truth (T), indeterminacy (I), and falsity (F) to model epistemic states that classical and fuzzy semantics cannot express. All previous applications, however, measured these components on outputs: answers, judgments, expert evaluations. This paper reports, to our knowledge, the first neutrosophic reading of components inside the internal representations of a large language model. Using the recently released Jacobian lens [9] on Qwen3.5-4B, we project intermediate-layer readouts onto lexicons of support, refutation, and hedging, obtaining layer-wise (T, I, F) profiles under three epistemic conditions (conflicting evidence, first-person false belief, factual control; n = 20 each). A first battery yielded three striking signatures — a surge of refutation mass under conflict, sustained indeterminacy under false belief, and an apparent "verdict collapse" before the output layer. A second, template-controlled battery showed that all three, as initially stated, were largely artifacts of prompt grammar: matched templates with agreeing evidence inflate the same lexicon masses, true beliefs elicit the same hedging, and the model verbalizes its verdict when allowed to generate (17/20 items). What survives the controls is stronger than what died: the neutrosophic verdict balance B = log10(F) − log10(T) tracks the polarity of the evidence under identical templates (separation of about 3 orders of magnitude, Cohen's d = 2.3–2.5, p < 10-6 in 9 of 9 layers; correct item-level classification 18/20 and 17/20), and residually discriminates false from true first-person beliefs (p < 2 × 10-4 at every layer). We argue that the independence of the neutrosophic components — the axiom that distinguishes neutrosophy from fuzzy and classical frameworks — is precisely what makes this measurement and its self-correction possible, and we distill the two-battery design into a reusable protocol for neutrosophic auditing of language model internals.
DOI 10.5281/zenodo.22249786
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Neutrosophic Sets and Systems

This work is licensed under a Creative Commons Attribution 4.0 International License.

