Typed Evaluation Models and the Collapse Between Conflict and Ignorance: A Case Study on Jev
Palabras clave:
typed evaluation models; System One models; LLM-as-judge; annotated paraconsistent logic; representational collapse; typed abstention; Jev.Resumen
Chat-based LLMs return free text that must be interpreted; a category of tools that emerged in September 2026, typed evaluation models ("System One models"), instead return a typed object directly usable by software. Jev (TypeSafe AI, reported by the vendor as launched on September 15, 2026) is the first commercial example of this category. This work situates it within a five-category landscape of AI model interaction and tests, with a minimal experiment, a prediction derived from annotated decision theory: that an output collapsed to a single calibrated probability cannot distinguish genuine conflict from genuine ignorance. For Jev's simplest question type (boolean/Noul), the experiment is consistent with that prediction: conflict cases (probability 0.50-0.57) and ignorance cases (0.46-0.48) fall within the same narrow band. But a second experiment, using the same model's Choice type with a schema that explicitly names "conflicting evidence" and "insufficient evidence" as options, separates both cases with probability 1.0 in all four completed cases. A third experiment, with Choice restricted to the two original options (no escape categories), reproduces neither the collapse nor the clean separation: an unanticipated directional bias appears, possibly tied to lexical cues in the input text. The central finding is therefore not that Jev collapses per se, but that the same model preserves, destroys, or unpredictably distorts the distinction depending on the question type and declared schema — an interface design risk more complex than a single follow-up experiment could anticipate.
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Neutrosophic Computing and Machine Learning

Esta obra está bajo una licencia internacional Creative Commons Atribución 4.0.
