NeutroStats-AI: A Neutrosophic Statistical Framework forEvaluating Uncertainty in Large Language Model Outputs
Palabras clave:
Neutrosophic Statistics; Large Language Models; Epistemic Uncertainty; Python; AI Auditing; Epistemic Calibration Score; Confidence Intervals.Resumen
The deployment of Large Language Models (LLMs) in high-stakes domains demands statistical frameworks capable of quantifying epistemic uncertainty beyond binary accuracy metrics. This paper introduces NeutroStats-AI, a Python-based framework implementing neutrosophic statistical inference for LLM evaluation. The framework provides neutrosophic confidence intervals, hypothesis tests, and a novel Epistemic Calibration Score (ECS) that penalizes indeterminacy suppression. Validated against six state-of-the-art LLMs across four benchmarks, NeutroStats-AI demonstrates that classical evaluation systematically underestimates model uncertainty by 34-67% relative to neutrosophic evaluation.
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Neutrosophic Computing and Machine Learning

Esta obra está bajo una licencia internacional Creative Commons Atribución 4.0.
