How AI systems make, explain, and justify judgments under ambiguity—including ethical reasoning, decision-making, and the relationship between an answer and its explanation.
Jon A. Chun
Jon Chun is a researcher and builder working at the frontier of AI evaluation. He combines engineering, humanistic methods, and real-world deployment experience to test how intelligent systems reason, behave, and interact with people.
His work spans frontier-model evaluation, computational research into difficult human questions, the study of human judgment, and systems tested beyond the lab. Experience in privacy, security, and adversarial environments informs how he turns abstract claims about AI into concrete, testable questions.
What he works to understand
How intelligent systems act in real-world, adversarial, and multi-agent conditions, with attention to robustness, agentic behavior, evaluation, and practical failure modes.
How AI systems interact with people, institutions, culture, and context—from human-AI interaction and narrative to institutional use and deployment.
Methods that meet the problem
Humanistic methods are not decorative context. They help define what should be measured and show how ambiguous human concepts can be operationalized without reducing away the problem itself.
Building systems, prototypes, technical infrastructure, and computational methods that make ideas testable.
Turning claims about AI into benchmarks, measurements, experiments, and comparative analysis.
Using interpretation, construct definition, narrative analysis, context, and ambiguity to improve evaluation.
Testing systems against users, institutions, privacy and security constraints, incentives, misuse, and failure modes.
Three connected bodies of work
Jon studies human judgment and ethical reasoning in frontier models and contributes to evaluation and standards work as Co-PI representing the Modern Language Association in the NIST CAISI consortium. His research develops concrete ways to compare what models decide, how firmly they decide it, and how they explain those judgments. Explore the evaluation research →
SentimentArcs, narrative modeling, and explainable-AI research reflect a longer project: turning complex human constructs into computationally testable forms while preserving their ambiguity and context. The work joins scalable analysis with interpretable measures. See SentimentArcs →
Building privacy and security systems at SafeWeb and Symantec taught Jon to test technological claims against actual users, incentives, misuse, and failure modes. That experience now informs his approach to AI evaluation. Read the building story →
A result that changes the comparison
Across healthcare, law, and finance scenarios, instruction-tuned language models were 110–300× more resistant to narrative manipulation than people. The finding shifts attention from whether models can be influenced in principle to how model and human vulnerabilities differ in practice. Read the research →
Where the work travels
Research on narrative AI, model openness, ethical reasoning, and human-centered AI has been taken up across computer science, digital humanities, governance, and education.
Selected scholarly reception and public discussionScholarly reception → · Grants and recognition → · Press coverage →
Choose a path
AI evaluation, human judgment, narrative modeling, publications, and current findings.
SafeWeb, privacy and security infrastructure, patents, and deployed AI systems.
NIST CAISI, Schmidt Sciences HAVI, academic partners, institutions, and cross-sector work.
The longer biography, background, teaching, speaking, and full professional record.