One practice, difficult questions

Jon Chun studies how AI systems reason, behave, and interact with people under ambiguity. His current work centers on frontier AI evaluation, while a longer line of computational research develops methods for making difficult human concepts measurable without stripping away context or interpretation.

Humanistic methods remain central to that practice: construct definition comes before measurement, domain knowledge shapes what counts as evidence, and evaluation must preserve ambiguity where ambiguity is part of the problem.


Reasoning and judgment

How do we test whether AI systems can reason, justify decisions, and handle ambiguous human values instead of producing only plausible answers? Jon studies ethical reasoning, human judgment, explanation, confidence, and decision-making in frontier models. The work treats an answer, the reasons offered for it, and the model's behavior across variations as distinct objects of evaluation.

Jon Chun serves as Co-PI representing the Modern Language Association in the NIST CAISI consortium. His standards-facing work includes LLM evaluation and red-teaming, with ethical auditing as a through-line. The team's results were presented during the opening keynote at the consortium's first plenary at the University of Maryland.

Systems, behavior, and adversarial conditions

How do intelligent systems behave outside controlled settings, especially under adversarial, strategic, or multi-agent conditions? This thread examines agent behavior, robustness, failure modes, and the gap between a system's intended behavior and what it does in use. Work on open-source generative AI asks how capabilities, incentives, and deployment choices redistribute risk; current directions include multi-agent evaluation and behavior under strategic interaction.

Jon's earlier privacy and security engineering informs the research question without substituting biography for evidence. Systems exposed to real users taught him to test claims against misuse, incentives, and practical failure modes. That security-informed habit now shapes adversarial evaluation of increasingly capable AI systems. See the deployment background →


Making interpretation measurable

What would it take to measure the shape of a story reliably enough to argue about it? Jon's computational humanities research has focused on building the instruments that make computational reading of narrative testable. Those methods make interpretive concepts empirically tractable without pretending they are simple. Middle Reading places computation between distant and close reading; SentimentArcs uses ensemble sentiment analysis and time-series comparison to model narrative trajectories; later work adds explainability, cross-cultural comparison, and multimodal coherence.

This work trained the same methodological muscles now needed for frontier AI evaluation: construct definition, ambiguity, context, interpretation, and the empirical testing of difficult human concepts. Jon and Katherine Elkins also evaluated GPT-3 for creative writing in the Writer's Turing Test and studied how AI changes narrative interpretation. Katherine Elkins used the companion AI-LIT workflows in In Search of a Translator (doi:10.3389/fcomp.2024.1444021) to compare an original and its translations across whole-narrative time.

  • Narrative sentiment methodology (2019–2022). Elkins and Chun introduced the first rigorous methodology for narrative sentiment analysis, establishing a reproducible framework for selecting, comparing, validating, and interpreting sentiment models across narrative texts. Earlier studies had applied sentiment analysis to literature; the contribution was the systematic method for deciding which models were appropriate and how robust conclusions were across methods.
  • Middle Reading (2019). Elkins and Chun introduced a methodological position connecting close and distant reading, in which computational patterns and domain knowledge can correct one another.
  • Writer's Turing Test (2020). Elkins and Chun conducted the first writer's Turing test of a large language model, published in the Journal of Cultural Analytics.
  • SentimentArcs (2021). Chun introduced SentimentArcs, a novel ensemble method for comparing narrative sentiment trajectories across texts. The study weighs dozens of models against one another and reports that state-of-the-art transformers can struggle to find narrative arcs.
  • Narrative trajectory comparison (2021). LTTB downsampling normalizes arcs to a common number of points while preserving peaks, valleys, and endpoints. Dynamic time warping then measures distance between whole arcs, making narratives of unequal length comparable and clusterable.
  • Explainable narrative analysis (2023). Chun and Elkins published the first application of explainable AI to narrative analysis, developing sentence-level story trajectories and comparability procedures for unequal-length narratives.
  • Ethics audit (2024). Chun and Elkins conducted the first ethics-based audit of moral reasoning in deployed large language models, comparing judgments, explanations, and confidence across ambiguous social decisions.
  • Comparative AI regulation (2024). Chun, Schroeder de Witt, and Elkins produced the first systematic comparison of AI regulation across the European Union, China, and the United States after passage of the EU AI Act.
  • Multimodal coherence (2024). Chun introduced the first multimodal method for measuring long-form film sentiment-arc coherence across modalities, comparing trajectories recovered from dialogue and images.

Read the full narrative-trajectories and dynamic-time-warping lineage →


Archival Intelligence

Archival Intelligence is part of the Schmidt Sciences Humanities and AI Virtual Institute. The project is building open computational infrastructure for rescuing endangered cultural archives in New Orleans, including Creole and Cajun multilingual newspapers and early jazz materials. The research asks how AI-assisted retrieval and interpretation can support preservation and access while retaining provenance, historical context, language, and community knowledge.

As Co-PI on Archival Intelligence, Jon contributes computational methods, AI workflows, and evaluation design. Cultural archives are a demanding test for AI: documents are noisy and incomplete; language changes across time and communities; categories carry historical assumptions; and a plausible retrieval can still be culturally wrong. Domain expertise and humanistic interpretation therefore become technical requirements for defining relevance, evaluating results, and designing archival systems.

The project extends Jon's longer research program from narrative trajectories to multilingual retrieval and archival infrastructure: construct definition comes before measurement, and evaluation must include what a system omits, distorts, or removes from context. Explore Archival Intelligence and its collaborators →


Systems, rules, and openness

With Christian Schroeder de Witt and Katherine Elkins, Jon developed a comparative framework for AI regulation across the European Union, China, and the United States. The work examines how institutions define risk, enforcement, innovation, and public responsibility differently.

Jon also contributed to the ICML position paper led by Francisco Eiras on the risks and opportunities of open-source generative AI. This governance work supports the evaluation program without displacing its central focus on methods, model behavior, and research infrastructure.


Publications across the program


Methods in use

  • The GPT-3 creative-writing study informed later work on human and machine persuasion in Science Advances and PNAS.
  • The narrative and computational-humanities work has been extended in narrative theory, education, translation, and multimodal cultural analysis.
  • The open-source generative-AI and comparative-regulation research has entered policy and governance discussions across multiple jurisdictions.

See full scholarly reception → · Google Scholar →


Evaluation beyond the benchmark

Current directions include frontier-model reasoning and judgment, agent behavior under strategic interaction, cultural and multilingual evaluation, retrieval and interpretation for archives, and standards that connect empirical results to governance. Across them, Jon asks whether a system succeeds, what the test assumes, whose expertise defines success, and how behavior changes outside controlled settings.


What does Jon Chun research?

Jon Chun studies frontier AI evaluation, human judgment, ethical reasoning, agent behavior, governance, and computational methods for narrative, culture, and archives. His research asks how difficult human phenomena can become technically testable without losing ambiguity, context, interpretation, or domain expertise.

What is his role at NIST CAISI?

Jon Chun serves as Co-PI representing the Modern Language Association in the NIST CAISI consortium, working on LLM evaluation and red-teaming.

Where has the research appeared?

Venues including ICML (a 2024 oral presentation), the Journal of Cultural Analytics, Narrative, and the International Journal of Humanities and Arts Computing.

What is Archival Intelligence?

Archival Intelligence is a project in the Schmidt Sciences Humanities and AI Virtual Institute (HAVI). It develops open computational infrastructure for endangered New Orleans cultural archives while preserving their languages, historical context, and domain expertise. Jon Chun is a Co-PI on the project.