Archival Intelligence

Archival Intelligence is part of the Schmidt Sciences Humanities and AI Virtual Institute. Jon Chun serves as Co-PI on the project. The team is building open computational infrastructure for rescuing endangered cultural archives in New Orleans, including Creole and Cajun multilingual newspapers and early jazz materials. Jon contributes computational methods, AI workflows, and evaluation design.

The system has to move from archival materials to usable AI-assisted retrieval and interpretation without stripping away the languages, histories, and community knowledge that make those materials meaningful. That is not a generic retrieval-augmented-generation or chatbot problem. Archival documents can be noisy, incomplete, multilingual, and historically unstable; a plausible retrieval can still be culturally wrong.

  • Retrieval and representation. Organize difficult archival material so people can find and work with it without treating historically specific terms and categories as interchangeable.
  • Provenance and source integrity. Preserve the evidence behind a result and keep its origin and historical context visible rather than allowing a fluent answer to replace the record.
  • Domain-expert workflows. Let archivists, historians, linguists, and community experts shape what gets represented, what counts as relevant, and how ambiguity is handled.
  • Evaluation in the build. Test whether the system retrieves something plausible alongside what it omits, distorts, mistranslates, or removes from context.

The project connects technical infrastructure to a practical rescue method: free, open AI tools that can work with smartphone photography of endangered cultural materials. See the Archival Intelligence collaboration →


How Jon builds

A system is not successful simply because it functions technically. It has to survive contact with users, incentives, ambiguity, misuse, institutional constraints, and the domain it claims to serve. Jon therefore treats evaluation as part of system design from the beginning, alongside implementation and interaction.

Real users and domain experts help expose requirements that a benchmark or controlled demo can miss. Adversarial and failure conditions reveal where stated capabilities break. Uncertainty has to remain visible when the evidence cannot support a single confident answer. The goal is not to remove interpretation from the system, but to design for it honestly.

Domain expertise is part of the architecture

In Archival Intelligence, humanistic expertise is not an external review layer applied after engineering. It shapes what the system represents, what evidence it must preserve, what counts as a valid answer, how retrieval is evaluated, how provenance and uncertainty are surfaced, and which failure modes matter. Domain knowledge therefore changes the architecture and workflow as well as the interpretation of outputs.


Building instruments for comparison

Jon’s building work also includes original computational methods. SentimentArcs combines ensemble sentiment analysis with time-series comparison; LTTB normalization and dynamic time warping make unequal-length narrative trajectories comparable.

Joint work on explainable narrative analysis connects model output to sentence-level evidence. His independently authored MultiSentimentArcs extends trajectory comparison across dialogue and images in film.

These methods connect production systems to current evaluation research: build an instrument, expose its assumptions, test it against difficult data, and keep the evidence available for inspection.


SafeWeb

Jon co-founded SafeWeb in 2000 with Stephen Hsu and James Hormuzdiar and later served as CEO. SafeWeb ran the world's largest privacy, anonymity, and anti-censorship web proxy at the relevant historical point and built Triangle Boy, a proxy system used to reach censored sites. It received the first security investment by In-Q-Tel, a nonprofit strategic investment firm affiliated with the CIA, and was acquired by Symantec in 2003 for $26 million. After the acquisition Jon served as director of development for Symantec's clientless VPN appliance line. SafeWeb on Wikipedia →

SafeWeb operated where privacy promises met adversarial use, censorship, browser behavior, and real deployment. Building privacy and security systems in that environment taught Jon to test technical claims against actual users, incentives, misuse, and failure modes. Those habits now shape how he approaches AI systems and evaluation.

What the sources show

The historical record includes both sides of the SafeWeb story: independent coverage of anti-censorship use, In-Q-Tel support, and the Symantec acquisition, alongside documented criticism of the consumer anonymizer. In 2002, Wired, Computerworld, and a USENIX Security paper described JavaScript- and cookie-based vulnerabilities that could deanonymize users, while later trade coverage tracked SafeWeb's shift toward enterprise SSL VPN products and Symantec's acquisition.

  • The Wall Street Journal covered In-Q-Tel's SafeWeb licensing and investment agreement.
  • The New York Times reported on SafeWeb proxy technology in the context of Chinese internet censorship and Voice of America access.
  • RAND analyzed SafeWeb and Triangle Boy in its report on Chinese dissident internet use.
  • Wired and USENIX Security documented anonymizer vulnerabilities; Wired's follow-up covered SafeWeb's response.
  • Le Monde and CRN add international and trade-press context to the In-Q-Tel and acquisition record.
  • Network World and The Register covered Symantec's purchase of SafeWeb as an SSL VPN maker.

SafeWeb to NIST

The through-line is not that web anonymizers, cultural archives, and frontier AI are the same problem. It is a builder's habit of testing technical claims against users, adversarial use, failure modes, and institutional consequences.

  • 2000. SafeWeb begins with privacy, anti-censorship access, and web-anonymization infrastructure.
  • 2001-2002. The public record includes both In-Q-Tel and anti-censorship coverage and independent vulnerability research.
  • 2003. Symantec acquires SafeWeb and folds the work into clientless SSL VPN products.
  • 2024-present. As Co-PI representing the Modern Language Association in the NIST CAISI consortium, Jon works on LLM evaluation, red-teaming, and ethical auditing: again asking how powerful systems behave under pressure.
  • 2025-present. With Archival Intelligence, he brings evaluation-aware engineering to open computational infrastructure for endangered cultural archives.

Patents: early VPN appliances

The SafeWeb work produced two US patents on what were among the first SSL / clientless VPN appliances. What was novel: secure remote access delivered through an ordinary web browser, with no pre-installed client software, by manipulating and re-encrypting traffic at a gateway. It is the architecture the enterprise "clientless VPN" market later standardized on.

  • US 7,730,528. “Intelligent secure data manipulation apparatus and method” (filed 2001; issued 2010). Originally assigned to SafeWeb, Inc., then to Symantec following the 2003 acquisition.
  • US 8,065,520. “Method and apparatus for encrypted communications to a secure server” (continuation of a May 2000 application; issued 2011). Assigned to Symantec.

Across different technical eras, Jon's building work has focused on systems that have to operate in messy, high-context environments rather than controlled demos.

Human-Centered AI extends that practice into capability building. Jon designed technical pathways through programming, data analysis, model evaluation, and research practice so students from different disciplines could build and test computational systems themselves. The nonprofit Human-Centered AI Lab, which he co-founded, supports collaboration among distributed researchers and domain experts.


What is Jon Chun building now?

Jon Chun is a Co-PI on Archival Intelligence, part of the Schmidt Sciences Humanities and AI Virtual Institute (HAVI). The project develops open computational infrastructure for endangered Creole and Cajun multilingual newspapers and early jazz materials in New Orleans.

How does Jon Chun incorporate domain expertise into AI system design?

Domain experts help define what the system represents, what counts as relevant evidence or a valid answer, how provenance and uncertainty appear, and which omissions, distortions, or losses of context count as failures.

What did Jon Chun build at SafeWeb?

Jon Chun co-founded SafeWeb in 2000 and later served as CEO. The company built web-anonymization, anti-censorship proxy, and enterprise SSL/clientless VPN systems. Symantec acquired SafeWeb in 2003 for $26 million, and the work produced US patents 7,730,528 and 8,065,520.

How does Jon Chun's security background influence his AI work?

Building privacy and security systems in adversarial environments taught Jon to test technical claims against actual users, incentives, misuse, and failure modes. Those habits now shape how he builds and evaluates AI systems.