Hi!:) I'm an undergrad studying Math and CS at Brown. Broadly, I hope to build the formal and empirical infrastructure for
trustworthy machine learning along two complementary dimensions:
-
AI4Math and automated reasoning: How can we design models that reason robustly and interpretably, grounded in formal theorem proving and program synthesis?
Formal theorem proving [1] [2] and program synthesis [3] [4]
are fully specified, unambiguous, machine-checkable substrates for studying models' reasoning processes.
Concretely, I'm interested in designing model architectures that leverage various learning paradigms (neurosymbolic programming in particular) to enable more robust and interpretable reasoning, in formal theorem proving and beyond.
-
Principled evaluation with formal guarantees: How do we substantiate trustworthiness claims about ML systems through systematic evaluation and formal guarantees?
I'm interested in designing evaluation and auditing approaches that attach formal or statistical evidence to model properties,
from specification conformance and robustness [4] [1] to value alignment [5] and supply-chain provenance.
These approaches should generalize across systems, development stages, and deployment contexts.
Below are selected publications;
more are available in my CV and Google Scholar
(* marks equal contribution).
-
Jiayi Wu, Robert Joseph George, Anima Anandkumar.
ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers.
ICML 2026 AI for Math Workshop.
[Paper]
[Code]
[Data]
[Project Page]
-
Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, Michiel A. Bakker.
Benchmarking Overton Pluralism in LLMs.
ICLR 2026.
[Paper]
[Code]
[Data]
[Project Page]
-
Ilija Ivanov, Jiayi Wu, Gavin Zhao, Zekai Li, Stephen Bach, Robert Lewis.
Mining Machine-Generated Lean Proofs: Advice for Users, Developers, and Testers.
AITP 2026.
-
Chance Jiajie Li*, Jiayi Wu*, Zhenze Mo, Ao Qu, Yuhan Tang, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Jinhua Zhao, Paul Pu Liang, Luis Alberto Alonso Pastor, Kent Larson.
Simulating Society Requires Simulating Thought.
NeurIPS 2025 Position Paper Track.
[Paper]
Outside of research, I work with communities that connect technology with governance and public-interest perspectives:
I’m co-directing the Technology & Science Policy (TASP) Summit at Center for Technological Responsibility (CNTR)
and the AI Auditing in Practice working group at Data Science Initiative (DSI);
previously, I was co-president of Brown’s AI Robotics Ethics Society (AIRES)
and co-director of the AI Governance Panel at Brown China Summit.
Misc: I enjoy cold brew, archery,
roaming through the neighborhoods and community spaces I'm involved in 🏘️🌳,
and the Rock's basement stacks
-- send me book recs!