Player One: Winner! ECE doctoral student’s theoretical game of AI attacks places at LLM conference

8/25/2026 Cassandra Smith

Xinbo Wu, a doctoral student in electrical and computer engineering at The Grainger College of Engineering at the University of Illinois Urbana-Champaign, had a paper accepted to ICML, one of the top artificial intelligence conferences, for research modeling AI “jailbreak” attacks as a game between an attacker and defender. The framework uncovers a structural advantage for attackers and derives a provably optimal defense work Wu sees as a step toward using theory to explain why AI systems fail. 

Written by Cassandra Smith

Every game has at least two players trying to outmaneuver each other. Xinbo Wu, a doctoral student within the electrical and computer engineering department in The Grainger College of Engineering at the University of Illinois Urbana-Champaign, has studied the moves attackers use to get large language models to break their own rules and the countermoves defenders can make. Now, Wu has won a different game: his paper on the subject was accepted to the International Conference on Machine Learning, one of the most competitive conferences in the artificial intelligence field. 

“I’m very proud of this acceptance,” Wu said. “This project involved both mathematical theory and extensive empirical work.” He said the conference validated the direction in which he wanted to take his research. “I think it’s a good platform to present my work, which demonstrates a way to use theoretical work to study empirical problems and show its powers to tackle a very challenging problem in AI safety.” 

Wu said this paper, “A Game-Theoretic Analysis of Attacks on Large Language Models via Compositional Skills,” was created after seeing attackers try to bypass safety mechanisms within large language models (LLMs). “Jailbreaks are techniques that manipulate an LLM to generate restricted, harmful, or unintended content and can  appear very different,” Wu said. “Some use role playing, translation, poetry, metaphors, hypothetical scenarios or complicated narratives. But underneath, they often follow the same principle: that LLMs are already good at disguising the users or their harmful intents.” That led him to ask whether these attacks could be understood as a part of one general mathematical problem, rather than a collection of unrelated prompting tricks. 

To solve this issue, Wu took this battle into a game-theoretic setting. “I modeled this situation as a game between an attacker and a defender,” said Wu. “In reality, the attack and defender are adapting to each other. The target searches for ways to effectively conceal its malicious intent while the defender decides where to concentrate its limited resources to defend against these attack patterns, such as fine tuning an underlying model.” 

There are many ways this project could have been framed, but Wu said the game theory worked best. “I try to use the lens of game theory, which is commonly used as a tool to study problems in economics, to investigate very different AI safety problems.” This way of thinking was influenced by Wu’s advisor, ECE adjunct associate professor Lav Varshney. “It was very enjoyable, working with Lav. He gave me a lot of freedom in choosing research directions,” said Wu. 

The result of this game is a theoretical “best-response” attack strategy. “Because this project involved both mathematical theory and extensive empirical work, it was challenging to make the abstraction general enough to capture many existing attacks while still being precise enough to describe the equilibria and optimal defense,” Wu said. 

Having his paper accepted into ICML was a huge moment for Wu. It was not just validation for one paper, but of a broader belief that more theoretical structures have real value in AI safety research. He wants to use theory to explain why failures happen, then turn that into testable, practical defenses. 

So, what’s next for this research? “Make the framework more realistic and more deployable,” Wu said. “I would want to study how the defense generalizes across related skills and how the attacker effectively changed as large number of skills are composed and how to account for practical trade-off such as false positives, model utilities and the computational cost.” 

For Wu, the ICML acceptance is a milestone in a career he's just beginning. It affirms not only the strength of this particular paper, but the value of the broader approach he wants to keep pursuing: using theory to explain why AI systems fail and turning those insights into defenses that hold up in practice. As large language models grow more capable and more widely used, that kind of grounded, rigorous thinking is likely to matter more, not less.  

For Wu, this paper isn't a finish line. It's a sign that he's headed in the right direction. 

The paper was co-authored with researchers from the ECE department including Assistant Professor Huan Zhang, Teaching Assistant Professor Abhishek K. Umrawal and Adjunct Associate Professor Lav Varshney.


Grainger Engineering Affiliations

Abhishek Umrawal is an Illinois Grainger Engineering teaching assistant professor of Electrical and Computer Engineering and is affiliated with the Coordinated Science Laboratory and the National Center for Supercomputing Applications.

Lav Varshney is an Illinois Grainger Engineering adjunct associate professor of Electrical and Computer Engineering and is affiliated with the Siebel School of Computing and Data Science.

Huan Zhang is an Illinois Grainger Engineering assistant professor of Electrical and Computer Engineering and is affiliated with the Coordinated Science Laboratory and the Siebel School of Computing and Data Science.


Share this story

This story was published August 25, 2026.