I lead a research group on AI safety and societal impacts at HPI.
I am also an Associate Member at Nuffield College, Oxford.

My current research focuses on risks and benefits that emerge as we integrate highly capable AI systems into the social and political processes our societies depend on (e.g. communication, coordination, competition).

Before coming to Berlin, I was a Lecturer at the Oxford Internet Institute and a Research Scientist at the UK AI Security Institute.

News

  • Sept 2026: I moved to Berlin to start a research group on AI safety and societal impacts at HPI.
  • Jun 2026: I joined UK AISI as a Research Scientist while continuing at Oxford.
  • Dec 2025: I joined the Oxford Internet Institute as Departmental Lecturer.
  • Jul 2025: Our HateDay project won Outstanding Paper at ACL 2025 🏆
  • Feb 2025: I won a UK AISI Grant to study the distortive effects of AI writing assistance.
  • Dec 2024: Our PRISM alignment dataset won Best Paper (D&B) at NeurIPS 2024 🏆
  • Dec 2024: I was elected Associate Member of Nuffield College at the University of Oxford.
  • Nov 2024: I joined an EU JRC expert panel to consult on the implementation of the EU AI Act.
  • Aug 2024: Our work on LLM values and opinions won Outstanding Paper at ACL 2024 🏆
  • Jun 2024: I am at NAACL 2024 in Mexico City to present XSTest and organise WOAH.
  • Jan 2024: I published SafetyPrompts.com, a catalogue of open datasets for LLM safety.
  • Jul 2023: The Sexism Detection Task we organised won Best Paper at SemEval 2023 🏆
  • Jun 2023: I joined Dirk Hovy's MilaNLP Lab as a postdoctoral researcher.
  • May 2023: I defended my PhD thesis in Oxford, assessed by Scott Hale and Maarten Sap.
  • May 2023: The HateCheck project that I led won the Stanford AI Audit Challenge 🏆
  • Mar 2023: My work on OpenAI's red team for GPT-4 was covered by various media outlets.
  • Mar 2023: The AI start-up Rewire that I co-founded in 2021 was acquired by ActiveFence.
  • ...

Group

In September 2026, I moved to Berlin to start a new research group on AI safety and societal impacts at HPI. I am now recruiting fully-funded PhD students, postdocs, and research interns to join the group. If you are interested, please fill out this form.

Research

AI systems are growing more and more capable. They are also becoming more widely adopted across society, and more integrated in our personal and professional lives. My group will take these trends seriously and engage with the risks and benefits they create.

Concretely, we will focus on risks and benefits that emerge as we integrate highly capable AI systems into the social and political processes our societies depend on (e.g. communication, coordination, competition). We will measure impacts through model evals, computational experiments, and human studies. When possible, we will develop technical mitigations.

For my work in this area, I won Outstanding Paper at ACL and Best Paper at NeurIPS D&B. For a complete record of my publications, please visit my Google Scholar profile.

Press

I enjoy talking about my work, and I have been fortunate to have it featured across many different media outlets. Here are a few examples:

If you want to chat, please get in touch via email or X.