Samuel Šimko

ELLIS PhD Student in AI Safety

samuel_simko_photo.png

I am an ELLIS PhD student working on AI safety, advised by Prof. Bernhard Schölkopf and co-advised by Prof. Zhijing Jin. Before that, I completed my MSc in Computer Science at ETH Zurich.

My research is about making language models hold up under adversarial pressure: tamper-resistant safeguards, defenses that survive fine-tuning, and robustness to jailbreaking and prompt injection attacks. More broadly, I am interested in adversarial defenses, causality, representation learning, and misalignment detection. My work has been featured at ICML, COLM, EMNLP and IASEAI. In the past, I have also worked on projects in computational cosmology and AI for healthcare.

Outside of research, I am an avid speedcuber. I have also participated and won prizes in several GraySwan Arena manual jailbreaking contests.

You can find my latest CV here

News

Sep 01, 2026 Excited to be starting my ELLIS PhD, advised by Prof. Bernhard Schölkopf and Prof. Zhijing Jin!
Jul 26, 2026 Happy to have taken part in the MARS 5.0 research sprint in Oxford, UK 🇬🇧.
Jul 08, 2026 Excited to be at ICML 2026 in Seoul, South Korea 🇰🇷, presenting Training with Honeypots!
Jun 07, 2026 Happy to have attended the CHAI 2026 workshop in Asilomar, California 🇺🇸.
Feb 26, 2026 Happy to have attended IASEAI’26 at UNESCO House in Paris, France 🇫🇷.

Selected Publications

  1. Improved Robustness Against Indirect Prompt Injection Can Be Built into LLMs
    M. Özdinçer, Samuel Simko, Bernhard Schölkopf, and Zhijing Jin
    In Conference on Language Modeling (COLM), 2026
  2. Training with Honeypots: Reshaping How LLMs Fail Under Adversarial Attacks
    Samuel Simko, Punya Syon Pandey, Zhijing Jin, and Bernhard Schölkopf
    In International Conference on Machine Learning (ICML), 2026
  3. Position: Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk
    David Guzman Piedrahita, Changling Li, Dave Banerjee, Terry Jingchen Zhang, and 7 more authors
    In International Conference on Machine Learning (ICML), Position Paper Track, 2026
  4. improving.gif
    Improving Large Language Model Safety with Contrastive Representation Learning
    Samuel Simko, Mrinmaya Sachan, Bernhard Schölkopf, and Zhijing Jin
    arXiv preprint arXiv:2506.11938, 2025