Publications

You can find my latest publications here

2026

  1. Last Translation Benchmark
    Vilém Zouhar, Niyati Bafna, Mukund Choudhary, Maike Züfle, and 1 more author
    arXiv preprint arXiv:2609.04173, 2026
  2. Improved Robustness Against Indirect Prompt Injection Can Be Built into LLMs
    M. Özdinçer, Samuel Simko, Bernhard Schölkopf, and Zhijing Jin
    In Conference on Language Modeling (COLM), 2026
  3. Training with Honeypots: Reshaping How LLMs Fail Under Adversarial Attacks
    Samuel Simko, Punya Syon Pandey, Zhijing Jin, and Bernhard Schölkopf
    In International Conference on Machine Learning (ICML), 2026
  4. Keeping Agents on Track: Using Context-Induced Failures as a Learning Signal
    M. Özdinçer, Samuel Simko, Bernhard Schölkopf, and Zhijing Jin
    In AI4GOOD Workshop at ICML, 2026
  5. Position: Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk
    David Guzman Piedrahita, Changling Li, Dave Banerjee, Terry Jingchen Zhang, and 7 more authors
    In International Conference on Machine Learning (ICML), Position Paper Track, 2026
  6. TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
    Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, and 7 more authors
    arXiv preprint arXiv:2602.06911, 2026

2025

  1. improving.gif
    Improving Large Language Model Safety with Contrastive Representation Learning
    Samuel Simko, Mrinmaya Sachan, Bernhard Schölkopf, and Zhijing Jin
    arXiv preprint arXiv:2506.11938, 2025
  2. Causal AI Scientist: Facilitating Causal Data Science with Large Language Models
    Vishal Verma, Sawal Acharya, Samuel Simko, Devansh Bhardwaj, and 5 more authors
    In AI for Science Workshop at NeurIPS, 2025
  3. accidental.gif
    Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
    Punya Syon Pandey, Samuel Simko, Kellin Pelrine, and Zhijing Jin
    arXiv preprint arXiv:2505.16789, 2025