Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration Paper • 2609.16204 • Published 7 days ago • 4
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks Paper • 2609.15029 • Published 7 days ago • 8
Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language Paper • 2510.23828 • Published Oct 27, 2025 • 2
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models Paper • 2510.10390 • Published Oct 12, 2025 • 5
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration Paper • 2609.16204 • Published 7 days ago • 4
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration Paper • 2609.16204 • Published 7 days ago • 4
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks Paper • 2609.15029 • Published 7 days ago • 8
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks Paper • 2609.15029 • Published 7 days ago • 8