Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

Yilong Wang, Qianli Wang, Bohao Chu, Yihong Liu, Jing Yang, Simon Ostermann

Jun 5, 2026 at 04:00

10 Views

0 Comments

arXiv:2605.11632v2 Announce Type: replace-cross Abstract: Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), offering a causally grounded approach to unraveling black-box LLM behavior. Yet extending them beyond English...

Read the full article at the source.

Read Original Article

Was this helpful?