Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

Yilong Wang, Qianli Wang, Bohao Chu, Yihong Liu, Jing Yang, Simon Ostermann

Jun 5, 2026 at 04:00

9 Visninger

0 Kommentarer

arXiv:2605.11632v2 Announce Type: replace-cross Abstract: Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), offering a causally grounded approach to unraveling black-box LLM behavior. Yet extending them beyond English...

Les hele artikkelen hos kilden.

Les original artikkel

Var dette nyttig?

Del: