CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

Zeyang Yue, Chenfei Yan, Feifei Zhao, Haibo Tong, Mengwen Xu, Xiaozhen Wang, Erliang Lin, Yi Zeng

Jun 5, 2026 at 04:00

4 Visninger

0 Kommentarer

arXiv:2606.06099v1 Announce Type: new Abstract: Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns. However, existing AI safety benchmarks remain largely restricted to explicit rule compliance and static prompts, failing to capture the dynamic and...

Les hele artikkelen hos kilden.

Les original artikkel

Var dette nyttig?

Del:

Kommentarer (0)

Vennligst logg inn for å skrive en kommentar

Ingen kommentarer ennå. Bli den første til å kommentere!

Relaterte nyheter

Lenke kopiert til utklippstavlen

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

Kommentarer (0)

Relaterte nyheter

Five big questions about the UK's under-16s social media ban

”Ericssons nya vd har jobbat där sedan 2g – nästa skifte blir existentiellt”

[Ekstra] KI gir foreløpig begrenset gevinst i norske virksomheter

[Ekstra] USAs Anthropic-stopp vekker debatt: : – Vårt ansvar, ikke Trumps

ShowCase: Mammotion Luba 3 tar sikte på stökiga tomter

Bla etter kategori