Kryptovaluta-ticker:
technology fra Ars Technica AI

LLMs respond differently to harmful prompts when AI watermarking is used

Dan Goodin
Thursday at 18:33
4 Visninger
0 Kommentarer
LLMs respond differently to harmful prompts when AI watermarking is used

In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be “cloudy,” the key might change it to “overcast.” Anyone who knows the key can determine if it was generated by the platform using it. New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to...

Les hele artikkelen hos kilden.

Delta i diskusjonen — kommenter, stem og del lenker.

Registrer
Var dette nyttig?
Del:

Kommentarer (0)

Vennligst logg inn eller registrer deg for å delta i diskusjonen

Ingen kommentarer ennå. Bli den første til å kommentere!