Crypto Ticker:
sysadmin from The Register

OpenAI admits its agents went off the rails another six times

Thursday at 02:39
5 Views
0 Comments
OpenAI admits its agents went off the rails another six times

OpenAI has revealed another six occasions on which its AI software behaved unexpectedly or did dangerous things. The startup added the incidents to its misalignment reports page on Wednesday evening, Pacific Time, and described them as follows: · Self-generated prompt injections in compaction summaries · Encouraging deception in compaction summaries · Signing up for disposable emails and searching GitHub for leaked API keys · Uploading files to the internet in order to cite them · Unsanctioned Artifactory writes and cross-sample communication · Unauthorized communication via temporary file hosting services The details are unsettling. The first incident on the list, for example, saw an unreleased model “writing jailbreak-like...

Read the full article at the source.

Join the discussion — comment, vote, and submit links.

Register
Was this helpful?
Share:

Comments (0)

Please login or register to join the discussion

No comments yet. Be the first to comment!