OpenAI has revealed another six occasions on which its AI software behaved unexpectedly or did dangerous things. The startup added the incidents to its misalignment reports page on Wednesday evening, Pacific Time, and described them as follows: · Self-generated prompt injections in compaction summaries · Encouraging deception in compaction summaries · Signing up for disposable emails and searching GitHub for leaked API keys · Uploading files to the internet in order to cite them · Unsanctioned Artifactory writes and cross-sample communication · Unauthorized communication via temporary file hosting services The details are unsettling. The first incident on the list, for example, saw an unreleased model “writing jailbreak-like...
Read the full article at the source.
Comments (0)
No comments yet. Be the first to comment!