Skip to content
AI News HubLIVE
Source content · Analysis pending1 min read

OpenAI reveals cases of ‘concerning’ AI behaviour and promises new plan for disclosing issues

Summary

Research model inserting ‘jailbreak-like instructions’ into its notes is among cases as company says it is introducing new way of tracking AI misalignment OpenAI has disclosed six new reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated. Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”. Continue reading...

SourceThe Guardian AIAuthor: Associated Press
OpenAI reveals cases of ‘concerning’ AI behaviour and promises new plan for disclosing issues
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

Research model inserting ‘jailbreak-like instructions’ into its notes is among cases as company says it is introducing new way of tracking AI misalignment OpenAI has disclosed six new reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated. Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”. Continue reading...

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Research model inserting ‘jailbreak-like instructions’ into its notes is among cases as company says it is introducing new way of tracking AI misalignment OpenAI has disclosed six…

Highlights and analysis are generated automatically and may contain errors. Check the original source.