Australia says OpenAI agent broke into government system, launches further checks
Sep 24, 7:08 AM • 2 min read
Safety
Safety covers harm, misuse, and the research that tries to understand or prevent it. Alignment and interpretability findings, jailbreaks, deepfakes, model-enabled attacks, security incidents involving AI systems, and evaluations that surface dangerous capabilities land here. A paper showing chain-of-thought is unfaithful is Safety; a paper claiming a new SOTA on SWE-bench is Models. Product abuse in the wild can be Safety when the harm or attack is the news, and Adoption when the story is deployment and labor effects without a misuse frame. This hub is for the risk surface, not a synonym for the legal /security page.
Sep 24, 7:08 AM • 2 min read
Sep 23, 9:38 PM • 1 min read
Sep 21, 11:42 AM
Sep 21, 4:16 AM • 2 min read
Sep 20, 9:54 PM
Sep 18, 9:08 PM
Sep 18, 7:21 AM • 2 min read
Sep 18, 4:14 AM • 3 min read
Sep 17, 3:01 PM
Sep 17, 9:12 AM
Sep 16, 2:44 PM
Sep 16, 2:18 PM
Sep 15, 9:06 AM • 2 min read
Sep 14, 9:05 PM
Sep 14, 9:42 AM
Sep 12, 10:16 PM
Sep 11, 9:07 AM • 2 min read
Sep 10, 8:36 PM
Sep 9, 7:40 AM • 3 min read
Sep 8, 3:49 PM
Sep 4, 8:49 PM
Sep 2, 2:09 PM
Sep 1, 3:01 PM
Aug 28, 2:46 PM • 2 min read