Safety

Safety

Safety covers harm, misuse, and the research that tries to understand or prevent it. Alignment and interpretability findings, jailbreaks, deepfakes, model-enabled attacks, security incidents involving AI systems, and evaluations that surface dangerous capabilities land here. A paper showing chain-of-thought is unfaithful is Safety; a paper claiming a new SOTA on SWE-bench is Models. Product abuse in the wild can be Safety when the harm or attack is the news, and Adoption when the story is deployment and labor effects without a misuse frame. This hub is for the risk surface, not a synonym for the legal /security page.