-
OpenAI와 Hugging Face, 모델 평가 중 보안 사건에 대응하기 위해 협력
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
-
장시간 실행 모델 시대의 안전성과 정렬
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
-
청소년이 안전한 AI에 접근할 권리
Why teens deserve access to safe AI
Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partnerships.
-
미국이 주 및 연방 차원의 조치를 통해 AI 안전을 진전시키고 있다
The US is advancing AI safety through state and federal action
OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
-
GPT-Red: 견고성을 위한 자가 개선의 활용
GPT-Red: Unlocking Self-Improvement for Robustness
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
-
우리의 프론티어 레드 팀으로부터의 진전
Progress from our Frontier Red Team
-
에이전틱 오정렬: LLM이 내부자 위협이 될 수 있는 방식
Agentic misalignment: How LLMs could be insider threats
-
AI 안전을 위한 프론티어 위협 레드팀
Frontier threats red teaming for AI safety
-
실제 AI 사용에서의 권한 박탈 패턴
Disempowerment patterns in real-world AI usage
-
Constitutional Classifiers: 보편적 탈옥으로부터의 방어
Constitutional Classifiers: Defending against universal jailbreaks
-
언어 모델의 설득력 측정하기
Measuring the Persuasiveness of Language Models
-
AI 모델의 이중 사용 지식을 위한 오프 스위치
An off switch for dual use knowledge in AI models
-
Claude를 위한 안전장치 구축
Building safeguards for Claude
-
Anthropic의 책임 있는 스케일링 정책 발표
Announcing Anthropic's Responsible Scaling Policy