-
OpenAI와 Hugging Face, 모델 평가 중 보안 사건에 대응하기 위해 협력
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
-
장시간 실행 모델 시대의 안전성과 정렬
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
-
청소년이 안전한 AI에 접근할 권리
Why teens deserve access to safe AI
Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partnerships.
-
미국이 주 및 연방 차원의 조치를 통해 AI 안전을 진전시키고 있다
The US is advancing AI safety through state and federal action
OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
-
GPT-Red: 견고성을 위한 자가 개선의 활용
GPT-Red: Unlocking Self-Improvement for Robustness
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
-
우리의 프론티어 레드 팀으로부터의 진전
Progress from our Frontier Red Team
-
에이전틱 오정렬: LLM이 내부자 위협이 될 수 있는 방식
Agentic misalignment: How LLMs could be insider threats
-
AI 안전을 위한 프론티어 위협 레드팀
Frontier threats red teaming for AI safety
-
실제 AI 사용에서의 권한 박탈 패턴
Disempowerment patterns in real-world AI usage
-
Constitutional Classifiers: 보편적 탈옥으로부터의 방어
Constitutional Classifiers: Defending against universal jailbreaks
-
언어 모델의 설득력 측정하기
Measuring the Persuasiveness of Language Models
-
AI 모델의 이중 사용 지식을 위한 오프 스위치
An off switch for dual use knowledge in AI models
-
Claude를 위한 안전장치 구축
Building safeguards for Claude
-
Anthropic의 책임 있는 스케일링 정책 발표
Announcing Anthropic's Responsible Scaling Policy
-
Fable 5의 사이버 보안 세부 사항과 탈옥 방지 프레임워크
More details on Fable 5’s cyber safeguards and our jailbreak framework
-
Anthropic의 AI 안전에 대한 핵심 견해
Anthropic's core views on AI safety
-
AI를 위한 핵 안전장치 개발
Developing Nuclear Safeguards for AI
-
고급 AI를 위한 공유 표준 구축 지원
Helping build shared standards for advanced AI
OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the Appia Foundation.
-
Import AI 462: 초설득, 자립 AI, ASI 경로
Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI
How religious are beliefs in the singularity?
-
공민 파트너십을 통한 AI 핵 안전장치 개발
Developing nuclear safeguards for AI through public-private partnership
-
배포 시뮬레이션으로 출시 전 모델 동작 예측
Predicting model behavior before release by simulating deployment
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.
-
Import AI 461: 정렬이 진행 중이 아님, FrontierCode, 그리고 합성 연구 인턴
Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research interns
Where are your agents right now?
-
[AINews] Fable과 Mythos, 공식적으로 출시하기에 너무 위험
[AINews] Fable and Mythos officially too dangerous to release
We are in the strangest timeline.
-
Anthropic의 첫 번째 공개 기록 결과
Results from first Anthropic Public Record
-
Anthropic Claude Fable 5 - Mythos급이지만 안전함, 논쟁의 여지가 있는 약관과 함께
[AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms
The much anticipated launch of the Mythos-class model was marred by some controversial usage policies
-
여러 제품에서 Claude를 제한하는 방법 - 에이전트가 더 강력해질수록 잠재적 위험 범위도 커집니다. 핵심 엔지니어링 과제는 이를 어떻게 제한할 것인가이며, claude.ai, Claude Code, Cowork 개발에서 배운 제한 기능 구축 방법입니다.
How we contain Claude across products
-
자동화된 정렬 연구자들: 대규모 언어 모델을 활용한 확장 가능한 감시
Automated Alignment Researchers: Using large language models to scale scalable oversight
-
OpenAI 공공정책 의제
OpenAI public policy agenda | OpenAI
OpenAI outlines its public policy agenda for AI, including safety, youth protection, workforce transition, and global standards to ensure AI benefits society.
-
프론티어 AI의 민주적 거버넌스를 위한 청사진 | OpenAI
A blueprint for democratic governance of frontier AI | OpenAI
OpenAI outlines a blueprint for U.S. governance of frontier AI, proposing a federal framework for safety, resilience, and national security.
-
글로벌 리더십을 통한 청소년 안전과 기회 증진
Advancing youth safety and opportunity through global leadership | OpenAI
OpenAI calls for global action on youth AI safety, proposing an international institute to strengthen safeguards, standards, and opportunities for young people.