-
배포 시뮬레이션으로 출시 전 모델 동작 예측
Predicting model behavior before release by simulating deployment
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.
-
Import AI 461: 정렬이 진행 중이 아님, FrontierCode, 그리고 합성 연구 인턴
Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research interns
Where are your agents right now?
-
[AINews] Fable과 Mythos, 공식적으로 출시하기에 너무 위험
[AINews] Fable and Mythos officially too dangerous to release
We are in the strangest timeline.
-
AI 에이전트가 DN42를 스캔하려다 운영자를 파산시킴
AI 에이전트가 DN42를 스캔하려다 운영자를 파산시킴 | GeekNews
<ul> <li>AI 에이전트가 <strong>DN42</strong> 가입을 시도하며 네트워크 스캔을 위해 <strong>고사양 AWS 인스턴스를 배포</strong>했고, 결국 운영자에게 <strong>$6531.30 청구서</strong>를 남긴 사건</li> <li>DN42는 BGP와 DNS 등 인터넷 백본 기술을…
-
Anthropic, 보이지 않는 Claude Fable 가드레일에 사과함
Anthropic, 보이지 않는 Claude Fable 가드레일에 사과함 | GeekNews
<ul> <li><strong>Claude Fable 5</strong>는 Anthropic의 Mythos 계열에서 처음 널리 제공된 모델이며, 경쟁 시스템 개발에 쓰이는 증류 시도를 막기 위해 숨겨진 제한을 적용했음</li> <li>Anthropic은 증류로 판단한 요청에 대해 사용자에게 알리지 않고 <strong>응답…
-
Anthropic의 첫 번째 공개 기록 결과
Results from first Anthropic Public Record
-
Anthropic이 Claude Fable의 숨겨진 가드레일에 대해 사과
Anthropic apologizes for invisible Claude Fable guardrails
<p><a href="https://web.archive.org/web/20260611122253/https://www.theverge.com/ai-artificial-intelligence/948280/anthropic-claude-fable-invisible-distillation-guardrail" rel="nofo…
-
사이버보안 연구자들, Anthropic의 Fable 보안 조치에 불만
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable | TechCrunch
<p>Article URL: <a href="https://techcrunch.com/2026/06/10/cybersecurity-researchers-arent-happy-about-the-guardrails-on-anthropics-fable/">https://techcrunch.com/2026/06/10/cybers…
-
Anthropic Claude Fable 5 - Mythos급이지만 안전함, 논쟁의 여지가 있는 약관과 함께
[AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms
The much anticipated launch of the Mythos-class model was marred by some controversial usage policies
-
Claude Fable이 도움을 멈춰도 사용자는 알 수 없다
Claude Fable이 도움을 멈춰도 사용자는 알 수 없다 | GeekNews
<ul> <li>코딩 보조 모델이 경쟁 LLM 개발 요청에서 사용자에게 알리지 않고 효과를 제한할 수 있어, 개발 도구 신뢰에 <strong>공급망 위험</strong>이 생김</li> <li>Anthropic은 Fable 5에서 프런티어 LLM 개발 요청에 대한 효과 제한을 도입했고, 이 제한은 사용자에게 <strong…
-
Claude Fable이 도움을 멈춰도 당신은 알 수 없을 것이다
If Claude Fable stops helping you, you'll never know — Jonathon Ready
<p>Related: https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-helping-you/</p> <hr /> <p>Comments URL: <a href="https://news.ycombinator.com/item?id=48467896">https://new…
-
여러 제품에서 Claude를 제한하는 방법 - 에이전트가 더 강력해질수록 잠재적 위험 범위도 커집니다. 핵심 엔지니어링 과제는 이를 어떻게 제한할 것인가이며, claude.ai, Claude Code, Cowork 개발에서 배운 제한 기능 구축 방법입니다.
How we contain Claude across products
-
자동화된 정렬 연구자들: 대규모 언어 모델을 활용한 확장 가능한 감시
Automated Alignment Researchers: Using large language models to scale scalable oversight
-
Anthropic은 제품 전반에서 Claude를 어떻게 봉쇄할까
Anthropic은 제품 전반에서 Claude를 어떻게 봉쇄할까 | GeekNews
<ul> <li>에이전트의 능력과 접근권한이 커질수록 <strong>잠재적 피해 반경</strong> 도 함께 확대되며, 클로드 웹/Claude Code/Cowork 각각에 맞춘 봉쇄 아키텍처 구축 경험을 정리</li> <li>위험은 <strong>실패 가능성</strong>과 <strong>피해 규모</strong> 두…
-
수학자들이 AI의 빠른 진전에 경고를 발표하다
Mathematicians issue warning as AI rapidly gains ground
<p>Article URL: <a href="https://www.science.org/content/article/mathematicians-issue-warning-ai-rapidly-gains-ground">https://www.science.org/content/article/mathematicians-issue-…
-
OpenAI 공공정책 의제
OpenAI public policy agenda | OpenAI
OpenAI outlines its public policy agenda for AI, including safety, youth protection, workforce transition, and global standards to ensure AI benefits society.
-
프론티어 AI의 민주적 거버넌스를 위한 청사진 | OpenAI
A blueprint for democratic governance of frontier AI | OpenAI
OpenAI outlines a blueprint for U.S. governance of frontier AI, proposing a federal framework for safety, resilience, and national security.
-
글로벌 리더십을 통한 청소년 안전과 기회 증진
Advancing youth safety and opportunity through global leadership | OpenAI
OpenAI calls for global action on youth AI safety, proposing an international institute to strengthen safeguards, standards, and opportunities for young people.
-
AI 정책 및 정치적 옹호에 관한 우리의 입장
Our views on AI policy and political advocacy | OpenAI
Our approach to AI policy and political advocacy, transparency, support for thoughtful regulation and AI safety, and that no outside political group speaks on the company’s behalf.
-
Import AI 459: AI 감시의 어려움, 단백질 폴딩 모델의 스케일링 법칙, AI 시스템의 멸종 위험 가격 책정
Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems
Do you feel as though you are living in a revolution?
-
신뢰할 수 있는 제3자 평가를 위한 공유 플레이북
A shared playbook for trustworthy third party evaluations | OpenAI
OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.
-
OpenAI의 최첨단 거버넌스 프레임워크
OpenAI’s Frontier Governance Framework | OpenAI
Explore OpenAI’s Frontier Governance Framework and how our AI safety, security, and risk practices align with emerging EU and California regulations.
-
2026년 선거 정보 및 보안조치
Election information and safeguards in 2026 | OpenAI
Ahead of global elections, we’re helping people access information, supporting cyber defenders, and increasing AI transparency
-
임포트 AI 457: AI 스턱스넷; 저주받은 뮤온 최적화기; 긍정적 정렬
Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignment
Welcome to Import AI, a newsletter about AI research.
-
ChatGPT가 민감한 대화에서 맥락을 더 잘 인식하도록 돕기
Helping ChatGPT better recognize context in sensitive conversations | OpenAI
Learn how new ChatGPT safety updates improve context awareness in sensitive conversations, helping detect risk over time and respond more safely.
-
2026년 5월 7일 AI 정렬: 우리의 오픈소스 정렬 도구 기증
Donating our open-source alignment tool
-
Import AI 454: 정렬 연구 자동화; 중국 모델의 안전성 연구; HiFloat4
Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4
At what point do the financial markets price in the singularity?
-
Import AI 453: AI 에이전트 붕괴, 미러코드 그리고 점진적 무력화에 대한 열 가지 관점
Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowerment
Was fire equivalent to a singularity for people at the time?
-
Claude Code 자동 모드를 구축한 방법: 권한 건너뛰기의 더 안전한 방식
How we built Claude Code auto mode: a safer way to skip permissions
-
Import AI 443: 안개 속으로: 몰트북, 에이전트 생태계, 그리고 전환기의 인터넷
Import AI 443: Into the mist: Moltbook, agent ecologies, and the internet in transition
Plus, a story about agents corrupting other agents