-
Claude 4의 사이버 평가
Cyber evaluations of Claude 4
-
배포 시뮬레이션으로 출시 전 모델 동작 예측
Predicting model behavior before release by simulating deployment
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.
-
TensorZero: LLM 게이트웨이, 관찰성, 평가, 최적화, 실험을 통합하는 오픈소스 LLMOps 플랫폼
GitHub - tensorzero/tensorzero: TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
<p>Article URL: <a href="https://github.com/tensorzero/tensorzero">https://github.com/tensorzero/tensorzero</a></p> <p>Comments URL: <a href="https://news.ycombinator.com/item?id=4…
-
실제 환경에서 AI 에이전트 자율성 측정
Measuring AI agent autonomy in practice