-
언어 모델의 설득력 측정하기
Measuring the Persuasiveness of Language Models
-
"프론티어 AI가 의료 전문 툴 이겼다"는 논문 재검증해보니 — 채점자간 일치도 0.10, 채점자가 곧 참가자
"프론티어 AI가 의료 전문 툴 이겼다"는 논문 재검증해보니 — 채점자간 일치도 0.10, 채점자가 곧 참가자
<h3>간략 요약</h3> <ul> <li>Nature Medicine에 2026년 6월 12일 게재된 논문 "General-purpose large language models outperform specialized clinical AI tools on medical benchmarks"에서 GPT-…
-
LLM 평가의 4가지 주요 접근법 이해하기 (기초부터)
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)
Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples