-
DeepSeek Vision 발표
DeepSeek Introduces Vision
<p>Article URL: <a href="https://chat.deepseek.com/">https://chat.deepseek.com/</a></p> <p>Comments URL: <a href="https://news.ycombinator.com/item?id=48581458">https://news.ycombi…
-
VLM은 한국 공공기관 문서를 얼마나 잘 읽을까? KOLongDoc 벤치마크 공개
VLM은 한국 공공기관 문서를 얼마나 잘 읽을까? KOLongDoc 벤치마크 공개 | GeekNews
<p>🔥 한국어 Long-Document VLM 벤치마크, <a href="https://github.com/Marker-Inc-Korea/KOLongDoc">KOLongDoc</a>를 공개했습니다!</p> <p>최근 ChatGPT, Claude, Gemini 같은 멀티모달 AI가 공공·행정 업무에도 활용되기 시작했지만,…
-
Gemma 4 12B: 인코더 없는 통합 멀티모달 모델
Gemma 4 12B: A unified, encoder-free multimodal model
<p>Article URL: <a href="https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/">https://blog.google/innovation-and-ai/technology/developers-too…
-
왜 비디오 에이전트 모델이 다음이 될까 — Ethan He, xAI Grok Imagine
Why Video Agent models are next — Ethan He, xAI Grok Imagine
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and why Grok Imagine is so underrated. For the first time, we do a deep dive with the guy who led it!
-
[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (비디오용 NanoBanana), Spark (백그라운드 에이전트), Antigravity 2.0
[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0
Google has been busy!