跳到正文
原文
Google DeepMind·· 2026-06-09精选AI 评分78

Google DeepMind 发布 Gemma 4 12B 无编码器多模态模型

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

AI 导读

Google DeepMind 发布 Gemma 4 12B,这是其首款支持原生音频输入的中型多模态模型,采用无编码器统一架构,视觉和音频输入直接进入 LLM 主干。

推荐理由

原文给出了架构变化、内存占用和开放入口,读者可据此判断它能否在本地硬件上跑起多模态智能体工作流。

来源:Google DeepMind · deepmind.google