🧊 前沿科技知识库
全部 / 人工智能(AI)

多模态 AI(Multimodal AI)

2026-09-26 · 人工智能(AI)
最后更新:2026-09-26 | 领域:人工智能 / 多模态模型与生成 | 说明:信息来源为公开网络资料,详见文末参考来源

一、概述

多模态 AI 指同时理解与生成文本、图像、音频、视频乃至 3D 世界的模型体系。2025–2026 年的主线是「原生多模态(native multimodal)」——不再把视觉/语音作为外挂模块,而是在同一模型内统一建模;与此同时,图像生成、视频生成、实时语音对话与「世界模型(world model)」四条产品线并行爆发,多模态能力成为前沿模型的默认配置而非附加项。

从能力评测看,多模态仍是「进展快但边界明显」的领域:一方面前沿模型在文档、图表、屏幕与视频理解上快速提升,另一方面在最难的空间推理任务上仍远低于人类——例如在视频空间智能基准 MMSI-Video-Bench 上,表现最好的 Gemini 3 Pro 仅得 38.0 分,而人类为 96.4 分(MMSI-Video-Bench);在 GST-Bench 上,最强的零样本模型 Gemini-3-Pro 得 42.68 分,人类基线为 79.08 分(GST-Bench)。这说明「看得懂」与「真正理解空间/物理」之间仍有明显鸿沟。

二、2025–2026 最新进展

三、核心技术与关键概念

四、代表性项目 / 产品

类别项目 / 产品官方链接
通用多模态 LLMOpenAI GPT-5 系列https://openai.com/index/introducing-gpt-5/
通用多模态 LLMGoogle Gemini 3 系列https://ai.google.dev/gemini-api/docs/gemini-3
开放权重 VLM阿里 Qwen3-VL / Qwen3-Omnihttps://qwen.ai/blog?id=99f0335c4ad9ff6153e517418d48535ab6d8afef
开放权重 VLMGoogle Gemma 3https://ai.google.dev/gemma
图像生成Google Nano Banana Prohttps://deepmind.google/models/gemini-image/
图像生成Black Forest Labs FLUX.2https://bfl.ai/
视频生成OpenAI Sora 2https://openai.com/sora/
视频生成Google DeepMind Veo 3.1https://deepmind.google/models/veo/
视频生成快手 可灵 Kling、字节 Seedancehttps://klingai.com/
实时语音OpenAI GPT-Livehttps://openai.com/index/introducing-gpt-live/
世界模型Google DeepMind Genie 3https://deepmind.google/models/genie/
世界模型World Labs Marble / Atlashttps://www.worldlabs.ai/
注:上表除检索结果直接给出的页面外,部分为厂商客观存在的官方站点域名,具体页面以厂商发布为准。

五、关键数据与评测结果

不同榜单因评测口径、工具调用与推理预算设置不同而给出不一致的排名,引用时需注明来源与日期。

六、趋势与争议

  1. 原生多模态 vs 模态专用模型之争:通用模型持续吞并图像/视频生成能力(如 Gemini 3 Pro 同时提供理解与图像生成),但专用工具(FLUX、Midjourney、可灵等)在特定质量维度仍领先,「通才够用、专才够好」并存。
  2. 世界模型是通向具身智能的桥梁还是营销概念:Genie 3、Cosmos、Atlas、Waymo World Model 等强调对物理世界的预测与因果,并已用于机器人与自动驾驶仿真,但可交互、可持久的 3D 世界在真实任务中的价值仍待验证。
  3. 评测可信度:MMMU-Pro 的提出本身即说明旧基准易被「刷分」;空间/视频基准上模型与人类的巨大差距(38.0 vs 96.4)说明能力被高估的风险,且不同排行榜结果差异明显,反映多模态评测尚未标准化(MMMU-Pro、MMSI-Video-Bench)。
  4. 版权与内容溯源:图像生成工具的差异化开始包含 C2PA 内容凭证、训练数据透明度与商业使用条款等治理维度(Best AI Image Generators Compared),但各平台披露程度不一(如 Midjourney 训练数据被指未披露)。
  5. 开放权重 vs 闭源前沿:Qwen3-VL、Gemma 3、Qwen3.8-Flash-Next 等以开放权重覆盖从端侧到超大 MoE 的档位,正在缩小与闭源前沿在多数实用多模态任务上的差距,但在最难的空间推理上仍与最强闭源模型存在落差(Qwen VL、Which Ollama Models Support Vision?)。

参考来源

  1. Introducing GPT-5 — OpenAI
  2. Introducing GPT-5.4 — OpenAI
  3. GPT-5.6: Frontier intelligence that scales with your ambition — OpenAI
  4. Gemini API 版本说明 — Google AI for Developers
  5. Gemini 3 开发者指南 — Google AI for Developers
  6. Gemini 3 Developer Guide
  7. Gemini 3 Flash — Google Cloud Docs
  8. Qwen VL: See, Read, Reason
  9. Qwen3-VL: Sharper Vision, Deeper Thought, Broader Action
  10. Qwen3-VL-Flash Launched on Model Studio
  11. Qwen3.7-Plus: Multimodal Agent Intelligence
  12. Best AI image generators 2026 (gpt-image-2, Nano Banana Pro)
  13. Best AI Image Generators: Midjourney vs DALL-E vs Stable Diffusion
  14. 10 Best AI Image Generators Compared
  15. Best AI Image Generators August 2026
  16. Sora 2 vs Veo 3 vs Seedance 2.0: Which AI Video Model Actually Wins?
  17. Sora 2 vs Veo 3 vs Runway Gen-4 vs Kling 3 — 2026 Comparison
  18. AI Video Generation — aiwiki.ai
  19. Seedance 2.0 Review (2026)
  20. Presentamos GPT-Live — OpenAI
  21. Genie 3 — Google DeepMind
  22. World Labs 官网
  23. World Labs Research & Insights (Marble / Atlas)
  24. Welcome to Marble — World Labs Docs
  25. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
  26. MMMU-Pro — BenchLM
  27. Artificial Analysis MMMU-Pro (AA-MMMU-Pro)
  28. MMMU-Pro Benchmark Leaderboard — Artificial Analysis
  29. A new era of intelligence with Gemini 3 — Google Blog
  30. Gemini 3 Pro: the frontier of vision AI — Google Blog
  31. Gemini 3.1 Pro — Google DeepMind Model Card
  32. Nano Banana 图片生成 — Google AI for Developers
  33. Nano Banana 🍌 — Google DeepMind
  34. Gemini 3 Pro Image (Nano Banana Pro) — Google AI Studio
  35. Gemini 3 Pro Image (Nano Banana Pro) — Google Cloud Docs
  36. Nano Banana Pro — aicharalab
  37. Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
  38. Which Ollama Models Support Vision?
  39. Meilleures APIs de Computer Vision & modèles open source 2026
  40. 10 Open World AI Models Actually Worth Following in 2026
  41. These AI research labs are building very different world models for robot training
  42. World Models in 2026: Why Google, NVIDIA, LeCun & Fei-Fei Li Are Betting Billions
  43. spatial intelligence — aiwiki
  44. Как устроены world models (Waymo World Model)
  45. MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
  46. GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
  47. MMMU-Pro (PDF)
  48. MMMU-Pro — Interfaze
  49. Veo 3.1 vs Kling 3.0 vs Sora 2: The Definitive April 2026 AI Video Comparison
  50. 2026 AI Video Generation Top 3 Compared: Veo 3.1 vs Kling 3.0 After Sora 2 Shutdown
  51. Best AI Video Generation Models in 2026 Compared
  52. Veo 3.1 vs Kling 3.0 vs Sora 2: Which AI Video Generator Should You Pick in 2026?
  53. Alibaba Launches Qwen3.8-Omni-Flash: Native Multimodal, Million-Context Audio Cost Cut by 98% — AIBase
  54. Qwen3.8-Omni-Flash: AI Model for Audio, Video and Autonomous Agents — AI Arabai
  55. Qwen3.7-Plus: Multimodal Agent Intelligence — Qwen
  56. Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
  57. 100 things we announced at Google I/O 2026 — Google Blog
  58. Text to Speech API — ElevenLabs
  59. Automate web and desktop apps with computer use — Microsoft Learn
  60. Toward Native Multimodal Modeling: A Roadmap — arXiv
  61. Tencent-Hunyuan/HY-World-2.0: Open 3D World Generation from Text, Images, and Video
  62. CC-OCR v2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing — arXiv
  63. DocAtlas: Multilingual Document Understanding — arXiv
  64. Multimodal search in Azure AI Search — Microsoft Learn
  65. Build AI-Ready Knowledge Systems Using 5 Essential Multimodal RAG Capabilities — NVIDIA
  66. EmbodiedBench: A Comprehensive Benchmark for Multimodal LLM-based Embodied Agents
  67. EmbodiedBench Challenge — CVPR 2026 Workshop
  68. Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis — arXiv
  69. OCR and Document AI Leaderboard 2026: Top Models Ranked — Awesome Agents
  70. TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents — arXiv
  71. EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents — arXiv
  72. Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms — arXiv
← MLOps 与 LLMOps开源 AI 工具链 →