原文开头
For decades, progress in the field of artificial intelligence was associated with the same principle: create larger models with more parameters, use more data, and increase computing power. Each time there were some achievements in the performance of AI, the latter grew in complexity, and this is particularly evident when we talk about multimodal AI. The result was amazing capabilities – but, equally, amazing complexity. From OpenAI to Anthropic, from Meta to thousands of open-source projects, multimodal models were always built upon specific encoders that would convert images, audio, and video to a representation understood by a language model. But then Google decided to do something different. …
摘自《科技人工智能杂志》(Tech AI Magazine)2026年7月刊。仅引用开头一小段供了解文章,版权归原刊所有,全文请阅读原刊。