🔍 Read the full analysis: The Coming Era Of Multimodal AI: Exclusive Insights From SenseTime’s Lead Scientist on ThorstenMeyerAI.com
TL;DR
SenseTime’s chief scientist Lin Dahua predicts that a major breakthrough in multimodal AI — systems understanding and generating across text, images, and video — is likely within one to two years. This forecast could influence industry innovation and investment, but remains unconfirmed by external benchmarks.
SenseTime’s chief scientist Lin Dahua has stated in an exclusive interview with 36Kr that a major breakthrough in multimodal AI systems is likely to occur within the next one to two years. This prediction points to a significant acceleration in AI systems capable of understanding and generating across multiple data modalities, including text, images, audio, and video. The statement, based on Lin’s perspective and ongoing research, highlights a potential paradigm shift in AI technology that could impact various industries and market players.
In the interview, Lin Dahua emphasized that the upcoming one-to-two-year window could mark a transition from incremental improvements to a decisive leap in multimodal AI capabilities. He explained that SenseTime’s research focus on foundation models—large AI systems trained to process and connect multiple data types—positions the company to be at the forefront of this development. While the full interview transcript was not publicly available, Lin’s statement is the most specific timeline shared by a senior researcher at a major Chinese AI firm about multimodal progress.
SenseTime, traditionally known for computer vision and facial recognition, has recently shifted toward developing its SenseNova foundation model platform. The company aims to compete with other Chinese tech giants like Baidu, Alibaba, and ByteDance, which are also investing heavily in large multimodal models. The forecast suggests that products built on these models—such as AI assistants capable of watching, listening, reading, and acting across formats—could become commercially viable before the end of the decade. However, the prediction is based on Lin’s assessment rather than published benchmarks or peer-reviewed research, making it a forecast rather than a confirmed milestone.
Implications for Industry and Market Leadership
This forecast signals a potential turning point in AI technology that could dramatically alter how machines understand and interact with human data. If realized, a breakthrough within one to two years could enable new applications in autonomous driving, content creation, virtual assistants, and multimedia analysis, reshaping competitive dynamics across the AI landscape. For Chinese AI firms like SenseTime, this timeline underscores a strategic push to accelerate innovation and differentiate from U.S. rivals such as OpenAI and Google, who are also advancing multimodal models. A rapid development cycle may influence investment flows, product launches, and global competitiveness, making Lin’s prediction a key indicator for industry watchers and investors.
As an affiliate, we earn on qualifying purchases.
Recent Trends and Industry Momentum in Multimodal AI
Over the past two years, progress in multimodal AI—integrating text, images, videos, and audio—has accelerated remarkably. Leading companies and research labs have demonstrated increasingly capable models, with notable advances in video understanding and unified multimodal systems. SenseTime’s shift from pure computer vision to foundation models reflects a broader industry trend: merging different data modalities into single, versatile systems. Chinese AI companies, including SenseTime, are investing heavily to catch up and compete with international leaders who have already showcased early multimodal capabilities. The industry’s rapid pace of progress fuels optimism about a potential breakthrough within the next year or two, aligning with Lin Dahua’s forecast.
“The multimodal AI breakthrough moment is coming in one to two years.”
— Lin Dahua, SenseTime chief scientist
As an affiliate, we earn on qualifying purchases.
Unconfirmed Benchmarks and Technical Details
The full technical basis for Lin Dahua’s prediction remains unclear. No published benchmarks, internal datasets, or peer-reviewed studies have been cited to substantiate the one-to-two-year timeline. It is uncertain whether SenseTime’s ongoing research is approaching a specific technical milestone or if this is a strategic projection based on industry trends. The prediction’s accuracy depends on future developments, and it could be affected by unforeseen technical challenges or slower-than-expected progress.

AI for Content Creation: The Ultimate Guide
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Monitoring SenseTime’s Model Releases and Industry Benchmarks
To assess the validity of Lin Dahua’s forecast, industry observers should watch for upcoming SenseTime model releases, especially updates to its SenseNova platform, and any published multimodal benchmarks. Industry-wide, successive demonstrations of video-understanding and multimodal reasoning capabilities over the next 12–24 months will serve as critical indicators. Additionally, a fuller transcript or official statements from SenseTime could clarify the scope and basis of the timeline. These developments will help determine whether the predicted breakthrough is on track or if the timeline needs adjustment.
virtual assistant with multimodal capabilities
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Who is Lin Dahua?
Lin Dahua is the chief scientist of SenseTime, leading its research efforts in artificial intelligence and foundation models, with a focus on multimodal systems.
What exactly did he predict?
He forecasted that a significant breakthrough in multimodal AI systems—capable of understanding and generating across text, images, audio, and video—will likely occur within one to two years.
Is this a confirmed fact?
No, it is a forecast based on Lin Dahua’s expert opinion; no external benchmarks or published data currently confirm this timeline.
Why does this prediction matter?
If accurate, it could signal a major shift in AI capabilities, enabling new applications and intensifying global competition, especially for Chinese AI firms seeking to lead in multimodal AI development.
Primary source: SenseTime · via ThorstenMeyerAI.com