Research Alert: China's Silicon Reawakening: Meituan's LongCat Lab and the Ascend-Powered Dawn of Sovereign AI
In the crucible of technological containment, China's AI ascent has assumed the character of a sophisticated national symphony—platform capital conducting hardware ingenuity, academic sparsity orchestrating around bandwidth austerity. The latest movement features LongCat-2.0, Meituan's audacious 1.6-trillion-parameter MoE titan: the first publicly chronicled frontier model whose entire pre-training and inference lifecycle transpired on 50,000 Ascend 910C cards within CloudMatrix 384 Superpods. Pre-trained on over 35 trillion tokens with dynamic activation of 33–56 billion parameters per token (averaging ~48B), native 1M context via LongCat Sparse Attention (LSA) —an evolution of DeepSeek DSA deploying lighter indexers for near-linear scaling—and 135B N-gram Embedding parameters at 5-gram depth, it delivers 59.5 on SWE-Bench Pro and 70.8% on Terminal-Bench 2.1. This is no mere parameter flex; it is the super-app's declaration that intelligence must be endogenous to command the physical world's chaotic data streams. - Meituan's dominion as China's lifestyle super-app—melding instantaneous delivery, mobility, hospitality, and discovery into one pulsating interface—renders generic.
In the crucible of technological containment, China's AI ascent has assumed the character of a sophisticated national symphony—platform capital conducting hardware ingenuity, academic sparsity orchestrating around bandwidth austerity. The latest movement features LongCat-2.0, Meituan's audacious 1.6-trillion-parameter MoE titan: the first publicly chronicled frontier model whose entire pre-training and inference lifecycle transpired on 50,000 Ascend 910C cards within CloudMatrix 384 Superpods. Pre-trained on over 35 trillion tokens with dynamic activation of 33–56 billion parameters per token (averaging ~48B), native 1M context via LongCat Sparse Attention (LSA) —an evolution of DeepSeek DSA deploying lighter indexers for near-linear scaling—and 135B N-gram Embedding parameters at 5-gram depth, it delivers 59.5 on SWE-Bench Pro and 70.8% on Terminal-Bench 2.1. This is no mere parameter flex; it is the super-app's declaration that intelligence must be endogenous to command the physical world's chaotic data streams. - - Meituan's dominion as China's lifestyle super-app—melding instantaneous delivery, mobility, hospitality, and discovery into one pulsating interface—renders generic frontier models insufficient. Its proprietary graph of real-time inventories, 1.3 billion reviews, and hyper-frequent user intents demands agents capable of autonomous orchestration: parsing vague spoken directives into executed bookings against live merchant states. Xiaomei already prototypes this; LongCat-2.0 elevates it. In an ecosystem where super-apps like WeChat, Douyin, and Taobao vie for primacy in daily autopilot existence, ceding the foundational layer equates to strategic surrender. Meituan's investment, born of survival calculus amid subsidy wars and margin compression, echoes yet surpasses Meta's proprietary stake: here, the LLM becomes the invisible conductor of commerce's physical symphony. - - The human filament threading this narrative glows with quiet poignancy. In early 2023, co-founder Wang Huiwen —fresh from Meituan tenure—poured personal fortune into Light Years Beyond, envisioning China's OpenAI and igniting a cohort of labs including DeepSeek and Moonshot. Mental health tribulations led to Meituan's ~ 2.065 billion RMB acquisition, absorbing talent and vision into LongCat Lab. What emerged transcends the original cadre: MOPD (Multi-Expert On-Policy Distill) fuses dedicated Agent (tool-calling, self-correction), Reasoning (STEM depth), and Interaction (nuanced fidelity) experts through on-policy distillation, dynamically gated within a single model. Paired with zero-computation experts that route simple tokens to negligible cost and N-gram embeddings—hashing sequences via decomposed sub-tables, linear projections, and amplification to sidestep collision while preserving residual pathway potency—this orthogonal sparsity dimension outperforms pure expert scaling on empirical Pareto frontiers. - - Huawei's CloudMatrix384 supplies the architectural loom
This is a real-time flash headline. A verified provider body is not available, so SquawkNews does not present it as a full article.