首页 > AI前沿 > A dark horse enters China's AI race: StartLux

A dark horse enters China's AI race: StartLux

Hacker News 2026-09-03 19:12 4 阅读 查看原文
9 min read Chen Danian returns, enters large model race Something big has happened - a dark horse has emerged among China's top domestic large model developers. A new company, with its first model having only 27B parameters, took second place overall in the CAICT MCP specialized test. Ranked ahead of it is the killer weapon Liang Wenfeng has kept under wraps for nearly a year: DeepSeek-V4-Pro, boasting a parameter count of 1.6 trillion. The difference between the two is a mere 1.3 percentage points. This dark horse is the StartLux-V1.0-27B-Preview, from StartLux (formerly Yuandian Xinghui). Looks a bit unfamiliar, doesn't it? Don't worry, its founder and CEO is an old acquaintance: Chen Danyan. Known as the "godfather" of programmers in the internet era, his achievements are a matter of public record: Shanda Network's co-founder and the head of Lian Shang Network, one of China's earliest programmers to introduce the concept of "shareware" After a decade of retirement, he has made a comeback, this time betting on local models. This has to do with Chen Dawei's recent rare public appearance. At the 18th anniversary reunion of Shanda Innovation Institute, he publicly declared "eight non-consensus views for the AI era," four of which center on local models. Local models will completely destroy the cloud market, catching up with Claude in three years and occupying 80% of the market. The model competition based on parameters is going to be outdated... StartLux is the best representation of his idea. The company's business has not followed the industry trend, instead focusing on the commercialization of small and beautiful local models, making it the world's first truly local model company in the true sense. As the first market-oriented scorecard, StartLux-V1.0-27B-Preview does not rely on the cloud and can run directly on consumer-grade PCs. In other words, this local model, which is nearly 60 times smaller, has Agent capabilities that can match those of a trillion-level cloud-based flagship model. What justification is there for this? Small Model Achieves Big Results, 27B Outperforms 1.6T Before the answer is revealed, let's take a look at who the comparison is being made to. It's often said that nobody remembers the second place, unless the first is DeepSeek. Moreover, the gap is minimal, making it well worth discussing. The results come from the authoritative institution, China Academy of Information and Communications Technology's trustworthy AI large model benchmark test MCP special item, which sets six types of tasks around real application scenarios: Location navigation, web search, browser automation, financial analysis, code repository management, 3D design. An additional comprehensive assessment will be added, focusing on evaluating Agent's multi-tool collaboration, complex task execution, and interaction in real-world environments. In simple terms, MCP-Universe doesn't evaluate models based on their responses, but rather on whether they can actually get things done. This is also the most fundamental aspect of judging an Agent's quality. The results showed that StartLux-V1.0-27B-Preview had a comprehensive score of 39.25%, ranking second. DeepSeek-V4-Flash-0731, with over 284 billion parameters, and Step-3.7-Flash, at 198 billion, trail DeepSeek-V4-Pro by just 1.3 percentage points. With the same 27B parameter scale, StartLux also surpasses Qwen 3.6 by 5.34 percentage points. It also excelled in individual subjects, ranking first in location navigation, financial analysis, and browser automation, with its other sub-items also ranking high. Let's take a look at two case studies, putting data aside. The first question is about a two-year Microsoft stock investment, with Claude Sonnet 4.6 as the competing topic. Claude's answer is: $47,254, 89.02%. Video link: https://mp.weixin.qq.com/s/365CtdgGYFlEKNDICtoCfg StartLux gave: $47,499.09, 90.00%. It may seem similar, but in the financial industry, a tiny difference can lead to enormous losses. Careful examination of the two models' reasoning processes shows that, due to missing raw data, Claude Sonnet 4.6 misidentified January 8, 2025 as a non-trading day and instead calculated the previous day's closing price. Under the same circumstances, StartLux retrospectively reviews the original data and verifies the market trends around the target date to confirm the accurate closing price before completing the calculation and generating visualization. Ultimately, the conclusion reached by StartLux proved correct, and it was fully verifiable and traceable. The second task was more straightforward: both models were asked to search for flight tickets in a browser at the same time. I'm unable to open a browser or interact with live websites like Google Flights. I can only process text and don't have real-time browsing capabilities. To find this flight yourself, here's what you'd do: 1. Go to google.com/flights 2. Enter Singapore (SIN) → Beijing 3. Select one-way, set departure date to 5 days from today 4. Filter by "Nonstop" and "Economy" 5. Check the results — note that flights to Beijing may land at either Capital (PEK) or Daxing (PKX). Exclude any Daxing arrivals. 6. Compare prices and pick the cheapest nonstop option landing at PEK. If you'd like, I can help you think through typical price ranges, airline options on this route, or general tips for finding cheap flights. Among them, StartLux-V1.0-27B-Preview completed the search in about 95 seconds, finding an Air China ticket priced at $299. By contrast, Claude Sonnet 4.6 required more screenshot confirmations when navigating date selection, popup dismissal, and filter menus, and even accidentally triggered the time filter panel at one point. More than 200 seconds later, it gave a lowest quote of $556. The price was higher, and the search time was still twice that of StartLux. In particular, in terms of operational pathways, StartLux is much more concise, requiring only 12 steps, whereas Sonnet requires a full 21 steps. This is enough to illustrate that, in Agent tasks, the scale of parameters is no longer the only decisive variable. New variables are being introduced through post-training. Cutting Prices, Not Capabilities StartLux-V1.0-27B-Preview was not trained from scratch. It is also based on Qwen3.6-27B, but the final test score is significantly higher than Qwen, and the reason lies in the task data and automated post-training methods. In simple terms, StartLux trains a more capable Agent, with training data focusing on reinforcing abilities such as task understanding, tool selection, parameter construction, multi-step execution, status checking, and result verification. The model needs to learn not only to output the final text, but also when to invoke which tool, how to adjust when a tool returns an exception, and under what circumstances it can declare the task complete. Further training will push this process even further. The team has independently developed a brand-new, multi-dimensional, verifiable, and scalable model iteration optimization technology, which uses the AI-trained AI (Auto Research) method to enable the model to autonomously execute tasks in a real-world tool environment and continuously adjust its strategy based on environmental feedback. For instance, the financial analysis case mentioned earlier, which involves backtesting and revision, as well as the constraint identification and path selection in browser tasks, are the most intuitive manifestations of post-training. According to official information, StartLux-V1.0-27B-Preview is also the country's first local Agent model to complete post-training using the Auto Research method. This does not mean that the Scaling Law is invalid. Large parameter cloud models are still the mainstream choice at present, but StartLux has also given a clear signal: this is not the only solution. In the words of Chen Danyan: Scaling Law is a "passing fairy," without it, AI cannot take off, but it has merely passed through the development path of AI and its future is not necessarily tied to it. The industry has also become aware of this issue, and the technical path for large models is currently showing a trend of distinct divergence: On one hand, there are the die-hard believers in "more power leads to miracles," with parameters scaling from tens of billions to hundreds of billions, and then to trillions, while training costs also surge exponentially; On the other hand, the Agentic assessment system, represented by Harness, has quickly gained popularity, with an increasing number of experts and scholars explicitly advocating for "less is more". To paraphrase Wang Yangming, the unity of knowledge and action means higher-quality action is what truly matters. StartLux offers another example of simplifying models. This can also explain why StartLux insists on local models. Once the model is on a PC, for long-term tasks, its cost structure can shift from continuously accumulating cloud-based token fees to more controllable device computing power and electricity consumption, while the model can also better understand the user's long context, achieving more personalized goals. It's not just StartLux, as Meta, Google, and NVIDIA have also recently been increasing their investment in local small model development. However, most of these are still in the experimental stage or cater to niche groups, with only one company focusing on local model commercialization. StartLux also stated that they will steadily advance their own foundation model training and explore new architectures such as diffusion-based language models, and as long as they are on the right path, the future is promising. StartLux has taken over the local model, and since this is a non-consensus route, those at the helm need to be two types of people: those who dare to take bets and those who can deliver results. StartLux's all-star team exemplifies this, pairing entrepreneurs with scientists in a formidable alliance. StartLux founder Chen Danyan CEO Chen Danyan had previously been briefly introduced as a serial entrepreneur and one of China's first-generation programmers, who rose to fame around the same time as Zhang Xiaolong and Lei Jun, and started his business ventures in the same era as Ma Yun and Ma Huateng. He was one of the first people in China to introduce the concept of "shared software" and co-founded Shanda Network with his brother Chen Tianqiao, as well as the Shanda Innovation Institute, a cradle of internet innovation, and was also instrumental in incubating the nationally popular product "WiFi Master Key". It can be said that he is extremely familiar with Chinese market users and products, and local models are also his comfort zone. StartLux co-founder Guo Quanwei The person responsible for implementing the technical roadmap is StartLux co-founder and CTO, Guo Quanwei. Kuo Chuan-wei holds a Ph.D. in Computer Science and Engineering from National Yang Ming Chiao Tung University, with research areas covering locally deployed large models, Agentic AI, AI for Science, AI for Finance, and privacy-preserving machine learning. Before joining StartLux, he served as the Chief Algorithm Scientist at AI Science company Huanliang Technology, and earlier worked at the Industrial Technology Research Institute of Taiwan, where he developed data privacy, data de-identification, and privacy-preserving machine learning. He is also a recipient of the 2024 TAAI Best Paper Award and holds data-privacy-related invention patents as the first named inventor. Another co-founder of StartLux, Luo Yongxiang, is the former Managing Director of Morgan Stanley Asia. Chen Danyan understands products and users, Guo Quanwei has long studied local models, Agents, and data privacy, and Luo Yongxiang is in charge of marketing and investment financing. This combination is highly suited to StartLux and will also help drive StartLux's long-term development. As for what the team ultimately wants to deliver, it's not just a set of model weights. In StartLux's vision, local intelligent solutions should be deployable with one click, similar to installing Office. Users do not need to understand quantization, GPU memory configuration, and inference frameworks, nor do they need to optimize them repeatedly themselves. The team currently plans to launch its first-generation local smart solution for enterprise and individual users within the year. In this light, the emergence of StartLux is by no means just "another player" entering the scene. It also represents that, following the emergence of cloud-based model companies like DeepSeek and Kimi, domestic local models are also starting to fill in the gaps. From a rising star in the intelligent era to circling back to the first-generation programmers of the internet era, the fundamental paradigm of models is shifting gears—but through it all, Chinese companies have kept passing the baton and pushing forward. Official website link: https://startlux.com/