For AI agents, the "environment" serves as both an examination hall and a training ground. This concept is rapidly reshaping how companies like TENCENT and Alibaba approach artificial intelligence development.
The shift from solving problems to defining them marks what some researchers call the "second half" of AI development. When general training methods can increasingly solve existing tasks faster, the more critical question becomes what AI should actually do and what constitutes genuine progress. Evaluation takes center stage in this new paradigm, with researchers questioning whether high scores on benchmarks can translate into real-world work capabilities.
In this evolving landscape, technical giants are investing heavily in creating sophisticated environments for their AI agents. These environments function as realistic testbeds where agents can interact with code repositories, use tools, observe state changes, and receive result validations鈥攁ll components essential for developing truly capable AI systems.
Recent academic publications reveal the scope of this investment. One major Chinese tech company has released six papers related to environment construction over the past six months, covering everything from executable task supply to long-duration tasks, mobile environments, automated generation with training loops, and environments that evolve alongside model capabilities. Another major player published research directly titled "Environment Scaling" at a prestigious academic conference, with their chief scientist listed among the authors.
Alibaba and BABA-W (09988) are not alone in this pursuit. ByteDance's Seed division, along with OpenAI and Anthropic across the Pacific, are all converging on the same battlefield. When debugging a software fault, the initial prompt might be a simple task description. But determining whether an agent genuinely possesses work capability requires more than a single test question鈥攊t requires building a world that closely mirrors real business operations.
The race to construct these training environments reveals significant challenges. Consider the quality filtration process: one team collected 47,678 public terminal environments, yet after screening for quality, solvability, and difficulty, only 127 passed鈥攁n elimination rate of 99.7 percent. A task whose requirements and detection standards don't align can be hacked by capable models, compromising the entire training process.
Quality issues plague even the most established players. An audit of 731 public tasks from a prominent benchmark found that roughly 30 percent had problems including overly strict tests, incomplete prompts, or insufficient test coverage. Another organization discovered that merely adjusting CPU and memory configurations in the run environment could swing results by six percentage points, demonstrating how environmental reliability directly impacts agent performance measurement.
Beyond quality, there is the question of timeliness. As model capabilities advance, previously challenging environments become trivial, and the effective learning signals diminish. Training environments have a shelf life. One research team's solution involves having environments evolve rather than building them from scratch, upgrading simple static pages into complex tasks requiring frontend-to-backend deployment and dynamic maintenance.
The economics of environment construction have already attracted venture capital. A company specializing in AI training and evaluation is reportedly in talks for a funding round of approximately $300 million at a $2.5 billion valuation, with a major tech firm potentially leading the investment. Another startup received $43 million in Series A funding to replicate real software environments like Salesforce and Slack as an "agent gym," only to be acquired four months later by a data platform company.
Incumbents are also adapting. One leading AI data company has launched a dedicated reinforcement learning environments division, noting that nearly half of its new training projects now involve such environments. The demand extends beyond training into actual business deployment, where security sandboxes, permission controls, authorization mechanisms, and execution environments at the harness and agent framework layers must be developed alongside the agents themselves.
The competitive landscape in China remains complex. TENCENT (00700) and TCEHY, along with Alibaba and other tech giants, are building environments in-house because these directly impact post-training effectiveness鈥攃ore model companies cannot easily outsource this critical capability. The vast sandbox infrastructure, immediate isolation requirements, and high-concurrency resets needed for environment operations also depend on cloud infrastructure that these companies already possess.
For companies aspiring to become the next Scale AI in the agent era, the path is narrow. Model teams understand what models need, cloud platforms understand machine scheduling, but who truly understands the intricate business workflows of financial institutions, law firms, or hospitals? The researchers who pioneered this direction predicted that AI's second half would shift from problem-solving to problem-defining. Now that major companies and capital are investing heavily in building worlds for their agents, this abstract judgment is becoming a concrete industrial reality. The answer continues to evolve, and the world itself is being rebuilt in the process.