A policy document not too long, which was implemented today in Beijing, has incorporated a series of terms that only recently appeared in academic papers and technical jargon into official government documents. This document, "Several Measures for Accelerating the Development of Intelligent Agents in Beijing," consists of ten points, including Agentic AI, Harness Engineering, AI OS, FDE, OPC, Token Economy, TaaS, AaaS, RaaS, Token Factory, and AIP. Especially when "Harness Engineering" is placed as the second point, the grand blueprint of the intelligent agent economy has, for the first time, gained a clear outline at the policy level.

image.png

The first point focuses on foundational models, with key wording being "continuously improve the actual task completion ability of large models." In recent years, large models were compared based on Benchmark scores, but when agents start working, they often only receive a vague goal, needing to complete the information themselves, break down the plan, adjust tools, read feedback, and modify the route until the result appears. There's a harsh calculation here: assuming an agent performs 100 steps continuously, with a 99% accuracy rate per step, the probability of getting everything right is only about 36.6%; even if the accuracy is increased to 99.9%, the chance of completing the whole process is only about 90.5%. Therefore, the practical value of an agent is more like a multiplication — model capability multiplied by long-term reliability, multiplied by environmental operability, and then multiplied by result verifiability. If any of these approaches zero, the total value also approaches zero. Frontier evaluations have also shifted in this direction, such as METR's Task Completion Time Horizon, which directly tests how long it takes for an agent to independently complete a task that would otherwise require a human expert, which also explains why code became the first area where agents exploded: it is highly digital, rollback-able, and verifiable.

The second point translates "Harness Engineering" as "Harness Engineering," requiring optimization around context engineering, task persistence, multi-agent collaboration, and system scalability to build a common foundation. The model provides potential capabilities, but whether these can be consistently delivered depends on the entire surrounding system — how to select context, how to adjust tools, how to break tasks down, how to control permissions, how to store states, how to retry after failure, and who verifies the results. These together form the "Harness." A paper from April this year, "Agentic Harness Engineering," demonstrated that by keeping the base model unchanged and only iterating on tools, middleware, long-term memory, and observable systems outside the model, after ten rounds, the Pass@1 score on Terminal-Bench2 increased from 69.7% to 77.0%, and there was a 5.1 to 10.1 percentage point improvement in other model families, with the gains mainly coming from tools, middleware, and long-term memory rather than just changing the system prompt. The policy also mentions skill markets, software stores, and AIP interoperability standards — in the future, many software users will become agents, and the competition will not only be about App Store rankings, but also about how agents discover them and make them default skills.

The third point encourages restructuring the underlying architecture based on intelligent agents, with the core being "demand intelligence." Software intelligence is divided into three layers: functional intelligence adds a generation button to a specific node; process intelligence connects several steps in a fixed workflow; demand intelligence allows the user to simply state the desired outcome, and the system automatically determines the steps, calls models and skills, connects software, and finds approvals, taking full responsibility for the final result. This is reflected in the concept of "super software" — its strength lies not in having more functions, but in crossing existing software boundaries, understanding requirements comprehensively, and recombining distributed capabilities. This point specifically mentions FDE (Frontier Deployment Engineer): entering customer sites, breaking down processes, connecting systems, and cleaning data, then bringing daily new problems back to the product. Five criteria determine benchmark scenarios — task frequency, manual cost, digitization level, result verifiability, and error reversibility. The more of these conditions are met, the more likely real ROI will be achieved. This also explains why the policy chose science, healthcare, education, government affairs, manufacturing, and culture.

The fourth point moves from software to terminals, requiring deep embedding of agents into phones, glasses, earphones, wearables, robots, and cars, promoting the "five-in-one integration of chip, model, cloud, and terminal usage." Terminal agents need continuous environmental perception, user status understanding, and long-term context remembering, making decisions at the right time and executing through devices. Sensors allow them to perceive the world, device identity and personal data help them understand "who they are," while the operating system and actuators enable them to change the world. The edge handles low-latency perception, identity verification, private data, and frequent simple tasks, while the cloud manages complex planning and reasoning. Device-to-device task state sharing allows one task to begin on earphones, continue on a phone, and be completed on a car or computer. The document also includes new products that meet the criteria in home appliance trade-in and smart product purchase programs, directly stimulating demand side.

The fifth point is the most direct for ordinary people: supporting innovative entrepreneurship models represented by OPC (One-Person Company). Behind it is the loosening of enterprise boundaries — Coase explained in 1937 why companies exist because internal coordination is cheaper than repeatedly going to the market to find people, negotiate, sign contracts, and supervise. However, agents are simultaneously lowering execution costs and coordination costs, allowing knowledge work that previously required multiple positions — writing code, designing, handling customers, responding to inquiries, and managing reports — to be connected by one person using multiple agents, thus reducing the minimum viable scale of enterprises. The policy doesn't just issue a slogan but continues with flexible computing power, professional incubation, entrepreneurial guidance, technology finance, intellectual property, policy consultation, OPC communities, full-cycle service stations, supply-demand matching, and convenient registration, effectively moving internal company back-end capabilities to the social public service layer. For individuals who want to seize opportunities, what they truly need are three assets: industry judgment, customer trust, and a reusable agent system.

The sixth point discusses the Token economy. Agents perform dynamic tasks with uncertain length, context, reasoning rounds, number of tools, and retries, making it common to spend 2 billion Tokens without achieving anything. This makes Tokens suitable as a cost measurement unit but difficult as a value currency. While the policy mentions Token service quality evaluation and billing standards, it also encourages establishing an evaluation system based on intelligent quality, usage volume, and conversion efficiency, and writes the heaviest sentence in the entire document: "Encouraging innovation entities to shift from billing by Token consumption to value-based billing." Three business models correspond to three value levels: TaaS sells basic Tokens and inference supply, competing on price, speed, stability, and model quality; AaaS packages models, tools, and Harness into capable agents, selling one capability; RaaS sells results directly, with customers focusing only on whether the contract is signed, the code is fixed, or the lead is converted. The higher up, the further away from Tokens and closer to customer value, but RaaS becomes much more difficult, as suppliers must bear the responsibility of task failure, repeated execution, cost fluctuations, quality acceptance, and compensation. Afterward, agent companies should calculate a new account called Cost per Successful Task (cost per successful task), and then Value per Successful Task (value per successful task). The document also proposes including eligible Token new products in small and medium enterprise service coupons, exploring Token coupons and intelligent agent service coupons, using Token factories to solve supply, service coupons to activate demand, and value-based billing to maintain the cycle.

The seventh point discusses security. Chatbots produce content, but agents gain action rights, able to edit code, initiate payments, control devices, and communicate on behalf of users, expanding risks from "saying something wrong" to "doing something wrong." Agent security must address at least five issues simultaneously: incorrect understanding of goals, permission overreach, prompt injection manipulation, pollution of long-term memory, and rapid amplification of errors between multiple agents. Connecting to third-party tools and skills also means that any plugin or interface failure can infiltrate the core along the task chain. The document's security measures have moved from content review to runtime governance; in May this year, the Cyberspace Administration of China and other departments issued the "Opinions on the Standardized Application and Innovative Development of Intelligent Agents," proposing to define the boundaries of decisions limited to oneself, authorized decisions, and autonomous decisions. Categorized regulation may actually help the industry accelerate — low-risk tasks can be tested quickly and launched, while high-risk tasks require strict inspection and manual confirmation. Here, safety is both a constraint and a market access certificate.

The eighth point implements the "Galaxy Computing Corridor" project, building a new computing infrastructure for high-frequency, low-latency demands of intelligent agents, promoting the integration of computing networks with 5G-A, 6G, and F5G, and tapping into existing computing power, aggregating scattered computing power through "zero savings and lump sum." Training large models is a centralized big project, but agent inference is like a long-running production network, responding to massive tasks around the clock, with loads fluctuating dramatically due to real business. Future scheduling will not only involve GPUs, but also different chips, models, data permissions, and network conditions across clouds, edges, and ends. The value of Galaxy Computing Corridor is similar to building a computing logistics system for intelligent production, creating opportunities for computing routing, heterogeneous scheduling, and edge computing. The document also supports banks and insurance institutions in developing financial products to support the implementation of intelligent agents, as the most valuable asset of agent companies in their early stage is code, data, and contracts, while the most important new asset in the new era is called "task trajectory" — a complete chain of goals, plans, actions, observations, corrections, and results, which can be used for training models, optimizing Harness, discovering security vulnerabilities, and analyzing business processes.

The ninth point promotes open source and openness, proposing the construction of a high-level China-South Asian Cooperation Center for AI applications, researching export paths for large models and intelligent agents for each country, and building a globally influential open-source community. The intelligent agent era especially needs open source and protocols, as it is inherently interconnected, and the ecological network effect is extremely strong — once a new skill is integrated, all compatible protocol agents may gain this capability, and the more agents there are, the more developers are willing to provide skills, the more skills there are, the more useful agents become. Whoever's protocol is widely adopted controls the development entry and ecological voice. The document specifically mentions evaluating open-source contributions, as open source is truly expensive in terms of long-term maintenance and community operations. The export part faces more complex privacy, responsibility, and access issues because intelligent agents are deeply embedded in local economies, hence the proposal for "one country, one strategy."

The tenth point addresses guarantees, coordinating national and municipal fiscal funds, government investment funds, and market-oriented funds to support technical breakthroughs, common platforms, and demonstration applications that meet the criteria. This policy covers a wide range, spilling over from a single technology industry to scientific research, software, manufacturing, consumer goods, entrepreneurship, finance, regulation, talent, and international cooperation. The intelligent agent era represents a direct capability reassessment for every individual and company: most jobs will not disappear overnight, but more realistically, they will be broken down into dozens of tasks, with repetitive, procedural, and easily verifiable parts being taken over by agents first; those who can only complete intermediate steps, cannot judge results, and cannot take responsibility for final delivery will face increasing pressure. Opportunities lie precisely in the redistribution of tasks and the revaluation of capabilities — in the past, it took capital, teams, and complex management to create a company, but now, with a mature agent system, it's possible to start from a very small real task and scale up. However, agents won't automatically bring success — they will only amplify a person's goals, judgments, expertise, and execution system together.