Both versions are aligned block by block, in reading order: title, the essentials, then paragraph by paragraph. Where the translation merged or split a paragraph, the matching cell stays empty — we never pair two passages by guesswork.
约300封邮件、5383条海关记录、多个彼此连接的业务系统,这不是聊天机器人熟悉的考卷。RealReplicaBench把这些材料放进电商工作流,要求AI完成供应商选择、邮件标签、回复草稿和日历安排,或把海关数据整理成采购控制塔。经过107个真实商业任务测试,参评的13个模型没有一个达到60分,最高分也只有56.1分。
Around 300 emails, 5,383 customs records and multiple interconnected business systems: this is not the kind of test paper familiar to chatbots. RealReplicaBench placed these materials into e-commerce workflows, requiring AI to select suppliers, label emails, draft replies and schedule calendar events, or organize customs data into a procurement control tower. After testing 107 real-world business tasks, none of the 13 models evaluated reached 60 points, and the highest score was just 56.1 points.
最高分来自Claude Opus 5。在Accio Work执行框架中,它通过了66/107项任务;在OpenClaw上是60/107,在PI上是65/107。这组结果显示,模型本身并不是唯一变量:执行框架如何连接工具、管理状态,也会改变任务能否走到终点。Accio Work是阿里旗下的AI工作台,目标是帮助商家统一管理和操作多个平台店铺。
The top score came from Claude Opus 5. Within the Accio Work execution framework, it passed 66/107 tasks; on OpenClaw, it passed 60/107, and on PI, 65/107. These results show that the model itself is not the only variable: how an execution framework connects tools and manages state can also determine whether a task reaches completion. Accio Work is an AI workspace under Alibaba, designed to help merchants centrally manage and operate stores across multiple platforms.
这套评测最严厉的地方,是它拒绝给“差不多完成”计分。供应商是否选对,会影响采购;价格是否算对,会影响商品发布;物流路线是否满足时限,会影响后续订舱。哪怕任务已经完成80%,只要剩下的20%仍需要用户手动处理,而且这20%可能正是关键环节,对用户来说就还没有形成可直接使用的交付。
The harshest aspect of this evaluation is its refusal to award points for being “more or less complete.” Whether the right supplier is selected can affect procurement; whether the price is calculated correctly can affect product listing; and whether the logistics route meets the deadline can affect subsequent booking. Even if a task is 80% complete, it has still not produced a directly usable deliverable for the user if the remaining 20% requires manual intervention—and that 20% may be the critical part.
因此,测试环境不只给出题目,还复刻了用户界面、浏览器操作、命令行工具、API、文件系统和后台状态。物流履约任务要求AI为一票从中国到美国的运单枚举路线,把海运、拖车、尾程、保险、清关、Bond和平台费纳入考虑,排除超过30天或港口衔接不成立的方案,最后完成Booking和Shipment Verification。判定依据不是模型说自己做完了,而是环境里是否真的生成了Shipment ID等结果。
For this reason, the test environment did more than provide prompts: it also replicated user interfaces, browser operations, command-line tools, APIs, file systems and backend states. A logistics fulfillment task required AI to enumerate routes for a shipment from China to the United States, taking ocean freight, trucking, final-mile delivery, insurance, customs clearance, Bond and platform fees into account; exclude options exceeding 30 days or with invalid port connections; and finally complete Booking and Shipment Verification. The basis for evaluation was not whether the model claimed to have finished, but whether results such as a Shipment ID were actually generated in the environment.
这对商家意味着什么?今后选择电商AI,不能只比较模型答题分数,还要检查它能否在真实系统里持续保持上下文、正确传递动态ID,并把结果交给下一环节。RealReplicaBench的107项任务来自约160万次完整对话、20万条商业执行轨迹和2000条高价值工作流;团队计划继续增加任务,并将评测用于模型评估、训练优化和模型路由。商家得到的是更接近实际工作的筛选工具,模型团队则必须面对一个更具体的标准:不是“会不会做”,而是“能不能交付”。
What does this mean for merchants? In the future, selecting e-commerce AI will require more than comparing models’ scores on questions. Merchants will also need to check whether an agent can consistently maintain context in real systems, correctly pass along dynamic IDs and hand results over to the next stage. RealReplicaBench’s 107 tasks were drawn from around 1.6 million complete conversations, 200,000 business execution traces and 2,000 high-value workflows. The team plans to continue adding tasks and use the evaluation for model assessment, training optimization and model routing. Merchants gain a screening tool closer to real-world work, while model teams must face a more concrete standard: not “can it do it?” but “can it deliver?”