E-commerce AI agents face a harsh real-world test: 107 tasks, top score just 56.1
Around 300 emails, 5,383 customs records and multiple interconnected business systems: this is not the kind of test paper familiar to chatbots. RealReplicaBench placed these materials into e-commerce workflows, requiring AI to select suppliers, label emails, draft replies and schedule calendar events, or organize customs data into a procurement control tower. After testing 107 real-world business tasks, none of the 13 models evaluated reached 60 points, and the highest score was just 56.1 points.
The top score came from Claude Opus 5. Within the Accio Work execution framework, it passed 66/107 tasks; on OpenClaw, it passed 60/107, and on PI, 65/107. These results show that the model itself is not the only variable: how an execution framework connects tools and manages state can also determine whether a task reaches completion. Accio Work is an AI workspace under Alibaba, designed to help merchants centrally manage and operate stores across multiple platforms.
The harshest aspect of this evaluation is its refusal to award points for being “more or less complete.” Whether the right supplier is selected can affect procurement; whether the price is calculated correctly can affect product listing; and whether the logistics route meets the deadline can affect subsequent booking. Even if a task is 80% complete, it has still not produced a directly usable deliverable for the user if the remaining 20% requires manual intervention—and that 20% may be the critical part.
For this reason, the test environment did more than provide prompts: it also replicated user interfaces, browser operations, command-line tools, APIs, file systems and backend states. A logistics fulfillment task required AI to enumerate routes for a shipment from China to the United States, taking ocean freight, trucking, final-mile delivery, insurance, customs clearance, Bond and platform fees into account; exclude options exceeding 30 days or with invalid port connections; and finally complete Booking and Shipment Verification. The basis for evaluation was not whether the model claimed to have finished, but whether results such as a Shipment ID were actually generated in the environment.
What does this mean for merchants? In the future, selecting e-commerce AI will require more than comparing models’ scores on questions. Merchants will also need to check whether an agent can consistently maintain context in real systems, correctly pass along dynamic IDs and hand results over to the next stage. RealReplicaBench’s 107 tasks were drawn from around 1.6 million complete conversations, 200,000 business execution traces and 2,000 high-value workflows. The team plans to continue adding tasks and use the evaluation for model assessment, training optimization and model routing. Merchants gain a screening tool closer to real-world work, while model teams must face a more concrete standard: not “can it do it?” but “can it deliver?”
Comentarios
Cargando el hilo…
Inicia sesión para escribir un comentario. Iniciar sesión