Command Palette
Search for a command to run...
MerchantBench: تقييم وكلاء نماذج اللغة الكبيرة من حيث التماسك طويل الأمد في عمليات التجارة الإلكترونية
MerchantBench: تقييم وكلاء نماذج اللغة الكبيرة من حيث التماسك طويل الأمد في عمليات التجارة الإلكترونية
الملخص
يتم تقييم وكلاء نماذج اللغة الكبيرة بشكل متزايد كمستخدمين مستقلين للأدوات، ومع ذلك تركز معظم المعايير المرجعية على مهام محدودة ذات معايير نجاح فورية. غالبًا ما تتطلب عمليات النشر في العالم الحقيقي تماسكًا طويل الأمد، وهو القدرة على الحفاظ على سلوك هادف عبر آفاق زمنية ممتدة مع تكييف القرارات بناءً على الأدلة المتراكمة. يتطلب تقييم هذه القدرة بيئة مستمرة تقيد فيها الإجراءات الخيارات المستقبلية، وتصل فيها التغذية الراجعة بتأخيرات غير متجانسة، وينتج عن السلوك غير المتماسك تأثيرات تراكمية قابلة للقياس. يوفر جانب البائع في التجارة الإلكترونية بيئة مناسبة لهذا التقييم من خلال قرارات متكررة ومترابطة تشمل توفير المنتجات، والتحكم في الإدراج والتسعير، وإدارة التدفق النقدي، والتكيف مع التغذية الراجعة مختلطة زمن الوصول. نقدم MerchantBench، وهو محاكاة على مستوى الطلب تمتد لـ 365 يومًا تستند إلى 98,843 سجلًا حقيقيًا لمنتجات التجارة الإلكترونية ومزودة بـ 26 أداة لتفاعل الوكيل. يقرن MerchantBench بين أحداث الموردين العلنية القابلة للملاحظة الفورية ونتائج الطلبات النهائية المتأخرة، مما يتطلب من الوكلاء تتبع دورات حياة الطلبات الفردية وإعادة النظر في القرارات السابقة. نقوم بتقييم ثمانية نماذج لغوية كبيرة ضمن إطارين للوكيل في 48 تجربة، تمتد كل منها على 365 يومًا محاكيًا. تكشف نتائجنا عن فجوة كبيرة حتى بين أحدث نماذج اللغة الكبيرة والمشاركين البشر، حيث حقق أفضل تكوين لنموذج لغوي كبير 27.3% فقط من متوسط صافي الأصول النهائية الذي حققه المشاركون البشر. الكود الخاص بنا متاح على https://github.com/KhanCold/merchantbench.
One-sentence Summary
Researchers from Zhejiang University, Peking University, and Fudan University introduce MerchantBench, a 365-day e-commerce simulation leveraging 98,843 real product records and 26 tools to assess LLM agents’ long-term coherence across sequential decisions such as product sourcing, pricing, cash-flow management, and mixed-latency feedback, finding that the best LLM attains only 27.3% of the mean final net assets of human participants.
Key Contributions
- MerchantBench is a 365-day order-level simulation benchmark for evaluating long-term agent coherence in seller-side e-commerce, constructed from 98,843 real product records and 26 interaction tools.
- The environment couples immediately observable upstream supplier events with delayed downstream order outcomes, requiring agents to track individual order lifecycles and adapt decisions across mixed-latency feedback.
- Evaluating eight LLMs across two agent frameworks in 48 runs shows the best-performing LLM configuration attains only 27.3% of the mean final net assets of human participants, highlighting a substantial gap in sustained autonomous decision-making.
Introduction
Evaluating LLM agents over extended realistic timescales is essential for deployment in persistent commercial settings such as seller-side e‑commerce, where sustained strategy, adaptation, and goal maintenance are critical. Existing benchmarks largely focus on short‑horizon, single‑session tasks and do not capture the challenges of long‑term operational coherence, including deteriorating activity, premature goal abandonment, and poorly calibrated strategy shifts. The authors introduce MerchantBench, a 365‑day order‑level simulation grounded in 98,843 real product records, designed specifically to measure long‑term coherence across eight LLMs and two agent frameworks and to reveal the performance gap relative to a human baseline.
Experiment
After 365 simulated days, ReAct GPT-5.6 Sol led all models in net assets, GMV, profit margin, and Sustained Window Rate, while incurring the fewest fines and lowest anomaly rate. Claude Opus 4.8 demonstrated strong store reliability with the highest average rating, and GLM-5.2 processed the most orders but paid the highest fines and earned the lowest margin. Sustained Window Rate differed sharply, from near-perfect for ReAct to just 11% for Qwen3.7-Max, highlighting large disparities in long-term operational consistency. ReAct GPT-5.6 Sol achieved the best overall business performance, with the highest net assets, GMV, and profit margin, while keeping fines and anomalies low. Claude Opus 4.8 secured the top average store rating and a low anomaly rate, though its financial metrics trailed ReAct's. GLM-5.2 drove the highest order volume but suffered the largest fines and the lowest profit margin among all agents. Sustained Window Rate ranged from 99.4% (ReAct) down to 11.1% (Qwen3.7-Max), revealing extreme variation in long-horizon reliability. Anomaly rates were lowest for ReAct (10.7%) and highest for Qwen3.7-Max (16.1%), indicating inconsistent store operation across models.
Monthly net profit profiles show that the human baseline attains far higher profits than any Hermes model, with a pronounced mid-year surge exceeding 35k. Among the Hermes variants, some demonstrate moderate profitability with peaks around 8-9k, while others remain near or below 1k throughout the year, often declining in later months. The human baseline exhibits a sharp seasonal spike, reaching its highest profits in months 6-8 and peaking above 39k. Hermes models display wide performance diversity: top performers achieve peak profits around 8-9k, whereas the weakest models rarely exceed 2k and show a downward trend. Several Hermes models end the year with profits under 1k, in contrast to the human baseline's sustained higher earnings.
A 365-day simulated store experiment evaluated multiple LLM agents, with ReAct GPT-5.6 Sol leading in net assets, profit margin, and long-term reliability (99.4% sustained window rate), while other models showed trade-offs such as high order volumes coupled with extreme fines and low reliability. A separate human-versus-Hermes profit study found that human traders achieved far higher profits with a seasonal peak exceeding 39k, whereas Hermes models delivered modest peaks around 8-9k at best and often fell below 1k, underscoring a large performance gap.