HyperAIHyperAI

Command Palette

Search for a command to run...

MerchantBench: 전자상거래 운영에서 대규모 언어 모델 에이전트의 장기 일관성 평가를 위한 벤치마크

초록

대규모 언어 모델 에이전트는 자율적인 도구 사용자로서 점점 더 많이 평가되고 있지만, 대부분의 벤치마크는 즉각적인 성공 기준이 있는 제한된 과업에 초점을 맞춘다. 실제 환경에 배포하려면 장기적 일관성, 즉 축적된 증거에 따라 의사 결정을 조정하면서도 확장된 시간 범위에 걸쳐 목적 지향적 행동을 유지하는 능력이 필요하다. 이러한 능력을 평가하려면 행동이 미래의 선택을 제약하고, 피드백이 이질적인 지연 시간을 두고 도착하며, 일관성 없는 행동이 측정 가능한 누적 효과를 발생시키는 지속적인 환경이 요구된다. 판매자 측면의 전자상거래는 제품 소싱, 리스팅 및 가격 통제, 현금 흐름 관리, 혼합 지연 피드백 적응에 걸친 반복적이고 상호 의존적인 의사 결정을 통해 이러한 평가에 적합한 환경을 제공한다. 우리는 98,843개의 실제 전자상거래 제품 기록을 기반으로 하고 에이전트 상호작용을 위한 26개의 도구를 갖춘 365일 단위의 주문 수준 시뮬레이션인 MerchantBench를 소개한다. MerchantBench는 즉시 관찰 가능한 상류 공급업체 이벤트와 지연된 하류 주문 결과를 결합하여, 에이전트가 개별 주문 생애 주기를 추적하고 이전 결정을 재검토하도록 요구한다. 우리는 두 가지 에이전트 프레임워크 하에서 8개의 대규모 언어 모델을 각각 365일의 시뮬레이션 기간 동안 총 48회 실행하며 평가했다. 그 결과, 최신 대규모 언어 모델조차 인간 참가자와 상당한 격차를 보였으며, 최고 성능의 대규모 언어 모델 구성은 인간 참가자가 달성한 평균 최종 순자산의 27.3%에 그쳤다. 우리의 코드는 https://github.com/KhanCold/merchantbench에서 확인할 수 있다.

One-sentence Summary

Researchers from Zhejiang University, Peking University, and Fudan University introduce MerchantBench, a 365-day e-commerce simulation leveraging 98,843 real product records and 26 tools to assess LLM agents’ long-term coherence across sequential decisions such as product sourcing, pricing, cash-flow management, and mixed-latency feedback, finding that the best LLM attains only 27.3%27.3\%27.3% of the mean final net assets of human participants.

Key Contributions

  • MerchantBench is a 365-day order-level simulation benchmark for evaluating long-term agent coherence in seller-side e-commerce, constructed from 98,843 real product records and 26 interaction tools.
  • The environment couples immediately observable upstream supplier events with delayed downstream order outcomes, requiring agents to track individual order lifecycles and adapt decisions across mixed-latency feedback.
  • Evaluating eight LLMs across two agent frameworks in 48 runs shows the best-performing LLM configuration attains only 27.3% of the mean final net assets of human participants, highlighting a substantial gap in sustained autonomous decision-making.

Introduction

Evaluating LLM agents over extended realistic timescales is essential for deployment in persistent commercial settings such as seller-side e‑commerce, where sustained strategy, adaptation, and goal maintenance are critical. Existing benchmarks largely focus on short‑horizon, single‑session tasks and do not capture the challenges of long‑term operational coherence, including deteriorating activity, premature goal abandonment, and poorly calibrated strategy shifts. The authors introduce MerchantBench, a 365‑day order‑level simulation grounded in 98,843 real product records, designed specifically to measure long‑term coherence across eight LLMs and two agent frameworks and to reveal the performance gap relative to a human baseline.

Experiment

After 365 simulated days, ReAct GPT-5.6 Sol led all models in net assets, GMV, profit margin, and Sustained Window Rate, while incurring the fewest fines and lowest anomaly rate. Claude Opus 4.8 demonstrated strong store reliability with the highest average rating, and GLM-5.2 processed the most orders but paid the highest fines and earned the lowest margin. Sustained Window Rate differed sharply, from near-perfect for ReAct to just 11% for Qwen3.7-Max, highlighting large disparities in long-term operational consistency. ReAct GPT-5.6 Sol achieved the best overall business performance, with the highest net assets, GMV, and profit margin, while keeping fines and anomalies low. Claude Opus 4.8 secured the top average store rating and a low anomaly rate, though its financial metrics trailed ReAct's. GLM-5.2 drove the highest order volume but suffered the largest fines and the lowest profit margin among all agents. Sustained Window Rate ranged from 99.4% (ReAct) down to 11.1% (Qwen3.7-Max), revealing extreme variation in long-horizon reliability. Anomaly rates were lowest for ReAct (10.7%) and highest for Qwen3.7-Max (16.1%), indicating inconsistent store operation across models.

Monthly net profit profiles show that the human baseline attains far higher profits than any Hermes model, with a pronounced mid-year surge exceeding 35k. Among the Hermes variants, some demonstrate moderate profitability with peaks around 8-9k, while others remain near or below 1k throughout the year, often declining in later months. The human baseline exhibits a sharp seasonal spike, reaching its highest profits in months 6-8 and peaking above 39k. Hermes models display wide performance diversity: top performers achieve peak profits around 8-9k, whereas the weakest models rarely exceed 2k and show a downward trend. Several Hermes models end the year with profits under 1k, in contrast to the human baseline's sustained higher earnings.

A 365-day simulated store experiment evaluated multiple LLM agents, with ReAct GPT-5.6 Sol leading in net assets, profit margin, and long-term reliability (99.4% sustained window rate), while other models showed trade-offs such as high order volumes coupled with extreme fines and low reliability. A separate human-versus-Hermes profit study found that human traders achieved far higher profits with a seasonal peak exceeding 39k, whereas Hermes models delivered modest peaks around 8-9k at best and often fell below 1k, underscoring a large performance gap.


AI로 AI 구축

아이디어에서 출시까지 — 무료 AI 코코딩, 즉시 사용 가능한 환경, 최적의 GPU 가격으로 AI 개발을 가속화하세요.

AI 협업 코딩
바로 사용 가능한 GPU
최적의 가격

HyperAI Newsletters

최신 정보 구독하기
한국 시간 매주 월요일 오전 9시 에 이번 주의 최신 업데이트를 메일로 발송합니다
이메일 서비스 제공: MailChimp