Commerce Gym · In development

Train and evaluate agents against real commerce.

We are developing environments grounded in authorized commercial state and observable outcomes, so teams can evaluate how agents make buying decisions.

State

What is available?

Products, prices, inventory, policies and permissions.

Action

What can the agent do?

Compare options and choose supported commercial actions.

Outcome

What actually happened?

Completion, failure and observable commercial results.

Why real-world context matters

Commerce changes while an agent is deciding.

01

Changing state

Prices, availability and delivery conditions change. An agent needs to reason against the relevant state.

02

Business constraints

Policies, authorization and supported actions determine what an agent can actually do.

03

Meaningful outcomes

A plausible answer and a completed commerce task are different observations. Evaluation needs a defined result.

Our development direction

From merchant context to useful evaluation.

Current foundation

Merchant context & outcomes

Agentic Page provides the starting point in merchant content, AI discovery and available attribution.

In development

Structured trajectories

Connect authorized state, supported actions and observable results into defined records.

Research direction

Evaluation & training environments

Use these records to calibrate tasks and evaluate agent behavior under commercial constraints.

For agent and model teams

Help define the tasks worth measuring.

We welcome discussions about evaluation tasks, environment requirements and authorized data partnerships. Availability and collaboration scope are agreed individually.