🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 4, 2026
Reinforcement Learning · Agents

Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets

First page
Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
The curator’s take

Shuze Daniel Liu (MIT), Claire Chen (Caltech), Jiuqi Wang (UVA), David Simchi-Levi (MIT) and Thorsten Joachims (Cornell) train a 30B LLM seller with RL to negotiate a catalog of substitutable products with several buyers at once under a shared turn budget, and report that it beats much larger frontier models on profit.

Ask this paper

Key points
01

Setting. One seller negotiates with multiple buyers over substitutable items. Buyers have private valuations and buy at most one item, and a global limit on dialogue turns forces the seller to choose whom to engage and what prices to offer.

02

Formulation. The task is cast as a POMDP with a four-part structured message protocol, so offers, buyer choice and matching can be optimized with RL on top of natural-language dialogue.

03

Training. Qwen3-30B-A3B-Instruct plays both sides. The untrained seller breaks the protocol in 43.3% of turns; RL brings this to 1.3% within training.

04

Results. The trained 30B seller leads every main economic metric, including reward and surplus extraction, over GPT-5.4, DeepSeek-V4-Pro and Kimi-K2.6 with thinking enabled, without using a hidden reasoning mode.

05

Behavior. GPT-5.4 with high reasoning closes the most deals but extracts less surplus. The trained seller offers to the best buyer-item pairs more often instead of accepting the first profitable offer.

Abstract

Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, where a seller concurrently negotiates a catalog of substitutable assets across a pool of independent buyers. Buyers hold private, heterogeneous valuations across products, and each can purchase at most one item. Facing limits on total communication turns, the seller must dynamically match buyers with the most profitable products considering their private valuations, while strategically allocating its limited interaction budget toward combinations of greater potential value. We formalize this problem as a Partially Observable Markov Decision Process using a structured, four-part message protocol that maps natural language into a parsable and regulated decision space. Using this formalization, we design a post-training method using Reinforcement Learning from Verifiable Rewards (RLVR). To evaluate this framework, we construct a multidimensional metric suite that quantifies constraint adherence, seller surplus extraction, and allocation quality. Our trained seller agent learns to match limited inventory to buyers more effectively, matching or outperforming trillion-parameter frontier models in both seller surplus extraction and buyer-product allocation quality. Finally, these learned strategies generalize robustly to unseen market structures, correlated valuation distributions, and price ranges not encountered during training.

Every Monday
Get next week’s papers.
Subscribe on Substack