🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 15, 2026
Agents

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

First page
Salesforce Koa: An Enterprise Language Model for Agentic Tool Use
The curator’s take

Salesforce post-trains the open-weight Nemotron-3-Super-120B with GRPO to produce Salesforce Koa, an enterprise model for agentic tool use whose training tasks and rewards are generated from the same Agent Script specifications that configure Agentforce agents.

Ask this paper

Key points
01

Specification to reward: Workflow specifications are expanded into persona-conditioned multi-turn tasks, with task-resolution rewards grounded in successful tool calls for data-dependent requests. Public tool-use domains use synthesized workflow structure through the same machinery.

02

Public benchmarks: On Tau2Bench Koa reaches a task-weighted average of 69.41 against 68.64 for its base and 54.48 for GPT-4.1, below Claude Opus 4.8 (74.00) and GPT-5.5 (83.99). On BFCL it scores 66.63% against 64.73% for the base and 53.96% for GPT-4.1.

03

Enterprise benchmark: On CRM Bench it scores 0.86 overall, close to Opus 4.8 (0.87), with function-call accuracy rising from 0.71 to 0.77 over the base.

04

Stated limitation: The base model was already RL post-trained, so the finding that RL helps multi-turn tool use far more than SFT may not hold when starting from a pre-RL checkpoint. No customer data was used.

Abstract

We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance. Its distinctive component is a simulation-to-reward pipeline that expands workflow specifications into persona-conditioned multi-turn tasks with task-resolution rewards grounded in successful tool use for data-dependent requests. For enterprise domains, these specifications are written in Agent Script, Salesforce's declarative language for building Agentforce agents; for public tool-use domains, we synthesize the workflow structure directly. The same simulation and grounded-reward machinery drives GRPO across both. Across public tool-use, agentic-reasoning, and enterprise Customer Relationship Management (CRM) benchmarks, Salesforce Koa improves over its open-weight base, with the clearest gains on multi-turn tool use, and surpasses a strong proprietary baseline while remaining below the strongest frontier models. These results show that specification-driven reinforcement learning is a practical path to specializing open-weight foundation models for enterprise agentic tasks.

Every Monday
Get next week’s papers.
Subscribe on Substack