Article

Anthropic Project Swap: Test User Preferences Before Agents Negotiate

What Anthropic’s book-trading experiment reveals about preference capture, negotiation metrics and the checks to run before delegating exchanges to AI agents.

Editorial illustration for Anthropic Project Swap: Test User Preferences Before Agents Negotiate: a controlled task flows from input to output. Not documentary evidence.

Anthropic published Project Swap on September 24, 2026: an internal book exchange involving 201 employees and their Claude agents. For teams designing agents that buy, match or negotiate for users, the research offers a reason to test preference capture separately from bargaining ability.

This is an analysis of that research, not a new marketplace launch. The practical question is whether an agent can recognize a deal its user would actually want—not just complete an exchange.

Check the preference model before the negotiating model

Participants described their reading tastes in a short intake conversation. Fable 5 inferred their book rankings. Separately, 188 participants supplied their own rankings, collected after trading but before allocations were revealed and withheld from the agents. The main analysis excluded Dublin’s three-person pool. The methods appendix explains the sampling and timing.

Anthropic reports 61% agreement on the ordering of book pairs between inferred and human rankings, against 50% for random guessing. That measures preference prediction, not transaction success. A completed trade can satisfy the agent’s ranking while disappointing the person it represents.

What the 85% shortfall means

The report compares three allocations using scores based on participants’ preferences. Its Figure 4 averages over the 188 ranking respondents; the decentralized result draws on 80 neutral-instruction simulation runs, not just the live exchange.

Allocation

Preference information used

Reported score on human preferences

Best feasible assignment

Participants’ rankings

0.89

Optimized assignment

Claude’s inferred rankings

0.60

Agent-to-agent bargaining

Claude’s inferred rankings

0.55

These are normalized rank scores: first choice scores 1 and last choice 0. The calculation assumes equally spaced value between ranks and imputes scores for books a participant did not rank. They are not monetary returns or probabilities of success. The report’s scoring notes describe these assumptions.

Calculation from the rounded reported scores: the preference-representation gap is 0.89 − 0.60 = 0.29; the total gap is 0.89 − 0.55 = 0.34. Dividing 0.29 by 0.34 gives approximately 85%, reproducing the authors’ decomposition.

That does not mean 85% of trades failed, or that a better questionnaire would recover all the missing value. It identifies where this experiment’s constructed score fell short. For a pilot team, it is a diagnostic: improving negotiation may have limited payoff while the objective remains inaccurate.

Better bargaining is not the same as better representation

The trading-model comparisons reused the same inferred rankings. In footnote 15, Anthropic reports an adjusted Haiku 4.5–Opus 4.8 gap of 0.12 when scoring against Claude’s rankings, but only 0.01 using people’s rankings in the same regression specification.

The implication is to keep two evaluations: how well the agent pursues its supplied objective, and how well that objective matches independently collected user choices. A gain on the first does not establish a gain on the second. This experiment is not a current-model buying guide or a matched-cost production comparison.

Four checks before delegating a real exchange

The following is a proposed pilot framework, not a tested intervention from Project Swap. Start with user preference review before granting transaction authority.

  1. Test choices, not just the profile summary. Ask the user to judge representative options and unacceptable outcomes. Use some examples to correct the intake, then separate held-out choices to assess it. Record disagreements and hard-limit violations. Let the user revise the profile or opt out; 61% is not a universal passing threshold.

  2. Make acceptable compromises explicit. Separate negotiable preferences from limits the agent must not cross. Include a no-deal example and ask whether helping another participant may justify a worse personal outcome. Swap’s prompt treated keeping the starting book as failure; that is an experimental rule, not a sensible default for every purchase or exchange. See the published prompts.

  3. Compare mechanisms on the same inputs. Where centralized allocation is feasible, compare it with agent bargaining using identical preference profiles. Change intake quality in a separate comparison. Score outcomes against user choices, including the worst-served users, constraint violations and abstentions, alongside cost and elapsed time.

  4. Count fulfillment separately. Define evidence for agreement, execution and delivery. Keep an access-controlled record connecting the preference version, model and prompt, proposal, authorization and final outcome. Assign someone to resolve incomplete exchanges; an accepted proposal alone should not count as a delivered benefit.

For intake design, compare free-form briefing with targeted follow-up questions about dislikes, context and compromises. The ICLR 2025 GATE paper reports benefits from interactive preference elicitation in recommendation and email-validation tasks. That motivates testing the interface; it does not establish an improvement for Swap or your own market.

Put transaction rules in the application

Swap’s transaction design required all agents in a multi-party exchange to accept before holdings changed together. A proposal became invalid if the relevant holdings changed first. Those mechanics are distinct from persuasive messages on the trading floor.

For a pilot, enforce participant identity, delegated authority, current holdings and accepted terms outside the negotiating model. Set admission rules and run limits explicitly. Claude’s current tool documentation says application code executes client-tool calls; that is where a team can check a proposed operation before committing it. Tool calling does not supply those checks automatically.

Give the user access to decision evidence without broadcasting their preference profile to counterparties. Define permitted disclosures and retention separately; the related agentic-privacy analysis explains how to map those data flows.

When to keep the pilot supervised

The study used employees, cooperative Claude agents and fixed market rules. Human rankings could be noisy, and some physical book deliveries were incomplete. These limitations leave open how the results transfer to unfamiliar counterparties or higher-stakes exchanges.

If users reject the agent’s sample choices, improve intake before tuning bargaining. If preferences are reliable but the agent performs poorly against them, investigate the model or negotiation mechanism. If agreements do not become fulfilled outcomes, fix execution and delivery before expanding delegation.

Where an operator can collect preferences and compute acceptable matches, start with a centralized baseline. Agent negotiation needs to justify its extra complexity—for example, by handling genuinely variable counterparty interactions. AI Chatbot, Workflow or Agent? Choose by the Task covers that architecture decision. Keep consequential commitments supervised until the pilot has evidence for both user fit and reliable completion.

Methodology: AI-assisted analysis of published research and documentation checked on October 7, 2026. Experimental results are Anthropic-reported, not independently reproduced. The arithmetic and pilot framework are analysis; no hands-on market test was conducted.