← All posts
Opinion·6 min read

Salesforce Will Charge You $2 a Resolution. It Won't Tell You Which Model Did It.

Agentforce Help Agent's new pay-per-resolution pricing is the cleanest example yet of AI billing that erases which model actually did the work. That makes blind, live model testing more necessary, not less.


Salesforce's Agentforce Help Agent reached general availability this month with a pricing model that sounds like relief: pay $2 when the AI resolves a support ticket, pay nothing if it doesn't. No token meter, no per-message anxiety, no invoice full of numbers nobody outside an ML team understands. Buy resolutions in blocks of 1,000 and stop thinking about the model underneath. That last part is the problem.

What "selling outcomes" actually buys you

The mechanics are specific enough to be worth stating plainly. A resolution costs a flat $2, sold in blocks of 1,000. A session window runs two hours for messaging channels and ten minutes for voice, and every action inside that window, however many questions get asked, counts as one resolution. If the customer asks for a human, or walks away unhappy, Salesforce doesn't charge. It's a genuinely well-built product. It's also being framed, correctly, as part of a broader industry pivot that a BigGo Finance analysis this week called moving "from selling tokens to selling outcomes."

The reasoning behind the pivot is real. Under straight token metering, an e-commerce operator cited in that same analysis paid for 5,000 tokens and got back "three to five lines" of usable output, and the piece cites research showing over 60% of enterprise users deliberately shorten their prompts to control cost, trading away exactly the reasoning depth they're paying for. Token pricing punishes exploration. Outcome pricing was built to fix that.

  • Agentforce Help Agent: $2 flat per resolution, sold in blocks of 1,000, GA July 2026
  • Session window: 2 hours for messaging, 10 minutes for voice, unlimited questions inside it
  • Cited failure mode of token billing: 5,000 tokens spent for 3-5 usable lines of output
  • Cited behavior change: over 60% of enterprise users shorten prompts to manage token cost

The receipt used to tell you something

Here is what a token-metered invoice still gives you, for all its flaws: a model name and a count. You can take that line item, go run the same model on the same kind of task somewhere you can watch it work, and decide for yourself whether it earned the money. That's the entire premise of running prompts across models live rather than trusting a vendor's summary of what happened, the same reason we test outputs live instead of screenshotting claims.

A per-resolution invoice gives you none of that. "Resolution: $2" is the whole line item, whether the ticket was a one-line FAQ answer a small cheap model could close in one turn, or a multi-step troubleshooting session that needed a frontier reasoning model and burned real compute to get there. Nothing in the pricing structure requires Salesforce, or anyone else selling outcomes, to disclose which model touched your ticket. That isn't an accusation that anyone is quietly downgrading customers to cheaper models behind a flat price. It's an observation about incentives: the bill no longer contains the information you'd need to even ask the question.

The dominant form of the future is likely to be tiered hybrid billing: the infrastructure layer maintains token pricing to ensure transparent recovery of computing costs; the application layer introduces outcome-oriented premiums, allowing AI service providers that genuinely solve problems to earn higher returns. - BigGo Finance, July 20, 2026

Compare it to a price cut you can actually audit

Meta's Muse Spark 1.1, priced this month at $1.25 per million input tokens and $4.25 per million output, against the $5-10 input and $30-50 output that OpenAI and Anthropic typically charge for comparable tiers, is a much bigger discount on paper. But it's still a legible discount. The model name is on the invoice. You can take "Muse Spark 1.1, $1.25/$4.25" straight into the coding arena, run it blind against Opus 4.8 or GPT-5.6 on the same prompt, and see with your own eyes whether the 75% cheaper model is actually worth 75% less for the work you need done. Cheap-but-legible and expensive-but-opaque are different problems, and outcome pricing is quietly choosing the second one for you.

Test before you sign, not after

The practical shift is this: as more vendors move billing away from tokens and toward outcomes, the burden of figuring out which model is doing your work moves entirely onto you, and it moves earlier, before the contract, not after. Once you're paying $2 a resolution in a block of 1,000, you have no per-ticket lever to say "route this class of question to a stronger model." Your only real leverage was upstream, at the point where you evaluated whether the underlying models were good enough in the first place.

A StepFun executive made a version of this same argument about benchmarks rather than pricing, and it applies just as well here.

We care far less about leaderboard scores and much more about whether our model actually works well across our own products, phones, cars and robots. The real test is just using it. - Yang Minghui, StepFun
Before you buy an outcome-priced product

Ask the vendor directly which model or models power the SLA, and whether that can change without notice. If they won't say, that's your answer. Then go run the models you suspect are involved through a blind arena yourself, on a task shaped like your actual workload, before the per-resolution or per-ticket price becomes the only number you ever see again.

None of this makes outcome pricing a bad idea. It genuinely fixes the metering-anxiety problem that made token bills a bad proxy for value. But it trades one opacity for another, and the fix for both is the same: don't take anyone's word, including ours, for which model is good at what. Watch the outputs. European and Polish readers following this same shift can track it in Polish at nowosci.ai.

Don’t take the post’s word for it

The arena runs every model’s real output live. Pick a challenge, go blind, and cast a vote that counts in the public tally.

Open the arena